Product problem
Receipt extraction is useful only when downstream systems can distinguish validated data from a model guess. Receipty treats the model as an untrusted extractor and makes uncertainty visible through typed outcomes instead of silently accepting malformed or unsupported fields.
Engineering approach
The FastAPI service sends schema-constrained requests, validates returned data with Pydantic, and applies deterministic locale-aware parsing. The same service path powers a labelled extraction harness. Indexed receipts can then be searched through keyword, dense, hybrid, or hybrid-rerank retrieval before the question-answering endpoint returns cited source IDs or a not-found response.
Validation and limits
The extraction suite contains 60 curated images: 55 receipts and five non-receipts. Receipt classification, total, and date checks passed on all 60 examples, but that result describes this limited dataset rather than universal accuracy. The 69-question retrieval suite contains 64 answerable questions and five not-found cases; recall and MRR apply only to the 64 questions with gold sources. Retrieval evaluation measures search quality, not answer correctness.
