All work

Independent project · LLM systems

Receipty — Evaluated LLM Extraction & RAG

A FastAPI receipt pipeline that treats model output as untrusted data, measures extraction quality, and grounds answers in indexed source receipts.

Role
Creator & AI Application Engineer
Period
2026
Platforms
API · Backend
Receipty app artwork showing a structured receipt inside a blue frame
Contribution

Built the extraction, validation, retrieval, grounded-answer, evaluation, and CI paths end to end.

60/60

receipt, total, and date checks

69 queries

labelled retrieval suite

80%

branch coverage gate

Product problem

Receipt extraction is useful only when downstream systems can distinguish validated data from a model guess. Receipty treats the model as an untrusted extractor and makes uncertainty visible through typed outcomes instead of silently accepting malformed or unsupported fields.

Engineering approach

The FastAPI service sends schema-constrained requests, validates returned data with Pydantic, and applies deterministic locale-aware parsing. The same service path powers a labelled extraction harness. Indexed receipts can then be searched through keyword, dense, hybrid, or hybrid-rerank retrieval before the question-answering endpoint returns cited source IDs or a not-found response.

Validation and limits

The extraction suite contains 60 curated images: 55 receipts and five non-receipts. Receipt classification, total, and date checks passed on all 60 examples, but that result describes this limited dataset rather than universal accuracy. The 69-question retrieval suite contains 64 answerable questions and five not-found cases; recall and MRR apply only to the 64 questions with gold sources. Retrieval evaluation measures search quality, not answer correctness.

Engineering challenges

  1. Prevent a vision model from silently turning uncertain receipt fields into confident application data.
  2. Compare retrieval strategies on labelled questions without confusing search quality with answer quality.
  3. Keep evaluation results reproducible across prompt, model, code, latency, token, and cost changes.

Key decisions

  1. Constrained extraction with strict JSON schemas, Pydantic validation, locale-aware parsing, and explicit success, review, and failure outcomes.
  2. Combined keyword and pgvector retrieval with lexical reranking, source citations, and a deliberate not-found response.
  3. Recorded prompt and commit hashes with model settings and operational measurements, then enforced deterministic quality checks in CI.