Skip to content
Radish
Home Guides Receipt scanner Spending tracker Benchmark Privacy Support

Transparent product evidence

A public benchmark for grocery receipt recognition.

Radish has preregistered a US grocery receipt-recognition benchmark covering 300–500 consented receipts from at least 10 retailers. The external one-shot study has not run, so no public field-accuracy percentage is available.

Last updated: August 16, 2026.

Read the frozen method See how receipt review works

Current status

Not yet measured

No external field result exists today. The study cannot read the hidden receipt set until every frozen development gate passes on the exact source revision and the external corpus completes consent, rights, privacy, and dual-truth review.

The current development baseline does not authorize holdout access. Publishing the protocol now makes later changes visible and prevents thresholds, metrics, or slices from being selected after seeing the field result.

Frozen scope

What the benchmark is designed to answer

Preregistered field-study population and measurement boundary
DimensionFrozen requirement
Population300–500 consented US grocery receipts from at least 10 retailers
Language and currencyEnglish-language receipts and USD totals
Product surfaceThe production iPhone app on a physical device
Measured momentThe exact pre-edit draft shown on the final Review screen
TruthIndependent dual review with exact asset and final-truth SHA-256 bindings
VisibilityUnseen locked holdout, consumed once only after development gates pass

Preregistered metrics

Accuracy, failure, correction, and time stay separate

Metrics frozen before the external holdout is inspected
MetricWhat it measures
Merchant exact matchWhether the reviewed merchant matches the verified receipt merchant after one frozen normalization rule
Local date exact matchWhether the proposed receipt date is exactly correct
Printed total exact matchWhether the proposed printed total matches to the cent
Line-item precisionHow many proposed grocery lines match verified receipt items
Line-item recallHow many verified receipt items appear in the proposed draft
Line-price errorMean absolute cent error for matched item occurrences
Printed-total reconciliationWhether the pre-edit draft is arithmetically balanced to the printed total
Hard-failure rateReceipts that do not reach a saveable Review draft
Corrections per receiptMedian visible field or line corrections before save, for saved receipts
Scan-to-save timeMedian seconds from capture start to durable save, for saved receipts

Required coverage

The aggregate cannot hide known hard receipt types

The same metrics are reported for preregistered slices with minimum receipt counts. They cover short, medium, long, and very-long receipts; multi-page receipts; clean, crumpled, faded, glare, shadow, cropped, and low-resolution images; and receipts containing discounts, weighted items, returns, tax, deposits, or fees.

A valid external corpus must meet every slice floor before scoring. Missing a difficult slice is a contract failure, not permission to publish a more favorable aggregate.

One-shot method

The runner fails closed before it can inspect private receipts

  1. Freeze the study population, metrics, slices, scoring rules, and producer bindings.
  2. Close every frozen synthetic development gate on one clean exact source revision.
  3. Validate active consent, rights, privacy review, and two independent truth reviews for each external receipt.
  4. Bind the app version, build, source, model, prompt, schema, scorer, corpus, gate policy, production path, physical-device evidence, and study period by digest.
  5. Create an irreversible consumption record before opening either private manifest or measurement file.
  6. Score the exact pre-edit Review drafts and export a new aggregate-only report. Any failure after consumption requires a newly reviewed holdout generation; it cannot be retried against the seen generation.

The benchmark adds no authentication bypass. Physical-device evidence must use the production app and its real protected receipt-processing path.

Publication contract

Counts first, percentages second

Every published rate must include its numerator, denominator, measured value, and Wilson 95% confidence interval. The report also carries the exact receipt, retailer, page, and item counts, the preregistered slice counts, the mean line-price error, and the saved-receipt medians.

Only aggregate and preregistered slice results can be published. The public report type cannot encode case IDs, receipt text, retailer identities, predictions, truth, timing records, correction records, image paths, or receipt bytes.

A favorable number is not enough. If source identity, development readiness, consent, corpus coverage, private-file integrity, producer binding, or one-shot state is invalid, no field result is eligible for publication or marketing.

Interpretation limits

What a future result will—and will not—mean

  • It will estimate performance for the frozen US grocery-receipt population and study period, not every receipt, language, country, currency, device, or future app version.
  • It will measure the proposed pre-edit Review draft. User corrections remain part of the workflow and are reported separately.
  • It will not establish financial, tax, accounting, nutrition, or universal store compatibility.
  • Provider, model, operating-system, camera, network, and receipt-quality changes can make a later version perform differently.
  • The consented private corpus will not become a downloadable public dataset.

Grocery receipt benchmark FAQ

Has Radish published a real-world receipt-recognition accuracy rate?

No. The external one-shot field study has not run, so Radish does not publish or market a field-accuracy percentage. This page publishes the frozen method and current status.

What will the grocery receipt benchmark measure?

It will measure exact merchant, date, and printed-total matches; line-item precision and recall; line-price error; total reconciliation; hard failures; corrections; and scan-to-save time at the pre-edit final Review stage.

Which receipts can enter the benchmark?

The frozen plan requires 300–500 consented US grocery receipts from at least 10 retailers, in English and USD, with documented rights, privacy review, and independent dual review of the final truth.

Why is the external holdout measured only once?

Single use prevents the team from repeatedly inspecting the same hidden receipts and tuning the product or scoring rules to improve that known result.

Will Radish publish receipt photos or per-receipt results?

No. Raw receipts, merchant-level identities, truth records, predictions, timings, and correction records stay outside Git, CI, and the public report. Only aggregate and preregistered slice results can be published.

Do synthetic development tests prove real-world accuracy?

No. Synthetic tests help find regressions and must pass before the hidden field study can run, but they are not a representative estimate of real-world grocery receipt recognition.

Evaluate the workflow yourself

Scan, review, and correct before you trust the record.

Free to download. Review and save one receipt before choosing a monthly or yearly plan for continued use. Requires iOS 18 or later. Optimized for English-language US grocery receipts and USD totals. Manual entry remains available.

Get Radish for iPhone Read the Privacy Policy

Related Radish pages

Home Grocery spending guides Grocery receipt scanner Grocery receipt to CSV Grocery spending tracker Grocery tracker for couples Track groceries without bank linking USDA grocery budget calculator Essential vs optional guide About Radish Privacy Policy Terms of Use Subscription terms Support

Radish publishes its product, evidence, privacy, legal, and support boundaries in public, canonical pages.