Back to walidsassi.com
Scoring a multimodal receipt extraction with code metrics and a model judge

Scoring What the Model Read: Metrics, a Judge, and 21 Receipts

Comparing store names with == fails almost every receipt. Here are the three code metrics and the model judge that turn a photo extraction into numbers, plus the results on 21 real receipts and the error patterns hiding behind them.

September 12, 2026 · 10 min · Walid Sassi
Reading a store receipt photo with Apple's multimodal on-device model

When the Model Can See: Reading Receipts with Multimodal Prompting

Multimodal prompting collapses OCR and understanding into one call. Here is the feature: a @Generable Receipt, one carefully shaped session with an image Attachment, and the photo dataset that has to be transcribed by hand before anything can be measured.

September 12, 2026 · 9 min · Walid Sassi