Back to walidsassi.com
Scoring a multimodal receipt extraction with code metrics and a model judge

Scoring What the Model Read: Metrics, a Judge, and 21 Receipts

Comparing store names with == fails almost every receipt. Here are the three code metrics and the model judge that turn a photo extraction into numbers, plus the results on 21 real receipts and the error patterns hiding behind them.

September 12, 2026 · 10 min · Walid Sassi
Reading a store receipt photo with Apple's multimodal on-device model

When the Model Can See: Reading Receipts with Multimodal Prompting

Multimodal prompting collapses OCR and understanding into one call. Here is the feature: a @Generable Receipt, one carefully shaped session with an image Attachment, and the photo dataset that has to be transcribed by hand before anything can be measured.

September 12, 2026 · 9 min · Walid Sassi

Mobile System Design Interview

Swift Academy Podcast · Series A mock-interview series with Artem Mirzabekian, Lead iOS Engineer. Real mobile system design problems, solved out loud the way a senior interview actually runs, with the roles swapping between episodes. Two episodes published. 🎧 Swift Academy Podcast · Series · 2 episodes · Requirements • APIs • Domain • Architecture “Most system design conversations stop at the box diagram. The interesting part starts after: what happens when the cursor is stale, the receipt URL has expired, and the user is on a train with no signal.” ...

September 5, 2026 · 6 min · Walid Sassi

Swift Academy Podcast, Season 1

September 5, 2026 · 0 min · Walid Sassi

Swift Academy Podcast, Season 2

September 5, 2026 · 0 min · Walid Sassi

Swift Academy Podcast, Season 3

September 5, 2026 · 0 min · Walid Sassi

Performance Is More Than Just Speed with Artem Mirzabekian and Bogdan Poplauschi

Swift Academy Podcast, Episode 2, Season 3 Performance Is More Than Just Speed Part two with Artem Mirzabekian and Bogdan Poplauschi, going underneath the abstractions: Swift Concurrency migration, the cost of the tools we use daily, and why application launch is an architecture problem. August 23, 2026 🎧 Swift Academy Podcast · Episode 2, Season 3 • Part 2 of 2 · Concurrency • Abstraction • App Launch Watch the Episode The Hidden Cost of the Tools We Use Daily Part one challenged the myths. This one goes underneath them, into Swift Concurrency, abstractions, architecture, application launch, frameworks and dependencies. ...

August 23, 2026 · 3 min · Walid Sassi
Refining a model-as-judge evaluator in Apple's Evaluations framework

Two Refinements That Sharpen a Model-as-Judge Evaluator

Two refinements to the NeighborhoodCoherence judge from part two: pass the metric name straight to ModelJudgeEvaluator to drop the ScoreDimension boilerplate, and shape what the judge actually reads with evaluationTarget and reference.

August 19, 2026 · 7 min · Walid Sassi
When code can't judge quality: model-as-judge evaluators in Swift

When Code Can't Judge Quality: Model-as-Judge Evaluators in Apple's Evaluations Framework

Code-based evaluators can only measure what code can measure. This companion article goes deep on ModelJudgeEvaluator: choosing a scoring scale, writing scoring levels that actually work, and wiring a model-as-judge into the TripPlanner evaluation from part one.

July 26, 2026 · 7 min · Walid Sassi
Apple's Evaluations framework: testing non-deterministic Foundation Models features

Testing the Untestable: A Deep Dive into Apple's Evaluations Framework

How to test generative, non-deterministic features with Apple’s Evaluations framework, the five steps of an evaluation, datasets and the Loader protocol, synthetic sample generation, evaluators and metrics, and the compiler gotchas hit along the way.

July 19, 2026 · 22 min · Walid Sassi