LAB/008 · Running · Self-directed, not client work

Structured Output Reliability

Four local models turning 200 receipts into JSON, with and without grammar-constrained decoding. The grammar makes every output valid. It does not make it right, and where data is missing it can make things worse.

Every system that puts a language model in front of other software depends on one thing: the model’s output must be data the next step can read. The usual fixes are “ask nicely for JSON” or “constrain the decoder to a schema”, which llama.cpp and most inference servers support. This Lab measures both on a realistic task, and looks at what the constraint does when the model has nothing to put in a field.

The test

200 shop receipts from CORD-v2 (CC BY 4.0), given to the model as text, one printed line per line, rebuilt from the dataset’s own word annotations. Using clean text instead of images separates the “text to JSON” step from reading the image (LAB/002 covers that). The target schema and the scoring are LAB/002’s: line items, subtotal, tax, service, discount, total, cash and change.

Two modes, same prompt, same four models from LAB/005:

  • Free: the prompt asks for JSON only.
  • Grammar: llama.cpp restricts decoding to the JSON schema, so the model cannot produce anything else.

Results

200 receipts as text → JSON, each model with free and grammar-constrained decoding.
Valid against schemaParseable JSONField accuracyLine items F1Receipt fully correctMedian time (s)
Qwen3-4B Instruct-2507 Q4_K_M · free92.5%100.0%87.6%84.5%44.5%1.06
Qwen3-4B Instruct-2507 Q4_K_M · grammar100.0%100.0%87.6%84.5%45.0%1.06
Qwen3-8B Q4_K_M · free94.0%97.5%87.2%77.2%47.0%1.50
Qwen3-8B Q4_K_M · grammar100.0%100.0%89.9%77.5%47.0%1.49
Qwen3-14B Q4_K_M · free96.0%100.0%92.3%87.1%56.5%2.85
Qwen3-14B Q4_K_M · grammar100.0%100.0%92.3%87.1%56.5%2.85
Phi-4-mini Q4_K_M · free88.5%93.0%77.7%73.3%36.0%0.83
Phi-4-mini Q4_K_M · grammar100.0%100.0%81.9%77.6%37.0%0.84
Valid against schema: exactly the required keys and value types · LAB/008 · measured 2026-10-08 · RTX 4070 12 GB
Output valid against the schemaScale 0–100% · free vs grammar-constrained

Qwen3-4B Instruct-2507 Q4_K_M

  1. free92.5%
  2. grammar100.0%

Qwen3-8B Q4_K_M

  1. free94.0%
  2. grammar100.0%

Qwen3-14B Q4_K_M

  1. free96.0%
  2. grammar100.0%

Phi-4-mini Q4_K_M

  1. free88.5%
  2. grammar100.0%

The grammar does its job: 100% valid output for every model. Without it, 4–11.5% of outputs broke the schema. Qwen3-8B produced invalid JSON 2.5% of the time. Phi-4-mini wrapped every answer in a Markdown code block, so none of its replies were bare JSON, and 7% could not be parsed even after stripping it. The constraint cost nothing in speed.

Field accuracyScale 0–100% · same fields, same receipts

Qwen3-4B Instruct-2507 Q4_K_M

  1. free87.6%
  2. grammar87.6%

Qwen3-8B Q4_K_M

  1. free87.2%
  2. grammar89.9%

Qwen3-14B Q4_K_M

  1. free92.3%
  2. grammar92.3%

Phi-4-mini Q4_K_M

  1. free77.7%
  2. grammar81.9%

Accuracy barely moved. For the two strongest models, field accuracy is identical with and without the grammar. It helped where the free output was broken (Phi-4-mini +4 points, Qwen3-8B +3), because a broken document scores zero. A valid structure is necessary, but the values inside it are as right or wrong as before.

What the grammar hides

The most common schema violation in free mode was honest: the model returned "price": null for a line with no price it could find, where the schema required a string. With the grammar on, null was not allowed, so the model had to write something:

Lines where free decoding said the price was missing (null), and what grammar-constrained decoding wrote there instead.
PlaceholderItem nameInvented amount
Qwen3-4B Instruct-2507 Q4_K_M1514
Qwen3-8B Q4_K_M300
Qwen3-14B Q4_K_M2200
Phi-4-mini Q4_K_M100
Placeholder: empty, "0" or the text "null". Counted only on receipts where both modes produced the same number of lines · LAB/008 · measured 2026-10-08 · RTX 4070 12 GB

Mostly placeholders: an empty string, "0", or the literal text "null". Some were worse: Qwen3-4B filled four lines with real-looking amounts and one with the item’s own name. A "0" price or an invented amount passes every type check downstream. The model’s “I don’t know” was turned into data that looks valid.

The fix is in the schema, not the model: make fields nullable wherever the information can legitimately be missing, and treat null as a signal for review rather than an error. Then run the same plain-code checks as in LAB/002 (totals add up, required fields present) on whatever comes out.

What this means in practice

  • Use grammar-constrained decoding for any model output that software will parse. It removes a whole class of failures for free.
  • Design the schema for missing data. A required field forces a value; allow null where real documents leave things out.
  • Validate values, not just structure. Schema validity says nothing about whether the total is right.
  • Measure per field, on your own documents, before automating.
Method and environment
dataset
CORD-v2 validation + test, 200 receipts (CC BY 4.0); input text rebuilt from the dataset's word annotations (no OCR errors)
task
receipt text → JSON with LAB/002's schema; scoring and normalisation from LAB/002
free
the prompt asks for JSON only; output parsed as-is, then leniently (code fences, surrounding text)
grammar
same prompt; llama.cpp constrains decoding to the JSON schema (json_schema request field)
prompt
Below is the text of a shop receipt, one printed line per line. Extract it as JSON with exactly these keys: {"items": [{"name": string, "quantity": integer or null, "price": string}], "subtotal": string or null, "tax": string or null, "service": string or null, "discount": string or null, "total": string or null, "cash": string or null, "change": string or null} One entry in "items" per purchased line; "price" is the line amount as printed. Copy amounts exactly as printed. Use null for anything not on the receipt. Return only the JSON. Receipt: {text}
models
llama.cpp b11461 (CUDA 13.4); greedy; Qwen3-8B/14B thinking off
environment
gpu: RTX 4070 12 GB · llama.cpp: b11461 (CUDA 13.4)
measured
2026-10-08

Limits

One document type (Indonesian receipts), clean text input rather than OCR output, one prompt, four models. The forced-value count compares line positions only on receipts where both modes produced the same number of lines, so it undercounts. The code is in labs/structured-output in this site’s repository and reuses LAB/002’s scoring.