Skip to content
Learn/

Structured Outputs

1 / 7

The output is a contract, not a suggestion

A model produces tokens. Your code needs a record: an order with a numeric quantity, a ticket with one of four priorities, a date that a database will accept. Somewhere between those two facts sits the single most common source of breakage in an LLM application — output that looks right to a human and fails JSON.parse.

There are three layers. Prompt and parse states the contract but relies on the model following it. Constrained decoding enforces a supported grammar or schema during generation. For a successful, complete generation within that supported subset, malformed structure can be ruled out; refusals, truncation, unsupported schema features, and provider failures still need handling. Validate and repair checks the result against your own contract and chooses what to do on failure.

Combine explicit instructions, constrained decoding where available, and validation. Structural enforcement does not remove semantic errors. A schema that says quantity: number may accept -4, and status: string accepts "probably shipped". Meaning and business policy remain yours.

prompt ──▶ model ──▶ raw output
                     │
        constrained  │  supported grammar enforces shape
        decoding     │  on a complete successful generation
                     ▼
                 validate  ── fails ──▶ repair prompt ──┐
                     │                                  │
                   passes                               │
                     ▼                        (another whole model call)
              typed record ──▶ your system  ◀───────────┘

A schema is a shape check, not a truth check. It can prove the field is an integer; it cannot prove it is the right integer.

Language Model is over capacity at 26%3.4 req/s failing·3.4% errors·96.6% available

4 components3 connections0:00

Traffic
100req/s
p50
2.36s
p99
5.79s
Errors
3.4%
Dropped
3.4req/s
Cost
$2.12M/mo