AI Architecture Pattern 03: Structured Output And Validation

A response can pass every presence-and-type check and still be useless downstream. Pattern 03 of the AI Architecture Patterns series shows why structured output needs schema-level enforcement — grounded in Amazon Bedrock Structured Outputs — so an invalid value becomes impossible to generate, not just easy to catch.

Dispute DSP-2026-0042 comes in. The assistant classifies it and returns status: "under review" — a space, not an underscore. Pattern 02’s contract checks it on the way through: present ✓, a string ✓. Clean pass. The routing queue does a literal match against three exact values — open, under_review, resolved. "under review" matches none of them. No exception fires. No alert goes out. The dispute just sits unrouted, indistinguishable from one nobody has looked at yet.

This is the pattern that closes that gap: constraining the model so this exact failure can’t be generated in the first place.

Dispute status ‘under review’ passes Pattern 02’s shape contract but fails the routing queue’s exact enum match and silently routes nowhere, versus ‘under_review’ which passes both checks cleanly

[!NOTE] This post is part of a continuing 16-pattern series: a payment-support assistant at a fictional bank, starting simple and accumulating exactly the complexity a real one would, in the order a real one would need it. Read the series overview or clone the GitHub repo to follow along.

Where the first two patterns left off

Pattern 01 — Grounded RAG. Mike asks why his payment is pending. The assistant gives a fluent, completely fabricated answer. The fix: retrieve the real record first, generate the answer second. Left open: retrieval finding the right document doesn’t guarantee the model draws the right conclusion from it.

Pattern 02 — Prompt and Context Contract. A second team plugs their own data in and the app crashes with a raw KeyError. The fix: validate the shape of every input before rendering a prompt with it. Left open: a contract only checks shape — present, correctly typed — not whether a correctly-typed value is actually the right value.

“Isn’t this the same check Pattern 02 already does?”

Worth answering directly. If the contract already validates that status is present and is a string, what does this pattern add?

Presence and type is a shape check. This pattern is a value check, and they’re genuinely different questions. "under review" for status passes Pattern 02’s checks — it’s present, and it’s a string. Pattern 02 has no concept of “this string must be one of these exact allowed values,” and it was never designed to.

A field can be present, correctly typed, and still wrong. That’s the exact space this pattern occupies.

The series so far

PatternMike asks…What brokeOne-line fixWhat kind of check
01 — Grounded RAG“Why is my payment pending?”Model made up a fluent, wrong answerLook up the real record first, generate secondGroundedness
02 — Contract(a second team plugs in their data)Raw KeyError, two layers deepValidate the shape before you renderShape (present, correct type)
03 — Structured outputDispute status comes back as "under review"Correct shape, wrong exact value — silently unroutedConstrain the model so that value can’t be generatedValue (one of the allowed exact values)

Why this isn’t a one-off typo

OpenAI’s own launch numbers for Structured Outputs show plain JSON mode — generate, then parse, then hope — fails 2–5% of the time across 500K calls; constrained generation falls under 0.1%, a 20–50x gap, measured by the company that shipped it.

Ramp’s engineering blog is the second proof point, in production: before constraining output, their system automated only 1.5–3% of transaction classifications. Constraining the output to the real chart-of-accounts vocabulary took automation to close to 100%, at roughly 99% accuracy.

The mechanism — schema, grammar, token

Worth naming a wrong turn too: the tempting first guess is Bedrock Guardrails. It isn’t the right tool — Guardrails is a content-safety policy layer (PII redaction, denied topics, profanity filters) with zero schema-enforcement capability. The service that actually compiles a JSON Schema into a token-level grammar is Bedrock Structured Outputs.

Correction: the tentative pick Bedrock Guardrails is wrong — it’s content-safety policy only, zero schema capability. The correct service is Bedrock Structured Outputs

Two real options exist. Option A — Bedrock Structured Outputs, constrained decoding, what this pattern builds. Option B — Pydantic, generate then validate then reject if wrong. The difference is when the bad value gets stopped:

  1. The schema is a fixed three-field JSON object: dispute_reference (string), status (string, enum: ["open", "under_review", "resolved"]), confidence (number).
  2. Bedrock compiles that schema into a grammar — a formal description of every token sequence that could ever satisfy it. Compiled once, cached 24 hours.
  3. Generation is restricted to that grammar, token by token. "under review" — the space variant that broke DSP-2026-0042 — is never a reachable output. Not rejected after the fact. Never generated.

Option B’s post-hoc validation lets the bad value exist for a moment, then rejects it. Constrained decoding never makes it a reachable output in the first place.

Option A, Bedrock Structured Outputs: strongest guarantee, no retry loop, narrower schema subset. Option B, Pydantic post-hoc validation: full JSON Schema support, zero cloud dependency, detect-and-reject not prevent

The dispute-routing schema this pattern actually validates against:

schema = OutputSchema(
    schema_id="dispute-classification-v1",
    description="Structured classification of a payment dispute for the routing queue.",
    fields=(
        SchemaField("dispute_reference", "string",
            description="The dispute's reference ID, e.g. DSP-2026-0042."),
        SchemaField("status", "string",
            enum=("open", "under_review", "resolved"),
            description="Must be one of the routing queue's three known states — nothing else."),
        SchemaField("confidence", "number",
            description="The model's confidence in this classification, 0.0-1.0."),
    ),
)

status’s enum tuple is the whole pattern in one line — Pattern 02’s contract would accept any string here; this schema is the first point in the series that checks the actual value, not just presence and type.

On real Bedrock, constrained decoding makes an enum violation unreachable — it literally cannot be generated. In production this whole exchange sits behind a single Lambda: receives the request over API Gateway, calls Bedrock Converse with the schema attached as outputConfig, and only ever writes a schema-valid status to DynamoDB’s routing table.

Production pipeline: dispute classifier through API Gateway to Lambda orchestration to Bedrock Converse with outputConfig=schema, returning a schema-valid status, written to DynamoDB’s routing table

Which option do you actually need?

Your situationReach forWhy not the alternative
Building on Bedrock directly, output has a fixed, finite vocabularyBedrock Structured Outputs (chosen)Constrained decoding makes the bad value unreachable at generation time
Schema needs a feature outside Bedrock’s supported subset, or not on Bedrock at allPydantic post-hoc validationThe fallback this pattern’s own mock borrows for exactly this reason
Output is free-form text with no finite enum to constrain againstNeitherNo fixed vocabulary to compile into a grammar — a prompting problem, not a schema problem
Value is schema-valid but still the wrong judgmentNeither — that’s Pattern 04A schema guarantees the value is allowed, not that it’s correct
The data needed to classify correctly isn’t in the prompt yetNeither — that’s Pattern 05Constraining the output shape doesn’t help if the model never had the right input

When this pattern is the right call, and when it isn’t

Reach for a schema when the output has a fixed, finite vocabulary a downstream system does an exact match against — status enums, category labels, routing keys. Constrained decoding is strictly better than validate-then-reject: the bad value never exists, instead of existing for one round trip before being caught.

Don’t reach for it when the output is genuinely open-ended — a summary, an explanation, free-form prose. There’s no finite set of allowed strings to compile into a grammar. And it doesn’t solve either of the harder problems next door: it constrains the shape and set of a value, not whether that value is the correct judgment for the case, and it does nothing for the input side — a model reasoning from stale or missing data will produce a schema-valid answer that is still wrong.

What’s next

Pattern 04 keeps the same schema-valid output but asks a harder question: is it actually good? A classification can be present, correctly typed, and within the right enum — and still be the wrong classification. Pattern 04 introduces the first second LLM call in this series: a judge that evaluates the first call’s output, not just its shape.

[!TIP] Star the GitHub repo and follow along with the series.


🎯 Interview Prep — Pattern 03

Scenario-based questions an interviewer would ask about this pattern. Try answering out loud before expanding each one. Full guide, all patterns: Interview Guide →

Q1 Pattern 02's contract validation passed. The dispute still didn't route. What happened?

The routing queue does a literal enum match against three labels: open, under_review, resolved. The classification returned status: "still being looked at" — present, a string, exactly what Pattern 02’s contract required. It doesn’t match any label. The dispute sat unrouted with no error raised.

Root cause: Pattern 02’s contract has no mechanism to express “this string must be one of these exact values.” It validates presence and Python type only.

Fix: Pattern 03 adds an enum constraint on the status field via Bedrock Structured Outputs (outputConfig.textFormat.structure.jsonSchema). The schema prevents the model from generating any value outside the allowed set — constrained at generation time, not caught after.

Q2 What is the difference between input validation and output schema validation?

Input validation (Pattern 02): checks data coming into the prompt — are required fields present? Are they the right Python types? Runs before the model call.

Output schema validation (Pattern 03): checks the model’s response — is it the expected structure? Are enum-constrained fields within the allowed set? Runs after the model call, or with Bedrock Structured Outputs, is enforced during generation.

Both are necessary: input validation protects the model from bad data (cheap, fast rejection). Output schema validation protects downstream systems from a model that produced a valid-looking but unusable value.

Q3 What is constrained decoding, and how is it different from post-hoc validation?

Post-hoc validation: model generates freely, a validation function checks the result. If wrong, reject and optionally retry. Problems: the bad value exists long enough to be logged, each retry costs another model call.

Constrained decoding (Bedrock Structured Outputs): the schema is compiled into a grammar. Each generation step is restricted to tokens that remain valid given the grammar. An enum-violating value is physically impossible to generate — no retry needed.

Use constrained decoding when the schema is within Bedrock’s supported subset (basic types, enum, required). Use post-hoc Pydantic for schema features outside that subset (numeric ranges, recursive schemas).

Q4 What JSON Schema features does Bedrock Structured Outputs support, and what doesn't it support?

Supported: basic types (string, number, integer, boolean, array, object), enum on string fields, required fields list, some format values.

Not supported: numeric minimum/maximum, string minLength/maxLength, recursive schemas, full JSON Schema Draft 2020-12.

The hybrid pattern: Bedrock Structured Outputs for the enum constraint (strongest guarantee, zero retry cost) + Pydantic post-hoc for range checks that Bedrock’s subset can’t express (e.g. confidence in [0.0, 1.0]).

Q5 The model starts returning confidence: 'high' instead of 0.91. How do you catch this before it reaches the routing queue?

With Bedrock Structured Outputs (real AWS): the schema declares confidence as type: "number". Constrained decoding prevents the model from generating the string "high" in that position — only tokens that parse as a JSON number are valid.

Post-hoc Pydantic (belt-and-suspenders or fallback):

class DisputeClassification(BaseModel):
    status: Literal["open", "under_review", "resolved"]
    confidence: float = Field(ge=0.0, le=1.0)

isinstance("high", float) is False — caught before touching the routing queue.

Best practice: both. Bedrock Structured Outputs at generation; Pydantic on the Lambda side for range constraints Bedrock can’t express.

Q6 Show me the Converse API call shape for Bedrock Structured Outputs.
response = client.converse(
    modelId="amazon.nova-micro-v1:0",
    messages=[{"role": "user", "content": [{"text": prompt}]}],
    outputConfig={
        "textFormat": {
            "structure": {
                "jsonSchema": {
                    "schema": {
                        "type": "object",
                        "properties": {
                            "dispute_reference": {"type": "string"},
                            "status": {
                                "type": "string",
                                "enum": ["open", "under_review", "resolved"]
                            },
                            "confidence": {"type": "number"}
                        },
                        "required": ["dispute_reference", "status", "confidence"]
                    }
                }
            }
        }
    }
)

Bedrock compiles the schema into a grammar on the first call (~few hundred ms) and caches it for 24 hours. The response is a JSON string in response["output"]["message"]["content"][0]["text"]. Parse with json.loads().

Q7 You're reviewing a PR where the developer validates model output with a bare try/except json.loads(). What's wrong?

Three problems:

  1. JSON validity ≠ schema validity. json.loads() confirms the response is parseable JSON. It does not check field presence, types, or enum values. {"status": "still being looked at"} is valid JSON.

  2. No field-level check. The routing queue does a literal enum match. An invalid status goes straight to the queue — no error raised, dispute unrouted.

  3. Unstructured error response. {"error": "model did not return valid JSON"} gives no field name, no allowed values, no way to fix it fast.

Fix: add a Pydantic model with Literal on status — or use Bedrock Structured Outputs so the invalid value can’t be generated at all.

Q8 When would you choose Pydantic post-hoc validation instead of Bedrock Structured Outputs?

Choose Pydantic when:

  • Numeric range constraints — e.g. confidence in [0.0, 1.0]. Bedrock has no minimum/maximum support.
  • Recursive schemas — Bedrock doesn’t support self-referencing schemas.
  • Not on Bedrock — Structured Outputs is a Bedrock feature; Pydantic works with any model output.
  • Richer error messages — Pydantic ValidationError gives field-level detail including which value violated which constraint.

In practice, use both: Bedrock Structured Outputs for enum and type constraints (strongest guarantee); Pydantic as a belt-and-suspenders check for constraints outside Bedrock’s supported subset.

Q9 How do you test that an output schema catches all the ways a model could return the wrong value?

Build a test matrix for each field:

  • Happy path — each valid enum value
  • Enum violations — free-text values, wrong case ("OPEN"), empty string, injection attempt
  • Type violations — string where number expected (confidence: "high"), bool where float expected
  • Missing required fields — each required field absent individually; all absent

For each failure case, verify:

  • Response type is schema_violation (not an unhandled exception)
  • reason field names the exact field and the violated constraint

For Bedrock Structured Outputs specifically: unit-test the local mock’s schema logic. Trust the real Bedrock service for constrained decoding guarantees — test the Lambda’s error-handling path separately.

Q10 You need to validate 10,000 model responses per minute against this schema. What actually changes at that volume?

Less than people expect, because the enforcement moved from your code to Bedrock’s generation step. The schema compiles into a grammar once and Bedrock caches it for 24 hours — you’re not paying a re-compilation cost per request, just normal Converse call volume.

What genuinely needs attention at 10k rpm:

  • Provisioned Throughput or a quota increase — same constraint as any high-volume Bedrock workload, not specific to structured outputs.
  • The Pydantic fallback path, if you’re running the hybrid pattern — post-hoc validation for range checks (confidence in [0.0, 1.0]) runs in your own compute, so it scales with your Lambda/ECS concurrency, not with Bedrock.
  • Schema-violation logging volume — even a low violation rate (say 0.5%) is 50 structured error events per minute at this scale; make sure the CloudWatch metric filter grouping by violated field doesn’t get lost in noise.

The one-line answer: constrained decoding pushes the expensive part (guaranteeing valid output) into a cached, one-time grammar compile — it’s the retry elimination, not raw throughput, that’s the real cost story at scale. A system doing post-hoc-only validation at 10k rpm would be paying for retries on every rejected generation; this pattern mostly avoids that cost by construction.

Q11 How do you explain Bedrock Structured Outputs to a product manager who asks why disputes keep misrouting?

“Think of a dropdown menu on a form. You can’t type ‘still under investigation’ in a dropdown — you can only pick from the options the form gives you.

Bedrock Structured Outputs does the same thing for the AI: instead of a free-text field, we give it a dropdown with exactly three options — open, under_review, or resolved. It has to pick one. It can’t write anything else.

We’re switching the status field from a text box to a dropdown. Once that’s deployed, the misrouting stops.”

Technical one-liner: Bedrock Structured Outputs compiles the JSON Schema into a generation grammar — the invalid token can’t be sampled. The routing queue never sees it.


Have questions or feedback? Drop a comment below or connect on LinkedIn.

💬 Comments

← Back to all posts