Skip to main content
PromptQuorumBuilt for humans. Structured for AI.
Home/Prompt Engineering/Best Tools for Structured Output and JSON Mode (2026)
Tools & Platforms

Best Tools for Structured Output and JSON Mode (2026)

·11 min read·By Hans Kuepper · Founder of PromptQuorum, multi-model AI dispatch tool · PromptQuorum

Seven tools dominate structured output in 2026: Instructor for Pydantic extraction, Outlines for constrained decoding, Pydantic AI for type-safe agents, BAML for schema-first prompt files, LangChain for unified APIs, Marvin for task-based extraction, and PromptQuorum for cross-model testing. Each solves a different workflow bottleneck.

Choose based on where your models run and which languages your team ships: Instructor and Pydantic AI for Python API workflows with retries and type safety; Outlines for guaranteed schema compliance on local models; BAML when the same schema must serve Python, TypeScript, and Go services; LangChain for teams already using chains or agents; Marvin for fast extract/cast/classify calls; PromptQuorum for consistency testing across GPT, Claude, and Gemini before production.

Best Tools for Structured Output and JSON Mode (2026)

Key Takeaways

  • Instructor is the most popular Python choice — Pydantic schemas, automatic retries, and official ports for TypeScript, Ruby, Go, Elixir, and Rust
  • Outlines guarantees schema compliance via constrained decoding on local models — zero hallucination risk on the structure itself
  • Pydantic AI adds type safety to multi-turn agent conversations and falls back from native structured output to tool calls to prompted JSON
  • BAML puts the schema and prompt in a versioned .baml file and generates typed clients, so polyglot teams share one contract
  • LangChain with_structured_output() unifies structured output across OpenAI, Anthropic, and Google APIs
  • Marvin 3.x sits on top of Pydantic AI and reduces extraction to a single extract or classify call
  • PromptQuorum tests structured output consistency across all models before production deployment

💡 TL;DR

Use Instructor for Python API extraction with retries. Use Outlines for guaranteed schema compliance on local models. Use Pydantic AI for type-safe multi-turn agents. Use BAML when Python, TypeScript, and Go services must share one schema. Use LangChain if you are already in that ecosystem. Use Marvin for one-line extract and classify calls. Use PromptQuorum to test structured output consistency across all models before production.

⚡ Quick Facts

  • ·Instructor is MIT-licensed and ships six official implementations: Python, TypeScript, Ruby, Go, Elixir, and Rust
  • ·Outlines 1.x constrains tokens at generation time and now also drives hosted APIs, not just local backends
  • ·Pydantic AI offers three output modes — native structured output, tool calls, and prompted JSON
  • ·BAML compiles one .baml schema file into typed clients and repairs malformed model output instead of retrying
  • ·LangChain 1.x reads each provider native structured-output support from its model profile
  • ·Marvin 3.x is built on Pydantic AI and exposes extract, cast, classify, and generate
  • ·PromptQuorum tests the same prompt across 25+ models for consistency

Problems Each Tool Solves

📍 In One Sentence

Structured-output tools solve three distinct problems — enforcing a schema at generation time, validating the result afterwards, and repairing malformed output — and most stacks need only the first two.

💬 In Plain Terms

Do not shop by feature list. Ask which failure you actually hit: the model ignores your format, or it obeys the format but the values are wrong, or it returns JSON that will not even parse. Each has a different answer.

Structured output requires solving three interdependent problems: schema definition, API enforcement, and validation. Different tools attack these problems differently. Instructor handles all three in Python with retries. Outlines eliminates the validation step via constrained decoding. Pydantic AI adds type safety for agents. BAML moves the schema into a compiled file and repairs imperfect output. LangChain wraps provider APIs. Marvin prioritises developer speed. PromptQuorum validates consistency across all models.

Problem
Instructor
Outlines
Pydantic AI
BAML
LangChain
Marvin
Define schemaPydantic modelsJSON Schema / GBNFPydantic models.baml class filesTool definitionsPython type hints
Enforce on API callRetry + validationToken-level constraintNative / tool / promptedGenerated prompt + parserProvider JSON modePydantic AI output types
Validate responseAutomaticGuaranteed at generationType-checkedSchema-aligned parsingManualAutomatic

Instructor: Pydantic Extraction

Instructor is the most widely adopted structured output library. It wraps any LLM API — OpenAI GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, Ollama, vLLM — and returns validated Pydantic models instead of raw text. Instructor handles retries automatically when validation fails, making it production-grade without extra error handling.

  • Works with every major provider (OpenAI, Anthropic, Google, Groq, Mistral) and local models via Ollama or vLLM
  • Pydantic v2 schemas: type hints, validation rules, docstring descriptions embedded in schema
  • Automatic retry with backoff on validation failure — no manual error handling needed
  • Six official implementations: Python, TypeScript, Ruby, Go, Elixir, and Rust
  • MIT-licensed open source, actively maintained, currently on the 1.x line
  • Pricing: Free (no additional cost beyond LLM API calls)
python
import instructor
from pydantic import BaseModel
from openai import OpenAI

class User(BaseModel):
    name: str
    age: int

client = instructor.from_openai(OpenAI())
user = client.chat.completions.create(
    model="gpt-5.6",
    response_model=User,
    messages=[{"role": "user", "content": "Extract: John is 25 years old"}]
)
# user.name == "John", user.age == 25

Outlines: Constrained Decoding

Outlines enforces schema compliance at token generation time via constrained decoding. Instead of generating tokens then validating, Outlines limits valid tokens at each step to match your schema. This guarantees the output parses against your schema with zero structural hallucination risk, which is what makes it the default choice for local models.

  • Local backends: transformers, llama.cpp, MLX, and any Hugging Face model
  • Server backends: vLLM, Ollama, and NVIDIA NIM
  • Hosted APIs are supported too (OpenAI, Gemini), so the same code moves between local and cloud
  • Schemas as Pydantic models, JSON Schema, regex patterns, literal choices, or context-free grammars
  • Guaranteed structural compliance — no post-generation validation or retries needed
  • Apache 2.0 open source, currently on the 1.x line, with a Rust core (outlines-core) for speed

Pydantic AI: Type-Safe Agents

Pydantic AI is the agent framework from the team behind Pydantic itself. It combines Pydantic models with first-class support for multi-turn agent conversations, adding full type safety to agent loops while enforcing structured output on each turn. It is past its 2.x line and used in production, not an experiment.

  • Pydantic v2 type system — full IDE support and static type checking on what an agent returns
  • Three output modes: provider-native structured output, tool calls, and prompted JSON as a fallback
  • Async-first design for high-throughput applications
  • Supports OpenAI, Anthropic, Google, Bedrock, Azure AI Foundry, Groq, Mistral, xAI, and Ollama
  • Durable execution integrations (Temporal, DBOS, Prefect) so long-running agents survive restarts
  • Tool calling baked in — define tools as Python functions with type hints
  • MIT-licensed and free to use (no additional cost beyond LLM API calls)

BAML: Schema-First Prompt Files

BAML takes the opposite approach to the Python libraries: the schema and the prompt live in a versioned .baml file, and a compiler generates a typed client for your language. Its schema-aligned parser repairs the mistakes models actually make — markdown fences around JSON, trailing commas, unquoted keys, reasoning text before the object — instead of throwing an error and burning a retry.

  • Schema and prompt live together in .baml files, versioned and reviewed like any other source
  • Generates typed clients for Python and TypeScript natively, plus Go, Java, Ruby, PHP, Rust, and C# via generated OpenAPI clients
  • Schema-aligned parsing (SAP) recovers valid objects from imperfect model output rather than failing
  • Works with models that have no native tool-use or JSON mode at all
  • Type-safe streaming — partial objects arrive typed, so you can render fields as they generate
  • Apache 2.0 open source; the hosted Boundary Studio observability product is a separate paid offering

LangChain: Unified APIs

LangChain exposes with_structured_output() on all major chat models, unifying structured output across OpenAI, Anthropic, Google, and local models behind a single method. Since the 1.x rewrite it reads each provider native structured-output capability from that model profile rather than hardcoding it, and agents built with create_agent accept a response_format directly.

  • Unified API: one .with_structured_output() method works across all providers
  • Automatically converts LangChain tool definitions to provider-specific schema formats
  • Agents created with create_agent take a response_format for their final answer
  • Native structured-output support is read per model from provider profile data on the 1.1+ line
  • Supports Pydantic models, TypedDict, dataclasses, and raw JSON Schema
  • Best for teams already invested in LangChain or LangGraph

Marvin: Task-Based Extraction

Marvin 3.x is the shortest path from unstructured text to a typed Python object. It is built on top of Pydantic AI, so you get the same provider coverage and validation with far less code. Note that the decorator-first API of Marvin 2 is gone: @marvin.fn was removed in 3.0 in favour of top-level helpers and a task-centric agent engine.

  • One-line helpers: marvin.extract, marvin.cast, marvin.classify, and marvin.generate
  • Built on Pydantic AI, so provider support and output validation are inherited, not reimplemented
  • Task-centric engine for multi-step work: marvin.run, marvin.Task, marvin.Agent, marvin.Thread
  • Python type hints become the schema — minimal boilerplate for extraction and classification
  • Migration note: the Marvin 2 @marvin.fn decorator no longer exists; rewrite those call sites
  • Apache 2.0 open source, maintained by Prefect, free to use

PromptQuorum: Cross-Model Testing

PromptQuorum is not a structured output library itself, but a testing platform for validating structured output consistency across models. Run the same prompt against GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, and 20+ other models simultaneously. Measure schema compliance, latency, and cost per model.

  • Multi-model dispatch in a single API call — test one prompt against 25+ models
  • Structured output compliance metrics — pass rate, latency, cost per model
  • Identify models that hallucinate on your schema — avoid deploying to unreliable models
  • Consensus mode — find agreements between independent model runs
  • Works with Instructor, Outlines, Pydantic AI, BAML, LangChain, or raw LLM APIs
  • Free tier available, enterprise pricing for high-volume testing

Side-by-Side Comparison

Tool
Best For
Schema Format
Language
Local Models
Licence
Learning Curve
InstructorPython APIs + retriesPydantic modelsPython, TS, Ruby, Go, Elixir, RustYes (Ollama, vLLM)MIT, freeLow
OutlinesLocal model deploymentPydantic, JSON Schema, regex, CFGPythonYes (native)Apache 2.0, freeMedium
Pydantic AIType-safe agentsPydantic modelsPythonYes (Ollama)MIT, freeLow
BAMLPolyglot teams, flaky models.baml class filesPython, TS + 6 via OpenAPIYes (OpenAI-compatible)Apache 2.0, paid observabilityMedium
LangChainChains + agentsTool definitionsPython, JSYesMIT, freeMedium
MarvinFast extract + classifyType hintsPythonYesApache 2.0, freeVery low
PromptQuorumMulti-model testingAPI-agnosticAPI-firstVia OpenAI proxyFree tier + enterpriseLow

Choosing the Right Tool

Start by answering three questions: (1) Which languages do the services that call the model actually ship in? (2) Do you need local model support? (3) How much validation complexity do you have?

  • Use Instructor if: You are building Python APIs and need automatic retries on validation failure. Best general-purpose choice.
  • Use Outlines if: You deploy local models (llama.cpp, vLLM, MLX) and want guaranteed schema compliance at generation time.
  • Use Pydantic AI if: You are building multi-turn agent workflows with type safety across all steps, or need durable execution.
  • Use BAML if: Python, TypeScript, and Go services must share one schema, or your model has no reliable native JSON mode.
  • Use LangChain if: You already use LangChain or LangGraph — with_structured_output() is the simplest addition.
  • Use Marvin if: You want a single extract or classify call and do not need custom validation logic.
  • Use PromptQuorum if: You need to test structured output consistency across GPT, Claude, and Gemini before production.

Adding Structured Output Step-by-Step

  1. 1
    Define your output schema — Create a Pydantic model (Python), a .baml class (BAML), a TypeScript interface, or JSON Schema describing the fields, types, and constraints you want the LLM to return.
  2. 2
    Choose a library — Instructor for Python APIs, Outlines for local models, Pydantic AI for agents, BAML for polyglot teams, LangChain if already in use, Marvin for one-line extraction.
  3. 3
    Install and wrap your LLM call — `pip install instructor` (Python), then pass your schema to the API call. Instructor handles validation and retries.
  4. 4
    Test with PromptQuorum — Deploy to PromptQuorum and run your prompt against GPT, Claude, and Gemini. Measure schema compliance per model.
  5. 5
    Refine schema based on failures — If a model fails validation, add examples to your prompt or adjust schema constraints. Iterate until all models pass.

Common Structured Output Mistakes

❌ Treating every JSON mode as a schema guarantee

Why it hurts: Plain JSON mode (response_format json_object, Anthropic JSON control) only guarantees the reply is valid JSON — not that it matches your fields or types. Strict schema modes go further and guarantee the shape, but neither guarantees the values are correct: a well-formed object can still contain an invented price or a hallucinated date.

Fix: Layer validation on top regardless: Instructor, Outlines, Pydantic AI, or BAML. Enforce business rules in Pydantic validators, not in the schema alone. Test with PromptQuorum to catch compliance failures per model.

❌ Designing schemas that are too strict

Why it hurts: Overly constrained schemas (tiny enum lists, very specific regex patterns) cause LLMs to fail validation frequently. High retry counts waste tokens and money.

Fix: Use PromptQuorum to test schema strictness across models. Loosen constraints to achieve 95%+ compliance. Use optional fields instead of required ones when possible.

❌ Not testing local vs. API model differences

Why it hurts: Outlines on llama.cpp behaves differently than Instructor on GPT-5.6. Schema compliance rates differ per model. Building only for a frontier API model, then deploying to a small local one, causes production failures.

Fix: Test all intended model backends early. Use PromptQuorum to run the same prompt across local (vLLM, Ollama) and hosted (OpenAI, Anthropic, Google) models.

❌ Ignoring latency and token cost impact

Why it hurts: Structured output with retries costs more tokens. Instructor retries on failure. Outlines constrained decoding adds per-token overhead compared with free generation. Not measuring per-model cost.

Fix: Use PromptQuorum cost tracking. Compare latency across models. For budget-conscious workflows, prefer Outlines or BAML (no retry loop). For accuracy on flexible schemas, accept Instructor retry cost.

❌ Mixing validation methods (no consistency)

Why it hurts: Some requests use Instructor, others use raw JSON parsing. Some models validated, others not. This leads to inconsistent errors in production.

Fix: Standardize on one validation approach per codebase. All requests use Instructor, or all use Outlines. Consistency reduces debugging time by 10x.

❌ Copying tutorials written against a superseded API

Why it hurts: Structured output libraries move fast. Marvin removed the @marvin.fn decorator in 3.0, LangChain reorganised its docs in the 1.x rewrite, and Outlines changed its import surface at 1.0. Code copied from an older tutorial fails on install.

Fix: Pin the major version you develop against and check the current docs for the API surface. Prefer the official repository README over blog posts, and re-check when you upgrade a major version.

What is structured output in LLMs?

Structured output constrains LLM responses to a specific schema — JSON format, defined fields, type constraints. Instead of free-text replies, structured output returns data your code can directly parse and validate without error handling.

Which tool is best for Python developers?

Instructor is the most popular Python choice. It uses Pydantic models to define schemas, automatically handles retries and validation, and supports every major LLM API plus local models via Ollama or vLLM. Pydantic AI is the better fit if you also want type-safe multi-turn agent conversations, and Marvin is the fastest option if you just need a one-line extract or classify call.

Can I use structured output with local models like Llama?

Yes. Outlines specialises in local model constrained decoding — it works with transformers, llama.cpp, MLX, vLLM, and Ollama, and guarantees the output parses against your schema at generation time. Instructor and Pydantic AI also support Ollama and vLLM if you run them as an API, and BAML works against any OpenAI-compatible endpoint.

What is the difference between Instructor and Marvin?

Instructor wraps your own LLM client and returns validated Pydantic models with automatic retries, so you control the call. Marvin 3.x is built on top of Pydantic AI and gives you one-line helpers instead — marvin.extract, marvin.cast, marvin.classify. Instructor is more explicit and better for complex validation; Marvin is more concise for straightforward extraction. Note that the @marvin.fn decorator from Marvin 2 was removed in Marvin 3.

Does LangChain support structured output?

Yes. LangChain exposes with_structured_output() on ChatOpenAI, ChatAnthropic, ChatGoogleGenerativeAI, and the other chat model classes, and agents built with create_agent accept a response_format. Since the 1.x line it reads each provider native structured-output support from model profile data rather than hardcoding it. Use this if you already run LangChain or LangGraph and want schema enforcement without switching libraries.

How do I test if structured output is reliable?

Use PromptQuorum to run the same prompt across multiple models and measure schema compliance. Different models — GPT-5.6, Claude Opus 5, Gemini 3.1 Pro — have different structured output reliability, and small local models differ more again. Test before deploying to production, and unit test with Instructor or Pydantic validation locally.

What does "constrained decoding" mean?

Constrained decoding limits token generation to only valid values according to your schema. Outlines does this by computing the set of valid next tokens at each step. This guarantees the output parses against your schema without post-generation validation or retries, making it more reliable than plain API-level JSON mode. It constrains structure, not truth — the fields will be right, the values still need checking.

What is BAML and when should I use it instead of Instructor?

BAML is a schema-first language: you write the schema and prompt in a .baml file and compile a typed client for your language. Choose it over Instructor when more than one language calls the same prompt — a Python worker and a TypeScript frontend sharing one contract — or when your model returns almost-valid JSON, because BAML schema-aligned parser repairs markdown fences, trailing commas, and leading reasoning text instead of burning a retry. Stay on Instructor if your stack is Python-only and you want to keep schemas in ordinary Pydantic code.

Can I use structured output without any library?

Technically, yes — you can prompt the model to return JSON and then parse it yourself. But parsing will fail on the malformed output models still produce, and nothing enforces your field names or types. All seven tools solve this by validating with retries (Instructor, Marvin), enforcing at decode time (Outlines), repairing output at parse time (BAML), or wrapping provider APIs (LangChain, Pydantic AI).

Which tool has the best documentation?

LangChain and Pydantic AI have the most comprehensive docs due to their corporate backing. BAML documentation is unusually good for a young project because the language needs teaching. Instructor has excellent tutorials and examples despite being community-maintained. Outlines docs are technical but thorough. Marvin docs are concise — check the 3.x pages specifically, as older Marvin 2 material still circulates.

Do I need all seven tools or just one?

Start with one. Python developers should try Instructor or Pydantic AI. Local model teams should try Outlines. Polyglot teams should try BAML. LangChain users should try with_structured_output(). Use PromptQuorum to validate consistency across all models. Most teams use one tool plus PromptQuorum for testing.

Sources

Apply these techniques with a local LLM or your own API keys — PromptQuorum works with any backend.

Try PromptQuorum free →

← Back to Prompt Engineering