Problems Each Tool Solves
📍 In One Sentence
Structured-output tools solve three distinct problems — enforcing a schema at generation time, validating the result afterwards, and repairing malformed output — and most stacks need only the first two.
💬 In Plain Terms
Do not shop by feature list. Ask which failure you actually hit: the model ignores your format, or it obeys the format but the values are wrong, or it returns JSON that will not even parse. Each has a different answer.
Structured output requires solving three interdependent problems: schema definition, API enforcement, and validation. Different tools attack these problems differently. Instructor handles all three in Python with retries. Outlines eliminates the validation step via constrained decoding. Pydantic AI adds type safety for agents. BAML moves the schema into a compiled file and repairs imperfect output. LangChain wraps provider APIs. Marvin prioritises developer speed. PromptQuorum validates consistency across all models.
Problem | Instructor | Outlines | Pydantic AI | BAML | LangChain | Marvin |
|---|---|---|---|---|---|---|
| Define schema | Pydantic models | JSON Schema / GBNF | Pydantic models | .baml class files | Tool definitions | Python type hints |
| Enforce on API call | Retry + validation | Token-level constraint | Native / tool / prompted | Generated prompt + parser | Provider JSON mode | Pydantic AI output types |
| Validate response | Automatic | Guaranteed at generation | Type-checked | Schema-aligned parsing | Manual | Automatic |
Instructor: Pydantic Extraction
Instructor is the most widely adopted structured output library. It wraps any LLM API — OpenAI GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, Ollama, vLLM — and returns validated Pydantic models instead of raw text. Instructor handles retries automatically when validation fails, making it production-grade without extra error handling.
- Works with every major provider (OpenAI, Anthropic, Google, Groq, Mistral) and local models via Ollama or vLLM
- Pydantic v2 schemas: type hints, validation rules, docstring descriptions embedded in schema
- Automatic retry with backoff on validation failure — no manual error handling needed
- Six official implementations: Python, TypeScript, Ruby, Go, Elixir, and Rust
- MIT-licensed open source, actively maintained, currently on the 1.x line
- Pricing: Free (no additional cost beyond LLM API calls)
import instructor
from pydantic import BaseModel
from openai import OpenAI
class User(BaseModel):
name: str
age: int
client = instructor.from_openai(OpenAI())
user = client.chat.completions.create(
model="gpt-5.6",
response_model=User,
messages=[{"role": "user", "content": "Extract: John is 25 years old"}]
)
# user.name == "John", user.age == 25Outlines: Constrained Decoding
Outlines enforces schema compliance at token generation time via constrained decoding. Instead of generating tokens then validating, Outlines limits valid tokens at each step to match your schema. This guarantees the output parses against your schema with zero structural hallucination risk, which is what makes it the default choice for local models.
- Local backends: transformers, llama.cpp, MLX, and any Hugging Face model
- Server backends: vLLM, Ollama, and NVIDIA NIM
- Hosted APIs are supported too (OpenAI, Gemini), so the same code moves between local and cloud
- Schemas as Pydantic models, JSON Schema, regex patterns, literal choices, or context-free grammars
- Guaranteed structural compliance — no post-generation validation or retries needed
- Apache 2.0 open source, currently on the 1.x line, with a Rust core (outlines-core) for speed
Pydantic AI: Type-Safe Agents
Pydantic AI is the agent framework from the team behind Pydantic itself. It combines Pydantic models with first-class support for multi-turn agent conversations, adding full type safety to agent loops while enforcing structured output on each turn. It is past its 2.x line and used in production, not an experiment.
- Pydantic v2 type system — full IDE support and static type checking on what an agent returns
- Three output modes: provider-native structured output, tool calls, and prompted JSON as a fallback
- Async-first design for high-throughput applications
- Supports OpenAI, Anthropic, Google, Bedrock, Azure AI Foundry, Groq, Mistral, xAI, and Ollama
- Durable execution integrations (Temporal, DBOS, Prefect) so long-running agents survive restarts
- Tool calling baked in — define tools as Python functions with type hints
- MIT-licensed and free to use (no additional cost beyond LLM API calls)
BAML: Schema-First Prompt Files
BAML takes the opposite approach to the Python libraries: the schema and the prompt live in a versioned .baml file, and a compiler generates a typed client for your language. Its schema-aligned parser repairs the mistakes models actually make — markdown fences around JSON, trailing commas, unquoted keys, reasoning text before the object — instead of throwing an error and burning a retry.
- Schema and prompt live together in .baml files, versioned and reviewed like any other source
- Generates typed clients for Python and TypeScript natively, plus Go, Java, Ruby, PHP, Rust, and C# via generated OpenAPI clients
- Schema-aligned parsing (SAP) recovers valid objects from imperfect model output rather than failing
- Works with models that have no native tool-use or JSON mode at all
- Type-safe streaming — partial objects arrive typed, so you can render fields as they generate
- Apache 2.0 open source; the hosted Boundary Studio observability product is a separate paid offering
LangChain: Unified APIs
LangChain exposes with_structured_output() on all major chat models, unifying structured output across OpenAI, Anthropic, Google, and local models behind a single method. Since the 1.x rewrite it reads each provider native structured-output capability from that model profile rather than hardcoding it, and agents built with create_agent accept a response_format directly.
- Unified API: one .with_structured_output() method works across all providers
- Automatically converts LangChain tool definitions to provider-specific schema formats
- Agents created with create_agent take a response_format for their final answer
- Native structured-output support is read per model from provider profile data on the 1.1+ line
- Supports Pydantic models, TypedDict, dataclasses, and raw JSON Schema
- Best for teams already invested in LangChain or LangGraph
Marvin: Task-Based Extraction
Marvin 3.x is the shortest path from unstructured text to a typed Python object. It is built on top of Pydantic AI, so you get the same provider coverage and validation with far less code. Note that the decorator-first API of Marvin 2 is gone: @marvin.fn was removed in 3.0 in favour of top-level helpers and a task-centric agent engine.
- One-line helpers: marvin.extract, marvin.cast, marvin.classify, and marvin.generate
- Built on Pydantic AI, so provider support and output validation are inherited, not reimplemented
- Task-centric engine for multi-step work: marvin.run, marvin.Task, marvin.Agent, marvin.Thread
- Python type hints become the schema — minimal boilerplate for extraction and classification
- Migration note: the Marvin 2 @marvin.fn decorator no longer exists; rewrite those call sites
- Apache 2.0 open source, maintained by Prefect, free to use
PromptQuorum: Cross-Model Testing
PromptQuorum is not a structured output library itself, but a testing platform for validating structured output consistency across models. Run the same prompt against GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, and 20+ other models simultaneously. Measure schema compliance, latency, and cost per model.
- Multi-model dispatch in a single API call — test one prompt against 25+ models
- Structured output compliance metrics — pass rate, latency, cost per model
- Identify models that hallucinate on your schema — avoid deploying to unreliable models
- Consensus mode — find agreements between independent model runs
- Works with Instructor, Outlines, Pydantic AI, BAML, LangChain, or raw LLM APIs
- Free tier available, enterprise pricing for high-volume testing
Side-by-Side Comparison
Tool | Best For | Schema Format | Language | Local Models | Licence | Learning Curve |
|---|---|---|---|---|---|---|
| Instructor | Python APIs + retries | Pydantic models | Python, TS, Ruby, Go, Elixir, Rust | Yes (Ollama, vLLM) | MIT, free | Low |
| Outlines | Local model deployment | Pydantic, JSON Schema, regex, CFG | Python | Yes (native) | Apache 2.0, free | Medium |
| Pydantic AI | Type-safe agents | Pydantic models | Python | Yes (Ollama) | MIT, free | Low |
| BAML | Polyglot teams, flaky models | .baml class files | Python, TS + 6 via OpenAPI | Yes (OpenAI-compatible) | Apache 2.0, paid observability | Medium |
| LangChain | Chains + agents | Tool definitions | Python, JS | Yes | MIT, free | Medium |
| Marvin | Fast extract + classify | Type hints | Python | Yes | Apache 2.0, free | Very low |
| PromptQuorum | Multi-model testing | API-agnostic | API-first | Via OpenAI proxy | Free tier + enterprise | Low |
Choosing the Right Tool
Start by answering three questions: (1) Which languages do the services that call the model actually ship in? (2) Do you need local model support? (3) How much validation complexity do you have?
- Use Instructor if: You are building Python APIs and need automatic retries on validation failure. Best general-purpose choice.
- Use Outlines if: You deploy local models (llama.cpp, vLLM, MLX) and want guaranteed schema compliance at generation time.
- Use Pydantic AI if: You are building multi-turn agent workflows with type safety across all steps, or need durable execution.
- Use BAML if: Python, TypeScript, and Go services must share one schema, or your model has no reliable native JSON mode.
- Use LangChain if: You already use LangChain or LangGraph — with_structured_output() is the simplest addition.
- Use Marvin if: You want a single extract or classify call and do not need custom validation logic.
- Use PromptQuorum if: You need to test structured output consistency across GPT, Claude, and Gemini before production.
Adding Structured Output Step-by-Step
- 1Define your output schema — Create a Pydantic model (Python), a .baml class (BAML), a TypeScript interface, or JSON Schema describing the fields, types, and constraints you want the LLM to return.
- 2Choose a library — Instructor for Python APIs, Outlines for local models, Pydantic AI for agents, BAML for polyglot teams, LangChain if already in use, Marvin for one-line extraction.
- 3Install and wrap your LLM call — `pip install instructor` (Python), then pass your schema to the API call. Instructor handles validation and retries.
- 4Test with PromptQuorum — Deploy to PromptQuorum and run your prompt against GPT, Claude, and Gemini. Measure schema compliance per model.
- 5Refine schema based on failures — If a model fails validation, add examples to your prompt or adjust schema constraints. Iterate until all models pass.
Common Structured Output Mistakes
❌ Treating every JSON mode as a schema guarantee
Why it hurts: Plain JSON mode (response_format json_object, Anthropic JSON control) only guarantees the reply is valid JSON — not that it matches your fields or types. Strict schema modes go further and guarantee the shape, but neither guarantees the values are correct: a well-formed object can still contain an invented price or a hallucinated date.
Fix: Layer validation on top regardless: Instructor, Outlines, Pydantic AI, or BAML. Enforce business rules in Pydantic validators, not in the schema alone. Test with PromptQuorum to catch compliance failures per model.
❌ Designing schemas that are too strict
Why it hurts: Overly constrained schemas (tiny enum lists, very specific regex patterns) cause LLMs to fail validation frequently. High retry counts waste tokens and money.
Fix: Use PromptQuorum to test schema strictness across models. Loosen constraints to achieve 95%+ compliance. Use optional fields instead of required ones when possible.
❌ Not testing local vs. API model differences
Why it hurts: Outlines on llama.cpp behaves differently than Instructor on GPT-5.6. Schema compliance rates differ per model. Building only for a frontier API model, then deploying to a small local one, causes production failures.
Fix: Test all intended model backends early. Use PromptQuorum to run the same prompt across local (vLLM, Ollama) and hosted (OpenAI, Anthropic, Google) models.
❌ Ignoring latency and token cost impact
Why it hurts: Structured output with retries costs more tokens. Instructor retries on failure. Outlines constrained decoding adds per-token overhead compared with free generation. Not measuring per-model cost.
Fix: Use PromptQuorum cost tracking. Compare latency across models. For budget-conscious workflows, prefer Outlines or BAML (no retry loop). For accuracy on flexible schemas, accept Instructor retry cost.
❌ Mixing validation methods (no consistency)
Why it hurts: Some requests use Instructor, others use raw JSON parsing. Some models validated, others not. This leads to inconsistent errors in production.
Fix: Standardize on one validation approach per codebase. All requests use Instructor, or all use Outlines. Consistency reduces debugging time by 10x.
❌ Copying tutorials written against a superseded API
Why it hurts: Structured output libraries move fast. Marvin removed the @marvin.fn decorator in 3.0, LangChain reorganised its docs in the 1.x rewrite, and Outlines changed its import surface at 1.0. Code copied from an older tutorial fails on install.
Fix: Pin the major version you develop against and check the current docs for the API surface. Prefer the official repository README over blog posts, and re-check when you upgrade a major version.
What is structured output in LLMs?
Structured output constrains LLM responses to a specific schema — JSON format, defined fields, type constraints. Instead of free-text replies, structured output returns data your code can directly parse and validate without error handling.
Which tool is best for Python developers?
Instructor is the most popular Python choice. It uses Pydantic models to define schemas, automatically handles retries and validation, and supports every major LLM API plus local models via Ollama or vLLM. Pydantic AI is the better fit if you also want type-safe multi-turn agent conversations, and Marvin is the fastest option if you just need a one-line extract or classify call.
Can I use structured output with local models like Llama?
Yes. Outlines specialises in local model constrained decoding — it works with transformers, llama.cpp, MLX, vLLM, and Ollama, and guarantees the output parses against your schema at generation time. Instructor and Pydantic AI also support Ollama and vLLM if you run them as an API, and BAML works against any OpenAI-compatible endpoint.
What is the difference between Instructor and Marvin?
Instructor wraps your own LLM client and returns validated Pydantic models with automatic retries, so you control the call. Marvin 3.x is built on top of Pydantic AI and gives you one-line helpers instead — marvin.extract, marvin.cast, marvin.classify. Instructor is more explicit and better for complex validation; Marvin is more concise for straightforward extraction. Note that the @marvin.fn decorator from Marvin 2 was removed in Marvin 3.
Does LangChain support structured output?
Yes. LangChain exposes with_structured_output() on ChatOpenAI, ChatAnthropic, ChatGoogleGenerativeAI, and the other chat model classes, and agents built with create_agent accept a response_format. Since the 1.x line it reads each provider native structured-output support from model profile data rather than hardcoding it. Use this if you already run LangChain or LangGraph and want schema enforcement without switching libraries.
How do I test if structured output is reliable?
Use PromptQuorum to run the same prompt across multiple models and measure schema compliance. Different models — GPT-5.6, Claude Opus 5, Gemini 3.1 Pro — have different structured output reliability, and small local models differ more again. Test before deploying to production, and unit test with Instructor or Pydantic validation locally.
What does "constrained decoding" mean?
Constrained decoding limits token generation to only valid values according to your schema. Outlines does this by computing the set of valid next tokens at each step. This guarantees the output parses against your schema without post-generation validation or retries, making it more reliable than plain API-level JSON mode. It constrains structure, not truth — the fields will be right, the values still need checking.
What is BAML and when should I use it instead of Instructor?
BAML is a schema-first language: you write the schema and prompt in a .baml file and compile a typed client for your language. Choose it over Instructor when more than one language calls the same prompt — a Python worker and a TypeScript frontend sharing one contract — or when your model returns almost-valid JSON, because BAML schema-aligned parser repairs markdown fences, trailing commas, and leading reasoning text instead of burning a retry. Stay on Instructor if your stack is Python-only and you want to keep schemas in ordinary Pydantic code.
Can I use structured output without any library?
Technically, yes — you can prompt the model to return JSON and then parse it yourself. But parsing will fail on the malformed output models still produce, and nothing enforces your field names or types. All seven tools solve this by validating with retries (Instructor, Marvin), enforcing at decode time (Outlines), repairing output at parse time (BAML), or wrapping provider APIs (LangChain, Pydantic AI).
Which tool has the best documentation?
LangChain and Pydantic AI have the most comprehensive docs due to their corporate backing. BAML documentation is unusually good for a young project because the language needs teaching. Instructor has excellent tutorials and examples despite being community-maintained. Outlines docs are technical but thorough. Marvin docs are concise — check the 3.x pages specifically, as older Marvin 2 material still circulates.
Do I need all seven tools or just one?
Start with one. Python developers should try Instructor or Pydantic AI. Local model teams should try Outlines. Polyglot teams should try BAML. LangChain users should try with_structured_output(). Use PromptQuorum to validate consistency across all models. Most teams use one tool plus PromptQuorum for testing.
Sources
- Instructor GitHub Repository — Official repository and docs for the Instructor library
- Outlines GitHub Repository — Constrained decoding for guaranteed schema compliance
- Pydantic AI Documentation — Type-safe agent framework with structured output
- LangChain Structured Output Guide — LangChain unified structured output API
- BAML Documentation — Schema-first prompt language and schema-aligned parsing
- Marvin GitHub Repository — Task-centric extraction library built on Pydantic AI
