What These Tools Do (and What They Don’t)
📍 In One Sentence
LangSmith, Helicone, Langfuse, and PromptLayer log and trace LLM calls already running in production — a different job from pre-deployment prompt testing and evals.
💬 In Plain Terms
Think of it like application monitoring: these tools are Datadog or Sentry for your LLM calls, not a test suite you run before deploying. You still need both — they just cover different stages.
LangSmith, Helicone, Langfuse, and PromptLayer are production LLM observability platforms — they log API calls after your application makes them, trace multi-step chains, and track cost and latency over time. They answer "what happened when this ran in production?"
That is a different question from "is this prompt good before I ship it?" — pre-deployment testing and evaluation is the job of tools like Promptfoo, Braintrust, and Vellum, covered in Prompt Testing & Evaluation Tools 2026. Prompt version control and team sharing before deployment is covered in Best Prompt Management Platforms. The tools on this page start earning their keep the moment a prompt goes live and real users start calling it.
The overlap: all four also offer some evaluation and dataset features, because knowing a prompt broke in production is more useful if you can also score how badly. But scoring is a secondary feature here, not the primary one — the primary job is tracing, logging, and cost visibility for live traffic.
Disclosure: every vendor link on this page (LangSmith, Langfuse, Helicone, PromptLayer) points directly to that vendor’s own site. These are plain product links — none of the four currently runs a public affiliate or commission program, so there is no current affiliate relationship to disclose beyond that.
How We Compared These Tools
We evaluated the four tools against five criteria that matter for production tracing specifically: integration effort, self-hosting availability, framework breadth, cost transparency at scale, and eval features layered on top of logging.
Criterion | What It Measures | Why It Matters |
|---|---|---|
| Integration effort | SDK wrapper vs. proxy vs. one-line header change | Determines how much of your codebase you touch to add tracing |
| Self-hosting | Whether the tool can run on your own infrastructure, and at what cost | Data residency, compliance, and long-term cost control |
| Framework breadth | Number of SDKs/frameworks with native integration | Teams using multiple providers need broad, not deep, coverage |
| Cost at scale | What happens to price past the free tier at real production volume | Free-tier pricing pages hide the number that actually matters |
| Eval depth | Dataset-driven scoring layered on top of raw logs | Logging tells you something broke; eval tells you how badly |
LangSmith: Deepest LangChain Integration
LangSmith is LangChain’s own observability platform — tracing is automatic for LangChain and LangGraph applications and requires the least setup of any tool here if you already build on those frameworks. Outside the LangChain ecosystem, LangSmith still works via its SDK, but the automatic-tracing advantage disappears. Start free with the Developer tier — no card required for the first 5k traces/month.
Free Developer tier: 1 seat, up to 5,000 base traces/month, then pay-as-you-go. Plus tier: $39/seat/month, unlimited seats, up to 10,000 base traces/month included, then usage-based billing on top — $1.50 per LangChain Compute Unit (LCU) and $1.00 per LangChain Storage Unit (LSU), which meter Engine, Fleet, Deployment, and Sandbox usage. Enterprise is custom-priced with self-hosted and hybrid deployment, SSO, and RBAC.
Best team features: dataset-driven evaluation with LLM-as-judge scoring, regression testing against saved traces, and prompt hub integration so evaluation and versioning live in one place. Tradeoff: the deepest value is specific to LangChain/LangGraph — teams on a different framework get a capable but generic tracer, not the integrated experience LangChain users see.
For teams evaluating whether to standardize on LangChain in the first place, see GPT, Claude, or Gemini? How to Pick the Right Model before locking into a framework-specific observability stack.
- Automatic tracing for LangChain and LangGraph — no manual instrumentation for chains built in those frameworks
- Free Developer tier: 1 seat, 5k base traces/month, then pay-as-you-go
- Plus tier: $39/seat/month, 10k base traces/month included, unlimited seats
- Usage-based: $1.50/LCU (compute) and $1.00/LSU (storage) on top of the base plan
- Dataset-driven evaluation with LLM-as-judge scoring built in
- Enterprise: self-hosted or hybrid deployment, custom SSO/RBAC, support SLA
⚠️ Usage Units Add Up
LCU and LSU charges apply on top of the $39/seat base — a team running heavy chains or high-volume traces can see the usage-based portion exceed the seat cost. Estimate LCU/LSU consumption from a pilot before committing to Plus at scale.
Langfuse: Free Self-Hosting, No Feature Paywall
Langfuse is MIT-licensed and free to self-host with every core feature included — there is no feature gate separating the free self-hosted version from the paid cloud tiers. Langfuse was acquired by ClickHouse, Inc. in January 2026; the product stays open source and self-hostable, and Langfuse Cloud continues running unchanged. Start free with self-hosting or the Cloud Hobby tier — no card required.
Self-hosted (Open Source): free, unlimited, MIT license, all core observability/evaluation/prompt-management features. Cloud Hobby: free, 50,000 units/month, 2 users, 30-day data retention. Cloud Core: $29/month, 100,000 units/month included, unlimited users, 90-day retention, $8 per additional 100k units. Cloud Pro: $199/month, 100,000 units/month included, unlimited users, 3-year data retention, plus a $300/month Teams add-on for SSO/RBAC. Cloud Enterprise: $2,499/month, dedicated support engineer, audit logs, custom SLA. Self-Hosted Enterprise (separately priced, bundled with a ClickHouse commercial plan) adds management APIs, project-level RBAC, and server-side data masking on top of the free self-hosted core.
Broadest framework integration of the four: OpenAI SDK, LangChain, LlamaIndex, Haystack, LiteLLM, Vercel AI SDK, Ollama, and Amazon Bedrock all have native integrations. Eval depth matches LangSmith — dataset-driven evaluation with LLM-as-judge scoring and regression datasets are both mature, not bolted-on logging extras.
Tradeoff: self-hosting means you own uptime, upgrades, and infrastructure cost — "free" software still needs someone to run it. The Cloud Hobby tier’s 2-user cap is tight for anything beyond a solo project, pushing teams to the $29/month Core tier fairly quickly once more than two people need access.
- Self-hosted, MIT license, free forever — no feature gate vs. paid cloud tiers
- Cloud Hobby: free, 50k units/month, 2 users, 30-day retention
- Cloud Core: $29/month, 100k units/month, unlimited users, 90-day retention
- Cloud Pro: $199/month, 3-year data retention, unlimited users
- Broadest integrations: OpenAI SDK, LangChain, LlamaIndex, Haystack, LiteLLM, Vercel AI SDK, Ollama, Amazon Bedrock
- Acquired by ClickHouse, Inc. (January 2026); 33.8k GitHub stars as of 2026-08-27
📌 Did You Know
Langfuse’s self-hosted version has no feature paywall — the same evaluation and tracing capabilities available on the $2,499/month Enterprise cloud plan run on infrastructure you control, for free, if you’re willing to operate it yourself.
Helicone: One-Line Proxy, Fastest Setup
Helicone is a reverse proxy — you point your API base URL at Helicone and add one header, and every call is logged automatically, with no SDK wrapper to install. This is the fastest of the four to add to an existing codebase, since most teams already have a single place where the API base URL is configured. Start free with the Hobby tier — no card required for the first 10k requests/month.
Hobby (free): 10,000 requests/month, 1 GB storage, 7-day retention, 1 seat. Pro: $79/month, unlimited seats, 1-month retention, alerts and reports, HQL query language. Team: $799/month, SOC-2 and HIPAA compliance, 3-month retention, dedicated Slack channel. Enterprise: custom pricing, on-prem deployment, SAML SSO, custom MSA.
Because it works at the proxy layer, Helicone is provider-agnostic in a different way than LangSmith or Langfuse’s SDK integrations — it works with any OpenAI-compatible endpoint without a framework-specific adapter. Eval and testing features exist but are lighter than LangSmith or Langfuse’s dataset-driven scoring — Helicone’s core strength is logging and cost visibility, not evaluation depth.
- Proxy-based: change base URL + one header, no SDK required
- Hobby (free): 10k requests/month, 1GB storage, 7-day retention, 1 seat
- Pro: $79/month, unlimited seats, 1-month retention
- Team: $799/month, SOC-2/HIPAA, 3-month retention
- Works with any OpenAI-compatible endpoint — no framework-specific adapter needed
- Eval features are lighter than LangSmith/Langfuse — logging and cost visibility are the core strength
💡 Fastest Path to First Trace
If your team just needs cost and latency visibility today and doesn’t want to touch application code beyond an environment variable, Helicone’s proxy setup gets you a working trace faster than any SDK-based tool on this page.
PromptLayer: The Lean Logger
PromptLayer is the leanest of the four — request logging and prompt versioning without a full evaluation framework, aimed at small teams that don’t need dataset-driven scoring. It wraps the OpenAI and Anthropic SDKs directly rather than working as a proxy or a deep framework integration. Start free — no card required for the first 2.5k requests/month.
Free: 5 users, 2,500 requests/month, 10 prompts, 1 workspace, 250 eval-cell executions/month. Pro: $49/month, unlimited playgrounds and workspaces, $0.003 per transaction overage. Team: $500/month, 25 users, 100,000+ requests/month, $0.002 per transaction overage, webhooks included. Enterprise: custom pricing, RBAC, deployment approvals, HIPAA with BAA, and the only self-hosted deployment option PromptLayer offers — available on GCP, AWS, and Azure at the Enterprise tier only.
Tradeoff: no self-hosted or open-source tier at any price below Enterprise — the only one of the four without one — and its framework integration list is the thinnest (OpenAI/Anthropic SDK wrappers, not the broad native-integration list Langfuse offers). Best fit is a small team that wants requests logged and prompts versioned without adopting a heavier evaluation platform.
- Free: 5 users, 2.5k requests/month, 10 prompts
- Pro: $49/month, $0.003/transaction overage
- Team: $500/month, 25 users, 100k+ requests/month
- No self-hosted/open-source tier below Enterprise — the only one of the four without one
- Thinnest framework list: OpenAI/Anthropic SDK wrappers only
- Enterprise: self-hosted on GCP/AWS/Azure, HIPAA with BAA, custom pricing
PromptQuorum: Compare Models Before You Trace Them
Before wiring LangSmith, Langfuse, Helicone, or PromptLayer into a specific model provider’s calls, use PromptQuorum to dispatch one prompt to 25+ models simultaneously and confirm you picked the right provider to trace in the first place. Free tier available.
Observability tools measure what you already chose to run in production — they don’t help you pick the model. PromptQuorum answers "which model handles this prompt best?" before you commit engineering time to instrumenting production tracing for it.
- 25+ models including GPT-5.6, Claude Opus 5, Gemini 3.1 Pro, and local models via Ollama and LM Studio
- 9 built-in prompt frameworks — TRACE, CO-STAR, CRAFT, and more
- Side-by-side response comparison with consensus scoring
- Token count per model — see cost differences before committing to production tracing for one provider
- Free tier — no engineering setup required
Head-to-Head: All 4 Tools Compared
No single tool wins every criterion. LangSmith goes deepest inside LangChain; Langfuse offers the most flexibility (free self-host, broadest integrations); Helicone is fastest to set up; PromptLayer is the leanest and cheapest at small scale.
Tool | Setup | Free Tier | Entry Paid | Self-Host |
|---|---|---|---|---|
| LangSmith | SDK / auto in LangChain | 5k traces, 1 seat | $39/seat + usage | Enterprise only |
| Langfuse | SDK, broadest integrations | 50k units, 2 users | $29/mo Core | Free (MIT), unlimited |
| Helicone | Proxy, one header | 10k requests, 1 seat | $79/mo Pro | Enterprise only |
| PromptLayer | SDK wrapper | 2.5k requests, 5 users | $49/mo Pro | Enterprise only |
| PromptQuorum | Web app, no code | Model comparison + credits | N/A | N/A |
📌 Free-Tier Ceiling
Three of the four tools now put real usage-based billing behind their free tiers. The "free forever" pitch holds up cleanly only for Langfuse’s self-hosted version — every cloud free tier here (LangSmith, Helicone, Langfuse Hobby) has a real ceiling around 5k–50k events/month.
Who Should Use Which Tool
Match the tool to your framework, compliance needs, and team size — not to which one has the most features on paper.
- 1Teams built on LangChain or LangGraph → LangSmith
Why it matters: Tracing is automatic with no manual instrumentation; evaluation and prompt hub live in the same product. - 2Teams that need self-hosting or the most generous free tier → Langfuse
Why it matters: MIT-licensed, free forever if self-hosted, broadest framework integration list of the four. - 3Teams wanting cost visibility with almost no code changes → Helicone
Why it matters: Proxy-based setup means a base-URL change and one header, not an SDK migration. - 4Small teams that just need logging and prompt versioning → PromptLayer
Why it matters: Leanest and cheapest at small scale; skip it if you need dataset-driven evaluation. - 5Teams still choosing a model provider → PromptQuorum first
Why it matters: Benchmark your prompt across 25+ models before investing engineering time in production tracing for one of them.
Skip This If…
Skip all four of these tools if you have no LLM feature in production yet — pre-deployment testing and evaluation (Promptfoo, Braintrust, Vellum — see Prompt Testing & Evaluation Tools 2026) solves a real problem before you ship; observability tooling solves a problem you don’t have yet. Wiring up tracing before you have production traffic to trace is wasted setup time, and the free tiers here are generous enough that adding one later, once you actually have live traffic, costs nothing but a day of integration work.
Common Mistakes
❌ Adding production tracing before the prompt has been tested pre-deployment
Why it hurts: You get clean traces of a prompt that was never systematically evaluated — you can see it ran, not whether the output was good.
Fix: Build an evaluation dataset and run it through Promptfoo or Braintrust before wiring up LangSmith, Langfuse, Helicone, or PromptLayer for production traffic.
❌ Choosing LangSmith for a non-LangChain stack
Why it hurts: You lose the automatic-tracing advantage that makes LangSmith worth its price and end up manually instrumenting a generic SDK anyway.
Fix: If you’re not on LangChain/LangGraph, evaluate Langfuse or Helicone first — both integrate as broadly without requiring the framework.
❌ Underestimating LCU/LSU usage-based charges on LangSmith Plus
Why it hurts: The $39/seat headline price does not include compute and storage units, which can exceed the seat cost for high-volume or long-chain workloads.
Fix: Run a usage estimate from your actual trace volume before committing to Plus at team scale, or start on the free Developer tier to gauge real consumption first.
❌ Self-hosting Langfuse without budgeting for the operational cost
Why it hurts: "Free to self-host" still requires someone to run upgrades, backups, and uptime — teams sometimes treat this as a zero-cost decision and are surprised by the ops burden.
Fix: Compare the engineering hours to run Langfuse yourself against the $29–$199/month Cloud tiers before assuming self-hosting is the cheaper option for your team.
❌ Picking PromptLayer for a team that needs dataset-driven evaluation
Why it hurts: PromptLayer’s eval-cell executions are lighter than a full evaluation framework — teams outgrow it fast once they need LLM-as-judge scoring or regression datasets.
Fix: Use PromptLayer for logging and versioning only; pair it with Promptfoo (free) if you need real evaluation depth, or move to LangSmith/Langfuse directly.
How to Choose
- 1Confirm you have live production traffic to trace — if not, start with pre-deployment testing first.
- 2Check your primary framework: LangChain/LangGraph → LangSmith gets automatic tracing; anything else → Langfuse or Helicone.
- 3Decide on self-hosting requirements early — only Langfuse offers a genuinely free, full-featured self-hosted tier.
- 4Estimate monthly event volume against each free tier’s real ceiling (5k–50k events) before assuming "free" covers you long-term.
- 5Start on the free tier of your top pick for 2–4 weeks with real traffic before upgrading to a paid plan.
- 6Re-evaluate at 90 days — pricing and tier limits on all four tools have changed at least once in 2026.
💡 Start Free, Confirm the Fit, Then Pay
All four tools have a usable free tier. Run real production traffic through it for a few weeks before paying for anything — the free tier alone is usually enough to tell you whether the tool’s framework fit and UI match how your team actually works.
Frequently Asked Questions
What is LLM observability?
LLM observability is the practice of logging, tracing, and monitoring large language model API calls after they run in production — tracking what was sent, what came back, how long it took, and what it cost. It answers "what happened in production?" rather than "is this prompt good?", which is the job of pre-deployment testing tools instead.
Is Langfuse really free, or is that just a trial?
Self-hosting Langfuse is genuinely free forever — it is MIT-licensed with no feature gate between the free self-hosted version and the paid cloud tiers. Langfuse Cloud also has a free Hobby tier (50,000 units/month, 2 users) that is not a time-limited trial, though it is capped rather than unlimited.
Should I choose LangSmith or Langfuse if I’m already using LangChain?
LangSmith gets automatic tracing with zero manual instrumentation if you’re on LangChain or LangGraph, which is a real setup-time advantage. Langfuse also integrates with LangChain natively and adds free self-hosting and a broader integration list beyond LangChain — choose LangSmith for the tightest LangChain-specific experience, Langfuse if you want flexibility to add other frameworks or providers later without switching tools.
What does a typical small team pay for LLM observability per month?
A 5-person team on Langfuse Core runs about $29/month. The same team on LangSmith Plus starts at $195/month in seats ($39 × 5) before usage-based LCU/LSU charges. On Helicone, a small team fits inside the $79/month Pro tier. On PromptLayer, a 5-person team fits the free tier up to 2,500 requests/month, or $49/month on Pro beyond that.
How do I set up basic tracing with Langfuse?
Create a free Langfuse Cloud project (or self-host with Docker), install the Langfuse SDK or use one of its native integrations (OpenAI SDK, LangChain, LlamaIndex, LiteLLM, Vercel AI SDK, and others), and add your Langfuse public/secret API keys as environment variables. Traces from instrumented calls appear in the Langfuse dashboard immediately — no separate ingestion pipeline to build.
Is PromptLayer the same as LangSmith or Langfuse?
No. PromptLayer covers request logging and prompt versioning without a full dataset-driven evaluation framework, and it has no self-hosted tier below Enterprise. LangSmith and Langfuse both add mature evaluation features (LLM-as-judge scoring, regression datasets) on top of logging. Use PromptLayer if logging and versioning is genuinely all you need; use LangSmith or Langfuse if you also need to score output quality systematically.
Can I use more than one of these tools at once?
Yes, though it is rarely worth the overlap in cost and maintenance. A common combination is Langfuse or LangSmith for tracing plus evaluation, alongside Helicone purely for its proxy-level cost dashboard on a specific high-volume endpoint. Running all four is redundant — pick one for your primary tracing and evaluation needs.
Does self-hosting Langfuse automatically meet GDPR or HIPAA requirements?
No. Self-hosting keeps data on infrastructure you control, which supports data-residency requirements, but GDPR and HIPAA compliance also depend on your own access controls, encryption, audit logging, and data-processing agreements — self-hosting is necessary for some compliance postures, not sufficient by itself. Langfuse’s Pro cloud tier lists HIPAA availability and SOC2/ISO27001 reports for teams that prefer a compliance-ready managed option instead.
Do these four tools have live affiliate or referral programs?
As of this article’s last fact-check (2026-08-27), none of the four publish a live public commission-based affiliate program on their own sites. LangChain runs a Partner Network for technology and services integrations, which is a partnership structure, not a per-referral commission program. Links in this article are plain product links to each vendor’s own site, with no current affiliate relationship.
Are these prices the same everywhere in the world?
Yes — all four are USD-billed global SaaS subscriptions. LangSmith, Langfuse, Helicone, and PromptLayer do not publish region-specific pricing; a team in the EU, Japan, or Brazil pays the same listed USD price as a team in the US, though currency conversion and card fees from your own bank or payment processor may add a small variable amount on top.
Sources
- LangSmith Pricing — LangChain — official pricing page; basis for Developer/Plus tier limits, $39/seat, and $1.50/LCU + $1.00/LSU usage pricing
- Helicone Pricing — official pricing page; basis for Hobby/Pro/Team tier limits and $79/$799 monthly pricing
- Langfuse Pricing — official cloud pricing page; basis for Hobby/Core/Pro/Enterprise tier limits and pricing
- Langfuse Self-Hosting Pricing — official self-host page; basis for MIT license, free self-hosted core, and Self-Hosted Enterprise feature list
- PromptLayer Pricing — official pricing page; basis for Free/Pro/Team tier limits and no self-hosted tier below Enterprise
- ClickHouse Acquires Langfuse — Official Announcement — basis for the January 2026 acquisition claim
- Langfuse GitHub Repository — basis for the 33.8k GitHub star count and MIT license, verified 2026-08-27
- PromptQuorum — Multi-Model Dispatch — multi-model comparison tool; basis for 25+ model dispatch claims
