GuidetoolingLLM observabilitycomparisonFinOps

Langfuse Alternatives: Tracing, Evals, and Who Can Actually Stop Spend

September 22, 2026 · Spendline

The short answer: if you want a like-for-like replacement for Langfuse, look at Comet Opik and Arize Phoenix, both open source and both self-hostable. If you are shopping because Langfuse showed you a cost number you did not like and you now want something that can act on it, no observability tool is the answer, including the ones that replaced it. The tool that stops spend has to sit on the request path, and tracing SDKs deliberately do not.

We sell in that second category, so read this with that in mind. Every claim below is checkable against a primary source, and we say plainly where Langfuse or one of its alternatives is the better buy.

Why people go looking

Langfuse is not in trouble. It is actively developed, MIT licensed for every core product feature including tracing, evaluations, prompt management, experiments, annotation, and the playground, with only SCIM, audit logging, and data retention policies held back as paid enterprise add-ons. That is one of the more generous open source positions in this category, and it removes the usual reason to leave.

So the searches that lead here are mostly one of three:

  1. Price at volume. Langfuse Cloud starts free at 50,000 units a month, then Core at $29 a month, Pro at $199, and Enterprise at $2,499, with usage above the included allowance billed per 100,000 units. Tracing is priced by how much you trace, so the bill grows with exactly the thing you are trying to watch.
  2. A different technical taste. Teams standardizing on OpenTelemetry often prefer Phoenix; teams already inside Comet or LangChain often prefer Opik or LangSmith.
  3. A question tracing cannot answer. Someone asked what a specific customer costs, or asked for the spending to stop at a number, and traces do not do either.

The first two are shopping. The third is a category error, and it is the expensive one, because you can migrate between observability tools for a year without getting closer to an answer.

What Langfuse does well, stated fairly

Langfuse is an LLM engineering platform: traces and spans across an agent run, versioned prompt management, datasets and experiments, automated and human evaluation, annotation queues, and a playground. Token and cost tracking is attached to each observation, so you can see what a trace cost and slice it by model, environment, tag, or user.

It also alerts on cost, which is more than many of its peers. Langfuse alerts evaluate metrics over a window, including cost aggregations, and route notifications to Slack, a webhook, or GitHub Actions. Separately, spend alerts watch your own Langfuse subscription invoice against a threshold you set.

That is useful, and it is still only notification: nothing in it blocks, caps, degrades, or reroutes a request. Which is not a criticism. Langfuse sits beside the request path as an SDK, so there is no moment at which it could refuse a call.

Alerting is a postmortem

This is the distinction the whole comparison turns on, so it is worth being concrete.

An alert fires after the tokens are bought. For a slow drift in monthly spend that is fine, because a human can act inside the feedback loop. For the failure that actually hurts it is useless: a retry loop or a runaway agent burns hundreds of dollars in minutes, and a notification at the 80% line describes an event that has already finished.

Enforcement is a different shape. The budget is evaluated before the request is forwarded, and reaching the cap has a defined consequence:

org "acme-inc"                 $9,000 / month   action: alert at 80%, block at 100%
├─ team "support-ai"           $4,000 / month   action: reroute to a cheaper model
│   └─ agent "ticket-triage"     $900 / month   action: block
└─ customer "cust_4821"          $250 / month   action: require approval to exceed

Two properties of that tree are doing the work. Every level is checked in the request path, so the tightest binding scope wins before any money is spent. And the last line is a customer, which is the scope an observability tool has no concept of. The mechanics, including the check-then-act race that makes naive versions overspend under concurrency, are worked through in LLM budget enforcement.

The comparison

A map of the Langfuse alternatives on two axes: whether a tool sits beside the request path or on it, and whether it answers the engineering question or the finance question. Langfuse, Opik and Phoenix cluster in the beside-and-engineering quadrant, LangSmith spans the path axis, and only the governance corner covers customers, margin and month close.

Langfuse Comet Opik Arize Phoenix LangSmith LiteLLM Spendline
Primary job Tracing, prompts, evals Tracing and agent testing OTEL-native tracing and evals Tracing plus an LLM gateway Self-hosted multi-provider proxy AI spend governance
Open source MIT core, EE add-ons Apache 2.0 Elastic Licence 2.0 No Yes No
Sits on the request path No, SDK beside it No No Gateway: yes Yes Yes
Cost caps None; alerts only None None Spend caps per org, workspace, user, key Budgets per key, user, team, model, customer Hierarchical budgets, org to customer
Action at the cap Notify Not applicable Not applicable Fallback model Request fails Block, reroute, degrade, or require approval
Per-customer margin No No No No No Yes, with revenue mapped in
Invoice reconciliation and period lock No No No No No Yes
Public entry price Free, Core $29/mo Free, Pro $19/mo Free, self-hosted $0 dev seat, Plus $39/seat Free, self-hosted Contact us

Two rows decide everything. "Sits on the request path" determines whether a tool can refuse a call before the money is spent. "Action at the cap" determines whether that refusal is useful to you. The rest is preference and price.

Three entries deserve more than a table cell.

LangSmith has quietly become the serious answer on this axis. Its LLM Gateway is a proxy: you swap the base URL and keep the rest of your code, and it sets spending caps for organizations, workspaces, users, and API keys, monitors spend at each level in real time, and can trigger a fallback model when a limit is violated. Anyone claiming the tracing vendors cannot enforce anything has not read that page. The honest limit is the scope list: organizations, workspaces, users, and keys are all internal units. None of them is the account that pays you, so it governs your engineering org rather than your customer base. Pricing is per seat, $0 for a single developer seat and $39 for Plus.

LiteLLM has the broadest budget scopes of the open source options, including an end-customer scope: max_budget with a budget_duration, set globally or per key, user, team, team member, model, or customer, with requests failing once a key crosses its budget (this needs the database-backed deployment). If you want to run the in-path route yourself and a hard failure at the cap is acceptable, it is the obvious starting point.

Opik and Phoenix are the true like-for-like swaps. Opik is Apache 2.0, self-hostable on the same codebase as the cloud product, free for 25,000 spans a month and $19 a month for Pro. Phoenix is OpenTelemetry-native under the Elastic License 2.0 and free to self-host. Neither adds cost control, and neither claims to.

What the finance question actually needs

If you moved to any tracing tool and still cannot answer your CFO, it is because three properties are structural rather than missing features.

Coverage is optional. Instrumentation is something each code path has to remember. A new microservice, a cron job, or a batch script that calls the provider directly never appears, and nothing tells you it is missing. For debugging that is survivable. For accounting it is fatal, because a number with an unknown error bar is not one finance can sign. If a call can reach the model without passing your metering point, you have monitoring, not accounting.

The grouping is wrong. Tracing tools model a user or a session. Finance needs the paying account, which is the join key to revenue. A tag can approximate it, but an optional tag is not a contract, and margin work needs cost per customer to be complete rather than mostly complete.

Traces are not a ledger. Trace stores are mutable and have a retention window (30 days on the Langfuse free tier, 90 on Core). A finance record is append-only, so corrections are new rows rather than edits, and it survives past the period it describes because someone will reopen it. That record is what gets compared against the provider invoice at month end and locked when the period closes, a workflow covered in the AI month close.

None of this makes observability tools bad. It makes them a different product answering a different question, which is how their own documentation presents them. Where each layer stops is mapped in gateway vs. observability vs. governance.

What a migration actually costs

The part that breaks quietly is attribution. Observability identifiers are optional metadata; governance identifiers are billing inputs. Map them deliberately rather than by search and replace, and set them server-side, because an identifier accepted from the client is a spoofable input to a number you will invoice on.

What it identifies Langfuse Spendline
The paying account metadata tag or userId x-customer-id
The product surface tags or trace name x-workflow-id
One user action's fan-out sessionId or trace grouping x-agent-id
Platform auth SDK public and secret keys x-spendline-key

The integration itself is a base URL change plus metadata injection for full attribution:

import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://www.spendline.ai/v1",
  apiKey: process.env.OPENAI_API_KEY,
  defaultHeaders: {
    "x-spendline-key": process.env.SPENDLINE_API_KEY,
  },
});

const response = await client.chat.completions.create(
  { model: "gpt-5.2", messages },
  {
    headers: {
      "x-customer-id": account.id,        // the paying account, set server-side
      "x-workflow-id": "support_summarizer",
      "x-agent-id": agentRun.id,          // groups this run's calls
    },
  }
);

Keeping both is the normal outcome, and it is not redundancy. Tracing answers why a response was wrong; the ledger answers what it cost and who owes it. They are not competing for the same seat.

Common failure modes

  • Migrating for the licence, then discovering the gap. Moving from Langfuse to Opik changes your licence and your bill. It does not get you a cap, because neither tool has one.
  • Treating an alert as a control. An 80% notification on a monthly budget cannot stop a retry storm that completes in four minutes. The controls that stop a single run work on the run, not the month.
  • Trusting tags for billing. Optional metadata set by whichever code path remembered it will not survive contact with an auditor.
  • Capping keys and calling it customer governance. Key and workspace limits protect your infrastructure. They say nothing about whether a specific account is profitable.
  • Letting retention quietly delete the evidence. A 30-day window is fine for debugging and short of what a reconciliation and close cycle needs.

How Spendline fits

Spendline is an AI spend governance proxy, which is the in-path route from the table above with the finance parts attached. Point your provider base URL at Spendline, send x-customer-id on every request, and each routed call is priced when it happens and written to an append-only ledger attributed to that customer, across OpenAI, Anthropic and eight other providers in one ledger. Budgets at org, team, agent, and customer level are evaluated before the request is forwarded, and reaching a cap blocks, reroutes, degrades, or requires approval rather than only sending mail. Map revenue in and you get margin per account; at month end the same ledger is what reconciles against the provider invoice and locks.

It is not a tracing tool and does not try to be. If you need to debug a bad agent response, keep Langfuse.

FAQ

What are the closest open source alternatives to Langfuse? Comet Opik (Apache 2.0, same codebase self-hosted as cloud) and Arize Phoenix (OpenTelemetry-native, Elastic License 2.0). Langfuse's own core is MIT licensed, so licensing alone is rarely the reason to move.

Can Langfuse enforce a spending limit? No. It alerts on cost to Slack, a webhook, or GitHub Actions, and watches your own subscription invoice. It sits beside the request path, so it has no point at which it could refuse a call.

Does LangSmith enforce budgets? Yes, within its own scopes. Its LLM Gateway is a proxy that caps spend per organization, workspace, user, and API key and can fall back to another model at the limit. None of those scopes is your paying customer.

Should we replace Langfuse or add to it? Usually add. Tracing answers why a response was wrong, which is a useful job and a different one from what finance asks.

What does an observability tool miss that finance needs? Guaranteed coverage, grouping by the paying account rather than the user or session, and an append-only record that reconciles against the invoice and locks at month end.

Sources and method

Written from building and operating Spendline's AI spend governance proxy, combined with public primary sources, all checked on 22 September 2026: plans and prices from Langfuse pricing, Comet pricing, and LangChain pricing; licensing from Langfuse's open source page, the Opik repository, and the Phoenix repository; alerting behaviour from Langfuse's alerts documentation; gateway behaviour from LangSmith's LLM Gateway page; budget scopes from LiteLLM's proxy documentation. Vendor prices change without notice, so check each page before relying on a figure. The two-axis categorization and the judgment about which gaps are structural rather than incidental are ours. Last updated: September 2026.


Tracing every call already, but still guessing what each customer costs you? Book a free 30-minute call. No integration is needed: we go through where your AI cost lands by customer and workflow, and how you would find the accounts that cost more than they pay. If the call turns up a real gap, we offer a free 60-day pilot on your own traffic. Book the 30-minute call

Not ready to talk? Take the 5-minute assessment instead.