Guide
Helicone Alternatives for Teams That Need Financial Governance
The short answer: if you used Helicone for tracing, prompts, and evaluation, the closest replacements are Langfuse and Arize Phoenix. If you used the gateway half for multi-provider routing and rate limits, look at LiteLLM and Portkey. And if the reason you are shopping is that your CFO wants cost per customer, budgets that actually stop spend, and a number that survives an audit, none of those categories is the answer, because none of them was built for it. This guide maps all three.
We sell in the third category, so read the comparison with that in mind. Every factual claim below is checkable against a primary source, and we say plainly where an observability tool is the better buy.
Why teams are looking in the first place
On March 3, 2026 Mintlify acquired Helicone. Both companies were direct about what that means for the product. Helicone's own post says its "services will remain live for the foreseeable future in maintenance mode," and that "security updates, new models, bug & performance fixes all keep shipping." Mintlify added that it would "work closely with every customer to support a smooth migration to another platform."
That is a clear statement, and it is not a shutdown. Read it accurately: the service works, it gets fixed, and the roadmap ended. The open source repository is still there under the Apache 2.0 licence, so self-hosting remains available.
So the honest reason to move is not panic. It is that a tool in maintenance mode will not grow into a requirement it does not already meet, which makes this the right moment to ask what you needed it to do.
What Helicone does well, stated fairly
Helicone is an LLM observability platform with an AI gateway attached, and it is good at both: request logging and dashboards, sessions and segmentation, prompt management, datasets, a playground, alerts, and a query language over your own request data. Pricing is public, with retention and ingestion rate stepping up per tier: a free Hobby tier at 10,000 requests a month, Pro at $79 a month, Team at $799, and custom enterprise terms.
It also does more cost control than most observability tools. Helicone's custom rate limits accept a policy header in the form [quota];w=[window];u=[unit];s=[segment], and the unit can be cents instead of requests. So 100;w=3600;u=cents;s=user caps a user at one dollar an hour. That is real spend control, it works through the gateway rather than async logging, and any comparison that claims observability tools cannot limit spend at all is wrong.
The distinction that matters is narrower, and it is the subject of the rest of this guide.
A rate limit is not a budget
A rolling window quota answers "how fast may this spend?" A budget answers "how much may this cost, in total, this month, and what happens at the cap?" They look similar and behave completely differently.
Take an agent product with 40 customers and a $9,000 monthly AI cost target. A limit of one dollar per hour per user is, at 40 customers running continuously, an implicit ceiling near $28,800 a month, and it never tells you that. It also has no view of the month: a customer that idles for three weeks then runs hard for four days stays inside every rolling window while blowing the plan for that account. The ceiling exists only per segment, so the organisation total is whatever the arithmetic happens to produce.
A budget is the other shape entirely:
org "acme-inc" $9,000 / month action: alert at 80%, block at 100%
├─ team "support-ai" $4,000 / month action: reroute to a cheaper model
│ └─ agent "ticket-triage" $900 / month action: block
└─ customer "cust_4821" $250 / month action: require approval to exceed
Each level is checked before the request is forwarded, the tightest binding scope wins, and hitting a cap has a defined consequence other than a 429. The scopes, the available actions, and the concurrency trap that makes naive implementations overspend are worked through in LLM budget enforcement.
The comparison
| Langfuse | Arize Phoenix | LiteLLM | Portkey | Spendline | |
|---|---|---|---|---|---|
| Primary job | Tracing, prompts, evals | OTEL-native tracing and evals | Self-hosted multi-provider proxy | Managed AI gateway | AI spend governance |
| Open source | Yes, self-hostable | Yes, Elastic Licence 2.0 | Yes | Core gateway | No |
| Sits on the request path | No, SDK beside it | No | Yes | Yes | Yes |
| Cost caps | Spend alerts by email | Not a feature | Budgets per key, user, team, model, customer | Budget limits on provider integrations, cost or token | Hierarchical budgets, org to customer |
| Action at the cap | Notify | Not applicable | Request fails | Key expires | Block, reroute, degrade, or require approval |
| Per-customer margin | No | No | No | No | Yes, with revenue mapped in |
| Invoice reconciliation and period lock | No | No | No | No | Yes |
| Public entry price | Free tier, Core $29/mo | Free, self-hosted | Free, self-hosted | Free tier | Contact us |
Two rows are the whole decision. "Sits on the request path" decides whether a tool can refuse a call before the money is spent; "action at the cap" decides whether the refusal is useful. Everything else is preference.
The cost-cap row deserves detail. LiteLLM has the broadest budget scopes of the open source options: max_budget with a budget_duration, set globally or per key, user, team, team member, model, or end customer, and requests fail once a key crosses its budget (budgets need a database-backed deployment). Portkey sets cost or token limits on provider integrations, resetting weekly, monthly, or not at all, and expires the key at the limit; it is on the enterprise plan and select Pro accounts. Langfuse alerts on cost metrics to Slack, a webhook, or GitHub Actions, and separately watches your own Langfuse invoice; both are monitoring rather than a control, and it says so. That trade-off, and the alternatives to Langfuse itself, are worked through in Langfuse alternatives.
Migrating the attribution, not just the traces
The part of a migration that quietly breaks is the metadata. Helicone attribution rides on its own headers, and everything downstream (per-customer numbers, margin analysis, month end) is built on those identifiers. Map them deliberately rather than by search and replace.
| What it identifies | Helicone header | Spendline header |
|---|---|---|
| The paying account | Helicone-User-Id (or a custom property) |
x-customer-id |
| The product surface | Helicone-Property-Feature |
x-workflow-id |
| One user action's fan-out | custom property | x-agent-id |
| Platform auth | Helicone-Auth |
x-spendline-key |
The integration itself stays a base URL change plus metadata injection:
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://www.spendline.ai/v1",
apiKey: process.env.OPENAI_API_KEY,
defaultHeaders: {
"x-spendline-key": process.env.SPENDLINE_API_KEY,
},
});
const response = await client.chat.completions.create(
{ model: "gpt-5.2", messages },
{
headers: {
"x-customer-id": account.id, // the paying account, set server-side
"x-workflow-id": "ticket_triage",
"x-agent-id": agentRun.id, // groups this run's calls
},
}
);
One rule survives every migration: set those identifiers server-side. An identifier accepted from the client is a spoofable billing input, and so is a client-supplied timestamp, which is how spend lands in the wrong month. The full metadata contract is in how to track LLM costs per customer.
Where this connects to the close
Enforcement and reconciliation are the two things that turn cost data into an accounting record, and they are the reason a governance layer exists at all next to a perfectly good observability tool.
Budgets get checked before the provider call, so a runaway agent is an in-flight problem with an in-flight answer rather than a line item you find on the fifteenth. Every routed call lands in an append-only ledger, so corrections are new rows and nothing is edited after the fact. At month end that ledger is compared against the OpenAI and Anthropic invoices, gaps are investigated, adjustments are posted, and the period is locked so the number stops moving. That workflow is described in the AI month close, and where each layer stops is mapped in gateway vs. observability vs. governance.
Common failure modes when migrating
- Replacing like for like without asking why. A maintenance-mode service that still ships fixes may be fine for another year. Migrate for a requirement.
- Losing the tag history. Export before you cut over. A January to September per-customer series with a gap in it is not a series.
- Assuming one tool covers both halves. Most replacements are observability or gateway, not both, and teams that plan for one discover the other three weeks later.
- Treating a rate limit as a cap. Covered above, and the most expensive misreading in this category.
- Leaving a side door open. A batch job calling the provider directly appears in no tool's numbers. Metering on the network path is the only structural answer.
- Client-set identifiers. They survive migrations because nobody re-reads them, and they corrupt every budget window downstream.
FAQ
Is Helicone shutting down? No. Mintlify acquired it on March 3, 2026 and both companies said the service stays live in maintenance mode, with security updates, bug fixes, and new model support continuing. Feature development stopped, and Mintlify offered migration support. The repository is still available under Apache 2.0.
What is the closest replacement for Helicone? It depends which half you used. Langfuse and Arize Phoenix for tracing, prompts, and evaluation; LiteLLM and Portkey for the gateway. No single product replaces both halves, so plan for two.
Can Helicone enforce a spending cap? Partially. Its cost-based rate limits set a quota in cents over a rolling window, segmented globally, per user, or per custom property, through the gateway integration. That caps a rate of spend rather than a monthly total across an organisation and its customers.
Do observability tools track cost per customer? Through optional tags, which is useful for debugging and weak for accounting. They model the session or end user rather than the paying account, and nothing flags a code path that forgot the identifier.
Should we migrate off Helicone right now? Not urgently. Move when you hit a requirement it was never built for: budgets enforced before the provider call, per-customer margin, or a ledger that reconciles against the invoice.
Sources and method
Acquisition facts and quotations come from the two primary announcements, Mintlify's and Helicone's, both dated March 3, 2026. Helicone's features, tiers, and rate limit behaviour come from its pricing page, its rate limits documentation, and its repository, all checked in September 2026. Budget behaviour for other tools comes from their own docs: LiteLLM, Portkey, Langfuse, and Arize Phoenix. Vendor features and prices move; check them before deciding. The two-axis framing and the rate-limit-versus-budget distinction are our categorisation judgment, not anyone's published position. Last updated: September 2026.
Comparing tools because you need cost per customer in dollars, not only per request? Book a free 30-minute call. No integration is needed: we go through where your AI cost lands by customer and workflow, and how you would find the accounts that cost more than they pay. If the call turns up a real gap, we offer a free 60-day pilot on your own traffic. Book the 30-minute call
Not ready to talk? Take the 5-minute assessment instead.