Guide
How to Reconcile OpenAI and Anthropic Invoices
The short answer: reconciling AI provider invoices takes two bridges, not one. Bridge one runs from your own per-request cost ledger to the provider's usage and cost API for the same UTC period. Bridge two runs from that API total to the invoice you actually pay. Most teams try to compare their internal number straight to the invoice, find a gap of a few percent, and have no way to tell whether the cause is missing instrumentation, a mispriced token type, or a committed-use credit. Splitting the comparison in two turns one unexplainable number into two explainable ones.
This matters more each quarter. In CloudZero's State of AI Costs research, 57% of surveyed companies still track AI costs in spreadsheets and 15% have no formal tracking at all. A spreadsheet can hold a total. It cannot hold a reconciliation.
Why comparing totals fails
A single total-to-total comparison mixes together two completely different classes of error, and they often move in opposite directions.
Errors in what you recorded make your ledger too low: a batch job that calls the provider with a raw key and never passes your metering point, console usage that has no API key at all, a streaming call that ended without a usage block. Errors in how you priced it can push either way: cache reads charged at the fresh-input rate, a price change that landed mid-month against a hard-coded rate table.
Errors in settlement live entirely on the provider's side: credits, committed-use discounts, tax, and spend that the cost API does not report. None of those are instrumentation problems, and none of them should ever be fixed in your code.
Net them together and a 1.3% gap can be a $318 unmetered cron job hiding behind a $212 retry over-count. You cannot fix what you cannot see separately.
Getting the provider's own numbers
Both providers publish organization-level usage and cost endpoints. Both require an admin credential rather than a normal API key, and both report cost in daily buckets only.
For Anthropic, per the Usage and Cost API documentation:
curl "https://api.anthropic.com/v1/organizations/cost_report?\
starting_at=2026-08-01T00:00:00Z&\
ending_at=2026-09-01T00:00:00Z&\
group_by[]=description&\
group_by[]=workspace_id" \
-H "anthropic-version: 2023-06-01" \
-H "x-api-key: $ANTHROPIC_ADMIN_KEY"
For OpenAI, per OpenAI's own usage API walkthrough:
curl "https://api.openai.com/v1/organization/costs?\
start_time=1754006400&\
bucket_width=1d&\
group_by[]=line_item&\
limit=31" \
-H "Authorization: Bearer $OPENAI_ADMIN_KEY"
The two responses do not look alike, and the differences are the kind that silently produce a hundredfold error:
| OpenAI | Anthropic | |
|---|---|---|
| Cost endpoint | /v1/organization/costs |
/v1/organizations/cost_report |
| Usage endpoint | /v1/organization/usage/completions |
/v1/organizations/usage_report/messages |
| Auth | Admin key, Authorization: Bearer |
Admin API key, x-api-key; workspace keys are rejected |
| Cost granularity | 1d only |
1d only |
| Usage granularity | 1m, 1h, 1d |
1m, 1h, 1d |
| Cost amount format | amount.value, a number, in dollars |
amount, a decimal string, in cents |
| Cost grouping | line_item, project_id |
description, workspace_id |
| Cached-token detail | input_cached_tokens |
uncached, cache read, and 5m / 1h cache creation, separately |
| Documented exclusion | check the invoice for credits and tax | Priority Tier costs are not in the cost endpoint |
Anthropic's cost report returns "amount": "123.45" meaning $1.23, and time buckets are snapped to the start of the day in UTC. OpenAI's costs result returns "amount": {"value": 0.1308, "currency": "usd"} in dollars. Anthropic's docs also note that usage and cost data typically appear within five minutes, that daily buckets are capped at 31 per page, and that console playground usage carries a null API key ID because it is not associated with one.
A worked monthly variance walk
August 2026, Anthropic, one organization. Start from the two totals and work until the residual is under tolerance.
Bridge one: your ledger to the provider's cost API.
| Line | Amount | Cause |
|---|---|---|
| Internal ledger, priced from our own records | $18,412.30 | |
| Console and playground usage | +$96.40 | no API key, never routed through the proxy |
| Nightly enrichment job on a raw provider key | +$318.75 | side-door traffic, never metered |
| Cache-write tokens priced at the base input rate | +$41.15 | cache tiers not captured as separate fields |
| Retries billed per attempt, recorded once | -$212.70 | ledger counted the logical call, not the attempts |
| Explained subtotal | $18,655.90 | |
| Provider cost report total | $18,657.44 | |
| Unexplained residual | $1.54 | 0.008%, within tolerance |
The raw gap here was $245.14, or 1.3% of the provider's number: over the 1% line where a variance stops being rounding and starts being a data problem. After the walk, the unexplained part is under a hundredth of a percent, and three of the four causes are engineering fixes rather than finance adjustments.
Bridge two: the provider's cost API to the invoice.
| Line | Amount |
|---|---|
| Provider cost report total | $18,657.44 |
| Priority Tier spend, not included in the cost endpoint | +$2,140.00 |
| Committed-use credits applied at settlement | -$1,500.00 |
| Invoice total | $19,297.44 |
Nothing on bridge two is a bug. It is the difference between a usage record and a settlement, and it belongs in a memo, not a ticket.
Running the ledger side is a single query when every routed call is one priced row:
SELECT provider, SUM(cost_usd) AS ledger_total, COUNT(*) AS calls
FROM ai_calls
WHERE created_at >= '2026-08-01T00:00:00Z'
AND created_at < '2026-09-01T00:00:00Z'
GROUP BY provider;
Two details decide whether that query is trustworthy. The timestamp has to be the server's, because a client-supplied one lets a caller move spend between months. And the boundaries have to be UTC on both sides, because the provider snaps its buckets to UTC days whether or not your warehouse does.
Where the numbers come from in the first place
Reconciliation is only as good as the recording underneath it. The gap causes above are almost all attribution failures showing up a month late: traffic that bypassed the metering point, token types that were never captured as separate fields, timestamps that were not server-controlled. That is the case for metering at a single choke point rather than in each service, which is worked through in how to track LLM costs per customer. Pricing each request correctly needs the current rate table and its multipliers, broken down for one provider in Anthropic API pricing.
Reconciliation is also step one of a larger process. Once the variance is explained, the period still needs an allocation completeness check, documented adjustments, and a lock. That full sequence is in the AI month close.
And the loop closes back into control. A reconciliation that keeps surfacing the same unmetered service is telling you that spend is reaching providers outside any budget, which means it cannot be capped before the call either. The scopes and the enforcement path are covered in LLM budget enforcement.
Common failure modes
- One bridge instead of two. A single ledger-versus-invoice number cannot distinguish an instrumentation bug from a credit.
- Unit mismatches. Cents-as-string on one provider, dollars-as-number on the other. This produces errors of exactly 100x, which are easy to spot and embarrassing to ship.
- Local-time month boundaries. Provider buckets are UTC. A warehouse running in local time moves the first and last day of every month.
- Explaining the gap in a Slack thread. If the explanation is not stored beside the closed period, next month's reconciler starts from zero and an auditor has nothing to read.
- Editing history to make it tie. Corrections belong as new adjustment rows against an append-only ledger. A ledger you can edit cannot prove anything, which defeats the point of reconciling it.
- Reconciling only the provider you worry about. Multi-provider stacks tend to have one well-instrumented path and one that nobody has looked at since it was added.
FAQ
How do you reconcile an OpenAI or Anthropic invoice with your own records? Two bridges. Ledger to the provider's usage and cost API for the same UTC period, then that API total to the invoice. Name a cause for every difference above your tolerance.
Why doesn't my computed LLM cost match the provider invoice? Usually unmetered side-door traffic, console usage with no API key, retries billed per attempt, cache token types priced at the fresh-input rate, a stale rate table, or mismatched period boundaries.
Can you get OpenAI and Anthropic cost data programmatically? Yes, through organization-level cost and usage endpoints on both providers. Both need an admin credential, both report cost in daily buckets, and their amount formats differ (dollars as a number versus cents as a string).
What variance between invoice and internal records is acceptable? Under about 0.5% is normal; above 1% points to a data collection problem. What matters is that everything above the threshold has a written cause.
Does the provider cost API total equal the invoice? Not always. Anthropic documents that Priority Tier costs are not in the cost endpoint, and credits, committed-use discounts, and tax land at settlement rather than in the API.
Sources and method
Endpoint paths, parameters, granularity limits, amount formats, and the Priority Tier exclusion come from Anthropic's Usage and Cost API documentation and its cost report API reference, and from OpenAI's usage and costs API walkthrough. Adoption figures are from CloudZero's State of AI Costs research. The two-bridge framing, the variance categories, the 0.5% and 1% thresholds, and the worked example are our own method and judgment from building and operating Spendline's AI spend governance proxy; the figures in the example are illustrative. Provider billing behaviour around credits and tax varies by contract, so verify yours against your own invoice rather than assuming the pattern above. Last updated: September 2026.
Want your provider invoices split by the customers who caused them? Book a free 30-minute call. No integration is needed: we go through where your AI cost lands by customer and workflow, and how you would find the accounts that cost more than they pay. If the call turns up a real gap, we offer a free 60-day pilot on your own traffic. Book the 30-minute call
Not ready to talk? Take the 5-minute assessment instead.