Guidefinanceunit economicsgross marginFinOps

How to Calculate AI Gross Margin by Customer

August 25, 2026 · Spendline

The short answer: AI gross margin for one customer is the revenue you recognized from that account in a period, minus the AI cost attributed to that account in the same period, minus your other variable costs of revenue, divided by the revenue.

ai_gross_margin_% = (recognized_revenue - attributed_ai_cost - other_variable_cost) / recognized_revenue

The arithmetic is a subtraction. (For a rough first pass before any of the plumbing below, the free AI margin calculator runs it from six numbers, in your browser.) The reason most teams cannot produce the number is that the two sides live in different systems, keyed differently, over different windows. This guide covers which revenue number to use, how to allocate AI cost that has no customer attached, a worked example on current model rates, how to track margin by cohort, and what to do about accounts that come back negative.

Why the naive approach fails

The usual first attempt divides the total AI bill by total revenue. That produces a real number and it is useless for every decision you would make with it, because aggregate margin describes a customer who does not exist: the average one. In an AI-native product the cost distribution is long-tailed. A handful of accounts drive most of the inference, and averaging spreads their cost onto customers who never incurred it. A company at a comfortable aggregate margin can contain accounts running at negative contribution, and the aggregate will never say so. Dividing by seat count is worse, because it assumes usage tracks the billing unit: two accounts on the same plan routinely differ by an order of magnitude, because one runs the agentic feature and the other does not.

The gap between AI and classic software economics is structural rather than a rounding error. Andreessen Horowitz's analysis of AI business models put AI gross margins "often in the 50-60% range, well below the 60-80%+ benchmark for comparable SaaS businesses," attributing part of it to "the 25% or more of revenue that AI companies often spend on cloud resources." Bessemer's State of AI 2025 reports its fastest-growing AI cohort at roughly 25% gross margins, with a more sustainable cohort near 60%. When the category sits in that band, the distribution inside your own base decides which customers are worth having. Survey averages from 2024 onwards, and what moves them, are collected in what AI did to software gross margins.

Getting the cost side right

The cost side is mechanical, provided every request reaching a model carries a server-set customer identifier. That contract is covered in how to track LLM costs per customer. With it in place, cost is one query. Two definitions still need deciding.

What counts as cost of revenue. Production inference does. Evaluation runs, model comparisons, and internal experiments generally do not. Fine-tuning is the contested case, and classifying it wrongly moves gross margin by points, so it needs a written rule rather than a habit (the decision tree).

What else is variable. Vector store reads and writes, embeddings, document storage, egress, and support cost that scales with usage. Small next to inference, and omitting them inflates every margin you report.

Getting the revenue side right

This is where the calculation actually breaks, and it breaks quietly.

Use recognized revenue, not invoiced. An annual contract invoiced in January, matched against January's inference, shows 96% margin in January and deep losses for the eleven months after. Amortize the subscription across the term; add usage revenue in the month it was earned.

Net out discounts and credits per account. A discount negotiated on one account changes that account's margin and nobody else's. An average discount rate moves margin from the customers who got it onto the ones who did not.

Do not report free accounts as negative-margin customers. Trials and design partners have real cost and zero revenue, so the ratio is negative by construction. That spend is acquisition or product development. Report it separately, and do report it, because in early AI products it is often material.

Split multi-product revenue. If an account buys two products and only one is AI-heavy, blending revenue makes the AI product look profitable on the strength of the other.

A worked example

One account, one month, published rates. The product runs a Claude Sonnet 5 agent with a Claude Haiku 4.5 pre-filter, prompt caching, and the server-side web search tool. Rates from Anthropic's published pricing as of August 2026.

Line Volume Rate Cost
Sonnet 5 uncached input 120M tokens $2 / MTok $240
Sonnet 5 cache reads 900M tokens $0.20 / MTok $180
Sonnet 5 cache writes (5m) 40M tokens $2.50 / MTok $100
Sonnet 5 output 150M tokens $10 / MTok $1,500
Haiku 4.5 input 200M tokens $1 / MTok $200
Haiku 4.5 output 14M tokens $5 / MTok $70
Web search 9,000 searches $10 / 1,000 $90
Attributed AI cost $2,380

The account is on a $4,000 monthly plan and earned $600 of usage revenue, so recognized revenue is $4,600. Other variable cost (vector store, storage, egress) is $265.

gross_profit = 4,600 - 2,380 - 265 = 1,955
gross_margin = 1,955 / 4,600 = 42.5%

Waterfall showing one customer for one month: $4,600 of recognized revenue reduced by $1,570 of output tokens, $440 of input tokens, $280 of cache reads and writes, $90 of server-side tools and $265 of other variable cost, leaving $1,955 of gross profit at 42.5% margin.

The composition is the useful part. Output tokens are $1,570 of the $2,380, about 66% of this account's AI cost, because output bills at five times input on every Claude tier and an agent that reasons and writes produces a lot of it. An optimization program aimed at prompt length would be working on a third of the problem, and no aggregate margin number would have told you that.

Allocating the cost you cannot attribute

Some AI spend has no customer on it: shared embedding jobs, index rebuilds, background classification, internal tooling. How you treat it changes every per-customer number you publish.

Approach Effect on per-customer margin When it is right What it hides
Leave unallocated Margin reported before shared AI cost; the pool is its own line The pool is small, or genuinely a platform investment Per-customer margins read higher than company gross margin
Spread pro rata by attributed usage Heavy users absorb most of it Shared cost scales with usage (index rebuilds, enrichment) Little, if the driver really is usage
Spread pro rata by revenue Large accounts absorb most of it Shared cost is a fixed capability serving everyone Whether small heavy users are unprofitable

The rule matters more than the choice: apply the same method every period and state which one, next to the number. A margin series that changes allocation method mid-year is not a series.

From formula to query

Once per-request cost is a ledger and revenue is a monthly table keyed to the same customer identifier, it is one join:

WITH ai AS (
  SELECT customer_id,
         date_trunc('month', created_at) AS period,
         SUM(cost_usd) AS ai_cost
  FROM ai_calls
  WHERE created_at >= '2026-08-01' AND created_at < '2026-09-01'
  GROUP BY 1, 2
)
SELECT r.customer_id,
       r.recognized_revenue_usd,
       COALESCE(ai.ai_cost, 0) AS ai_cost,
       r.other_variable_usd,
       r.recognized_revenue_usd - COALESCE(ai.ai_cost, 0)
         - r.other_variable_usd AS gross_profit,
       ROUND(100.0 * (r.recognized_revenue_usd - COALESCE(ai.ai_cost, 0)
         - r.other_variable_usd)
         / NULLIF(r.recognized_revenue_usd, 0), 1) AS gross_margin_pct
FROM revenue_by_month r
LEFT JOIN ai ON ai.customer_id = r.customer_id AND ai.period = r.period
WHERE r.period = '2026-08-01'
ORDER BY gross_margin_pct ASC;

Two details carry the correctness. created_at must be a server-set timestamp, or a client can move cost between months and your margin series drifts from your books. And the ledger must be append-only, corrections posted as new rows, or a margin reported in August can silently change in October.

Tracking margin by cohort

A single month tells you where you are. Cohorts tell you where you are going, and for AI products the direction is usually down, because adoption of the expensive features grows faster than the contract does. Track a signup cohort forward, computing margin after AI cost at each month of tenure. An illustrative shape, built to mirror what adoption curves do:

Cohort month Accounts Avg recognized revenue Avg AI cost Margin after AI cost
Month 1 38 $3,900 $1,010 74%
Month 3 38 $4,150 $1,660 60%
Month 6 35 $4,320 $2,540 41%

Revenue per account rose 11% over six months. AI cost per account rose 151%. Every snapshot looked acceptable alone, and the trend is what ends the business. That is why the ratio of AI cost growth to revenue growth is a better early warning than the margin level, and why this review belongs on a monthly cadence. Calibration ranges for what healthy looks like are in AI has broken SaaS unit economics.

What to do with a negative-margin account

Work down this list. Skipping to the bottom is how teams reprice a contract to fix what was actually a bug.

  1. Rule out a defect. A retry loop, a runaway agent, or a workflow re-embedding the same corpus nightly is an engineering problem wearing a pricing problem's clothes. Look at cost per successful outcome, not cost per month.
  2. Change the model and the caching. Route classification and extraction to a cheaper tier, cache the stable prefix, batch what is not latency-sensitive. Usually the largest reversible win.
  3. Cap the account. A hard budget converts an unbounded loss into a known one, and it takes effect on the next request rather than at renewal.
  4. Reprice at renewal. If the account uses the product exactly as designed and still loses money, the pricing is wrong for that usage pattern.
  5. Offboard. Rarely necessary if the account was caught in month two rather than month twenty.

How this connects to enforcement and month close

A margin number is a measurement, and measurements stop nothing. Two things turn it into control.

Enforcement. Budgets evaluated before the provider call, scoped to org, team, agent, or customer, so an account heading through its margin floor is blocked, rerouted, or sent for approval rather than reported on later. The scopes, the actions available at a cap, and the concurrency race that defeats naive implementations are in LLM budget enforcement.

Month close. A margin you cannot tie to the provider invoice is an estimate. Closing the month means reconciling the ledger against each bill, investigating the gap, posting corrections as new records, and locking the period so August's number still reads the same in October (the workflow).

Common failure modes

  • Invoiced revenue against monthly cost. Annual prepay matched to one month of inference. The most common way this goes wrong.
  • Averaging the AI bill across accounts. A number for every customer and a true number for none.
  • Free accounts reported as negative margin. They are acquisition or R&D spend, not unprofitable customers.
  • An unattributed pool nobody quantifies. If 30% of AI spend has no customer on it, every margin you publish is overstated by an unknown amount.
  • Changing allocation method between periods. Makes the trend, which is the useful part, unreadable.
  • Mutable cost records. If August's ledger can be edited in October, the series cannot be audited.

How Spendline does this

Spendline produces the cost side of this calculation automatically. Every provider call routed through Spendline carries an x-customer-id, is priced at the moment it happens, and lands in an append-only ledger, so AI cost per customer for any month is a query, not a spreadsheet exercise. Load revenue per customer and the dashboard shows the margin left after AI cost on each account, ranks the negative ones, and lets you set a per-customer budget so an unprofitable account is capped in the request path rather than discovered at close.

The month-close workflow locks the period once the ledger reconciles to the provider invoice, so the margin figure you report is the one that survives audit. Setup is a base URL change plus headers; see the integration page.

FAQ

How do you calculate AI gross margin for a single customer? Recognized revenue for the period, minus AI cost attributed to that customer in the same period, minus other variable cost of revenue, divided by revenue. The difficulty is getting both sides keyed to the same customer and window.

Which revenue number should you use: invoiced, billed, or recognized? Recognized. Invoiced revenue puts an annual contract's full value against one month of inference and makes eleven subsequent months look like losses.

How should you handle AI costs that cannot be attributed to a customer? Leave them in a separate platform line, spread them pro rata by attributed usage, or spread them pro rata by revenue. Any of the three is defensible. Changing between them, or never quantifying the pool, is not.

What is a healthy AI gross margin per customer? Judge the distribution, the trend, and your floor rather than a single benchmark. Published figures put AI-native gross margins meaningfully below classic SaaS, so a healthy average can still hide accounts below zero.

What do you do about a customer with negative AI gross margin? Rule out a defect, then optimize model choice and caching, then cap the account, then reprice at renewal. Offboarding is last and usually avoidable if the account is caught early.

Sources and method

Written from building and operating Spendline, an AI spend governance proxy that records per-request cost with customer attribution. External figures: the AI and SaaS gross margin comparison and cloud spend as a share of revenue from Andreessen Horowitz's "The New Business of AI"; AI cohort gross margins from Bessemer's State of AI 2025; model and tool rates from Anthropic's published pricing, checked August 2026. The worked example, the cohort table, the allocation comparison, and the negative-margin sequence are our own construction and judgment, built to mirror patterns we see in practice rather than drawn from survey data. Last updated: August 2026.


Want your margin by customer before you build the pipeline? Book a free 30-minute call. No integration is needed: we go through where your AI cost lands by customer and workflow, and how you would find the accounts that cost more than they pay. If the call turns up a real gap, we offer a free 60-day pilot on your own traffic. Book the 30-minute call

Not ready to talk? Take the 5-minute assessment instead.