In October 2023, the Wall Street Journal reported that GitHub Copilot — which charged individuals $10 a month — was losing an average of $20 a month per user in the first months of that year. Some users cost Microsoft as much as $80 a month.
That story became the founding myth of AI unit economics: the more people love it, the more you lose. It's still a useful warning. But it hides the more interesting part — how much of your margin is decided by choices nobody files under "pricing."
The margin gap is real
The data has caught up with the anecdote.
- Bessemer, State of AI 2025 — the fastest-growing AI startups, which it calls "Supernovas," average about 25% gross margin. Steadier "Shooting Stars" average about 60%.
- ICONIQ, State of AI 2026 — AI product gross margins rising from 45% in 2025 to a projected 53% in 2026 and 59% in 2027.
- Duolingo, Q1 2025 — gross margin of 71.1%, down from 73.0% a year earlier, attributed to "increased generative AI costs related to the expansion of our Duolingo Max tier."
- Microsoft, quarter ending December 2025 — Microsoft Cloud gross margin fell to 67%, "driven by continued investments in AI infrastructure and growing AI product usage."
For years, SaaS didn't have to think about this. One more user cost almost nothing to serve. I wrote about what breaking that assumption does to your pricing model in Pricing is the feature you forgot to ship. This piece is the other side of the same coin: the cost line — and who actually controls it.
The levers are on your provider's price list
Model providers sell the same tokens at very different prices depending on how you ask for them. From Anthropic's pricing page:
- Prompt caching — store a repeated block of context once, then reuse it. Writing to the cache costs 1.25x the normal input price for a 5-minute cache, or 2x for a 1-hour cache. Every read after that costs 0.1x — and on Anthropic's newest models, 0.05x or less.
- Batch processing — anything that can wait gets "a 50% discount on both input and output tokens." And the two discounts "can be combined."
OpenAI's pricing page shows the same shape: cached input at a tenth of the normal price on its current GPT-5.6 models, and "Save 50% on inputs and outputs with the Batch API."
Here's what that does to one workload. Say your product sends the same large context — a codebase, a contract, a long set of instructions — 100 times within the cache window.
| Cost, in units of one normal read | |
|---|---|
| No caching | 100 |
| Caching: 1 write at 1.25 + 99 reads at 0.1 | 11.15 |
Same work. Same output. About 89% less on that input.
Agent products live inside this pattern. The model doesn't remember the conversation between calls, so every turn sends the whole history again. That's repeated context — the exact thing caching discounts.
The decision nobody puts on the roadmap
Once engineering ships caching, the cost of serving a customer drops. And a question appears that nobody scheduled: who gets the savings?
Keep the arithmetic simple and assume the whole bill is input. You charge a customer $120 for a workload that costs you $100 at standard rates — a 17% margin. Now 90% of that input is served from cache. The provider bill drops to about $21.50: $9 for the cached reads, $12.50 for writing the rest to cache. Charge the same $120, and your margin is about 82%.
You have two honest options.
| Pass the savings down | Keep the savings | |
|---|---|---|
| What you do | Lower the price, or charge cost plus a fixed markup | Charge the standard rate; keep the efficiency |
| What the customer sees | A bill that shrinks over time | No change |
| Your margin | Thin, driven by volume | Wide |
| When it fits | You're fighting for share; your buyers are developers who benchmark unit costs | You're differentiated; your buyers pay for outcomes, not tokens |
Neither is wrong. The mistake is not choosing — and most teams don't, because something else chooses for them: the pricing unit.
Charge per token at cost plus a markup, and the savings flow to the customer automatically. Charge per seat, per task, or per outcome, and you keep them. Nobody sat in a room and decided that. It fell out of the unit.
Your pricing unit decides who keeps your engineering savings — so pick the unit knowing that.
Two costs that move the other way
Margins don't only improve on their own. Two things push in the other direction, and they rarely show up in a pricing review.
- New tokenizers. When Anthropic launched Claude Opus 4.7, it noted that "the same input can map to more tokens—roughly 1.0–1.35× depending on the content type." Same text, same price per token, bigger bill. Every model upgrade needs a cost-per-task benchmark, not just a quality check.
- Tool fees. Agents pay for their tools on top of tokens. Anthropic's web search costs $10 per 1,000 searches, plus the tokens it pulls in. Code execution runs $0.05 per container-hour once the 1,550 free monthly hours are used up. Thousands of agents searching all day is a line item nobody modeled.
What to bring to your next pricing review
- Cache hit rate per product surface — is repeated context actually being reused?
- Workloads that could run in batch and don't — reports, indexing, anything nobody waits for.
- Gross margin per feature, not just per account — know which features earn and which leak.
- A cost-per-task benchmark after every model change.
- A written decision — pass the savings down or keep them, and which pricing unit enforces it.
Your next margin gain might not come from a price increase. It might already be sitting on your provider's price list, waiting for someone in product to claim it.
Who on your team decided what happens to the savings from caching — and did anyone write it down?
Free companion
The Credit Engine — Infrastructure Playbook
The infrastructure behind credits-based pricing in a single PDF — the ledger, metering, and enforcement patterns that decide whether your margins survive once engineering ships the savings.
Get the playbook →This is one piece of a longer framework I teach in Chapter 5 of Product Strategy in the AI Era — including IQ tiering: gating your most expensive models behind higher tiers, so a power user on a cheap plan can't run up your bill.
Sources
- AI Business, "GitHub Copilot loses $20 a month per user" (October 11, 2023): aibusiness.com
- TechRadar, "Microsoft is reportedly losing huge amounts of money on GitHub Copilot" (October 10, 2023): techradar.com
- Bessemer Venture Partners, "The State of AI 2025" (August 13, 2025): bvp.com
- ICONIQ, "State of AI 2026": iconiq.com
- Duolingo Q1 2025 shareholder letter (SEC filing): sec.gov
- Microsoft FY26 Q2 performance: microsoft.com
- Anthropic, Claude API pricing: claude.com
- Anthropic, "Introducing Claude Opus 4.7" (April 16, 2026): anthropic.com
- OpenAI, API pricing: openai.com