Verdict: OpenAI is not yet profitable, but for the first time it has a credible, multi-layered path to get there. In the second half of 2026 it stopped renting the entire AI stack and started owning two layers of it — the model (with GPT‑5.6 Sol) and the silicon (with the Broadcom‑built Jalapeño inference ASIC). The combination of a flagship model that matches Anthropic's Claude Fable 5 at roughly half the price, plus a custom chip Broadcom's CEO says halves inference cost per token, is the lever that turns a $14 billion inference bill into something OpenAI could carry. The catch — and it is a real one — is Jevons' paradox: cheaper tokens do not mean lower bills when usage explodes, and OpenAI's own Codex agent is a 400%-growth engine built to make usage explode. For builders and small businesses, the real signal is not whether OpenAI turns a profit in Q3 2027 — it is that the token price you pay for frontier intelligence is falling roughly 10× a year and is about to fall again. Plan for an AI bill that is 90% cheaper in 12 months, not one that is stable.
Last verified: 2026-08-03
- OpenAI's annualized revenue run rate sat near $25 billion in May 2026; its projected 2026 operating loss (non‑GAAP) is roughly $14 billion, with a GAAP figure forecasted near $33 billion (FutureSearch, Reuters).
- Sam Altman publicly admitted on January 5, 2025 that the $200/month ChatGPT Pro plan was losing money because people used it more than expected (TechCrunch, Fortune).
- The price of fixed-quality AI performance fell roughly 1,000× between late 2021 and late 2024 — GPT‑3 at $60 per million tokens became Llama 3.2 3B at $0.06 (a16z, LLMflation).
- The GPT‑5.6 family — Sol, Terra, Luna — was publicly released on July 9, 2026 at $5/$30, $2.50/$15, and $1/$6 per million tokens respectively (OpenAI).
- Jalapeño, OpenAI's first custom inference ASIC, was unveiled with Broadcom on June 24, 2026; Broadcom CEO Hock Tan said early testing showed ~50% lower cost per inference token vs. current Nvidia GPUs (Bloomberg, OpenAI).
- Pricing and chip claims are volatile — re-check before any procurement decision.
Why does OpenAI lose money even with $200/month subscribers?
OpenAI loses money on its most engaged subscribers because the marginal cost of serving a heavy user is higher than the subscription fee they pay. In December 2024 OpenAI launched ChatGPT Pro at $200 per month. Two weeks later, on January 5, 2025, Sam Altman wrote on X: "I personally chose the price… and thought we would make some money" — and then confirmed the plan was in fact losing money because people were using it far more than projected (TechCrunch, Fortune). The pattern is not a one-off: OpenAI's Q1 2026 revenue was $5.7 billion on a $3.7 billion cash burn, an operating margin of roughly −122% (Reuters, The Decoder).
The reason is inference cost, which is fundamentally different from training cost. Training a model is a one-time capital expense — you spend billions building the smartest model you can, and then you own it. Inference is the cost that fires every single time a customer asks a question. It scales linearly with usage, and usage is exactly what ChatGPT's 900 million weekly users generate. FutureSearch and The Information both report OpenAI's internal 2026 projection puts inference alone at roughly $14 billion of cost (FutureSearch, The Information). The more a heavy Pro subscriber uses the product, the deeper that subscriber's unit economics go negative.
This is the inverted economics of frontier AI in 2026: in a normal software business, your best customers are your most profitable ones. In an AI subscription business where marginal cost scales with usage and price is fixed, your best customers are your least profitable ones — until you either raise price or radically lower the cost of inference.
Why can't OpenAI just raise the subscription price?
OpenAI cannot raise subscription prices because the AI token market is the most aggressively commoditizing market in software history. The cost of a fixed level of AI performance has dropped approximately 10× per year for the last three years. When GPT‑3 became publicly accessible in November 2021, achieving an MMLU score of 42 cost $60 per million tokens. By late 2024, the cheapest model hitting that same quality bar — Llama 3.2 3B served by Together.ai — cost $0.06 per million tokens, a 1,000× collapse in three years (a16z, LLMflation).
In that market, any provider that raises its consumer price hands its heaviest users to a competitor in a single billing cycle. Google made Gemini free at the low tier; Anthropic carved out usage-based metering on Claude Fable 5 at $10/$50 per million tokens (Wired, Anthropic support); Chinese open-weight providers like Qwen and DeepSeek compress list prices further. OpenAI sits inside a price ceiling set by every one of its rivals.
So the question becomes: if OpenAI cannot raise price, what can it do? The answer in late 2026 is the same answer every hyperscaler eventually reaches — stop renting the kitchen and build your own oven.
What is the five-layer AI stack, and which layers does OpenAI own?
Jensen Huang has described the AI industry as a five-layer cake: energy, chips, infrastructure, models, and applications. For most of its history OpenAI owned only one layer — the model — and rented everything else.
| Layer | Owner | OpenAI's position |
|---|---|---|
| Energy | Power plants | Zero control |
| Chips | Nvidia (~75% gross margin) | Rented from Nvidia |
| Infrastructure | Microsoft data centers | Rented capacity |
| Models | OpenAI, Anthropic, Google, Meta | Owned |
| Applications | Builders like you | Sold to via API + ChatGPT |
The structural problem is visible in that table. When every ChatGPT answer ran on Nvidia's chips inside Microsoft's data centers powered by somebody else's electricity, OpenAI captured the model layer margin and paid Nvidia's 75-cent-on-the-dollar gross margin for the privilege of running it. OpenAI's 2026 strategy is to own a second layer — chips — so that the inference cost it pays itself stays inside the OpenAI/Microsoft economic circle rather than leaking to Nvidia.
That is the whole thesis of the Jalapeño chip. It is not a training chip. It is an inference-only ASIC, co-designed with Broadcom and manufactured on TSMC's 3nm node, designed from scratch around the mathematical patterns of running large language models rather than training them (OpenAI, TechCrunch). If Broadcom's 50% claim holds at production scale, OpenAI's unit cost of serving any given model roughly halves — which is the same as saying its $14-billion-a-year inference bill gets cut in half once Jalapeño systems go live in volume in 2027.
What is GPT-5.6 Sol and why does it matter for the profit story?
GPT‑5.6 is OpenAI's newest model family, launched publicly on July 9, 2026 after a 30-day U.S. government safety review. It ships as three tiers rather than one model: Sol (the flagship for the hardest reasoning and agentic work), Terra (a balanced everyday model at about half Sol's price), and Luna (the fastest and cheapest) (OpenAI, OpenAI community).
| Model | Input (\(/Mtok) | Output (\)/Mtok) | Best for | |
|---|---|---|---|
| GPT‑5.6 Sol | $5 | $30 | Frontier reasoning, long-horizon agents |
| GPT‑5.6 Terra | $2.50 | $15 | Everyday work, matches GPT‑5.5 at 2× lower cost |
| GPT‑5.6 Luna | $1 | $6 | Fastest, cheapest tier |
Sol matters for the profitability story because it competes directly with Anthropic's Claude Fable 5 — the model most analysts credit with handing Anthropic the coding crown in early 2026. Anthropic's Fable 5 carries list pricing of $10/$50 per million tokens (Wired, Anthropic support). Sol's $5/$30 rate is half of Fable 5's input price and 60% lower on output, with OpenAI claiming it tops Fable 5 on coding and agent benchmarks. That is not a coincidence with the profitability story. The cheapest frontier model wins the API share war, and OpenAI is using Sol to claw back developers who had been migrating to Anthropic. Anthropic's own June 2026 ARR hit roughly $47 billion (Anthropic Series H press release, Axis Intelligence), so the gap is real and OpenAI has the incentive — and now the silicon — to close it.
How does Codex accelerate the profitability problem?
Cheaper models attract users, but they do not by themselves reduce total cost — because of what economists call Jevons' paradox: when the price of a resource falls, people consume more of it, often so much more that total spend rises rather than falls. OpenAI's Codex agent is the purest expression of that paradox in 2026.
Codex is no longer a coding assistant that autocompletes lines. It is an autonomous agent that writes code, builds tools, ships websites, runs multi-step jobs, and only returns when the task is done. In June 2026 OpenAI reported Codex had crossed 5 million weekly active users, up roughly 400% in 2026, with non-developers (marketers, lawyers, PMs, researchers) making up about 20% of usage and growing three times faster than the developer segment (OpenAI, "How agents are transforming work", 24 AI News). The OpenAI API now processes more than 15 billion tokens per minute.
The killer for unit economics is this: a single Codex task can consume what a hundred ordinary ChatGPT questions cost in compute. So the same user who was burning a $10/month hole through Pro at $20/month is now burning a hole that is 10 to 100 times larger when they hand Codex a real engineering job. That is why the cheap-flagship model and the cheap-inference chip are not three separate stories — they are one story. Sol creates the demand, Codex blows up the usage, and Jalapeño is the only thing that keeps the unit cost affordable enough for the math to work. Remove any one of those three pieces and OpenAI's profitability case collapses.
What is Jalapeño and why is it OpenAI's most important 2026 launch?
Jalapeño is OpenAI's first custom inference ASIC, unveiled with Broadcom on June 24, 2026. It went from blank-slate design to manufacturing tape-out in nine months, a cycle OpenAI and Broadcom claim is the fastest in high-performance semiconductors — they used OpenAI's own models to accelerate parts of the chip's design and optimization process (OpenAI, TechCrunch).
Three things make it the single most important move in OpenAI's profitability story:
- It is inference-only, not training. Training still runs on Nvidia; OpenAI did not try to compete with Nvidia on the hardest workload. It picked the workload where it has the biggest bill — decoding tokens for 900 million users — and built a chip purpose-fit for that math alone.
- The 50% cost claim is a unit-economics swing, not a marketing claim. Broadcom CEO Hock Tan told Reuters and Bloomberg that early testing shows Jalapeño delivers performance on par with Nvidia's Blackwell chips and Google's TPUs at roughly 50% lower cost per inference token (Bloomberg). If that holds at production scale, a $14 billion annual inference bill becomes ~$7 billion — which is the difference between losing $14 billion a year and losing a manageable amount. Note the caveat: this is an early-test, vendor-reported figure, not an independently audited production result. Treat it as a directional claim pending 2027 volume ramp.
- It lets OpenAI pass savings to API users without destroying its own margin. This is the piece builders actually feel. When your provider's cost base halves, your token price can halve without the provider going bankrupt — and that is the mechanism by which the AI token market keeps compressing 10× a year instead of stalling out.
Engineering samples of Jalapeño are already running GPT‑5.3‑Codex‑Spark workloads in OpenAI's lab at production-target frequency and power (OpenAI). Initial deployment is targeted for end of 2026, with volume production in 2027–2028.
Will cheaper AI actually make OpenAI profitable?
This is the part of the story most analysts undercall. Cheaper AI does not automatically mean lower bills — it means more usage. Jevons' paradox is the entire reason OpenAI is pouring money into Agents and Apps as fast as it builds chips: every 10× drop in inference cost opens up new use cases that were not commercially viable before (think video generation, autonomous browsing, long-horizon research agents), each of which consumes 10–100× more tokens per session than legacy chat.
So OpenAI's profitability in 2027 will depend on whether the rate of cost reduction outpaces the rate of usage growth. The ingredients that tilt that ratio toward profitability:
- Jalapeño volume ramp halving per-token cost (target 2027).
- Sol/Terra/Luna tiering that routes routine work to the cheapest acceptable model instead of always paying Sol prices.
- Codex monetization (it is now included on every ChatGPT plan and is the fastest-growing surface in the product).
- The $122 billion war chest that buys runway through the transition.
The reason I would not bet against the case is that the same dynamic is what turned cloud computing from a 2010s loss-leader into a 2020s profit engine — once hyperscalers owned the silicon underneath their services, per-unit cost fell faster than usage rose, and the business turned. OpenAI is now repeating that exact pattern one stack-layer down.
What does OpenAI's cost war mean for builders and small businesses?
For anyone building with AI in 2026 — a small business automating customer support, a solo founder shipping an agent product, a developer integrating LLMs into production — the takeaway is not to track OpenAI's Q3 earnings. It is to plan your own AI stack for the cost curve you are actually on.
- Model to the cheapest tier that clears the quality bar. Most production workloads do not need Sol or Fable 5 — they need Terra-class or Luna-class throughput. Our DeepSeek V4 Flash vs GPT-5.6 Luna cost comparison walks through the routing decision; the short version is that the cheapest tier is now good enough for ~80% of agent tasks, and the gap is closing fast.
- Build a routing layer, not a model loyalty. Every time a new model drops, the cheapest-acceptable tier shifts. We wrote a full guide on how to set this up: How to Set Up AI Model Routing With OmniRoute and Google Antigravity (2026 Guide). The pattern is identical to multi-cloud in the cloud era — you route each workload to the provider that is cheapest at that moment, not the one you started with.
- Assume your AI bill is 90% cheaper in 12 months, not stable. The a16z "LLMflation" curve shows ~10× decline per year at fixed quality. Build pricing and unit economics for that curve. If you are quoting a customer a per-seat AI price today, do not lock it in for three years — you will be charging them 10× the marginal cost in 2027 and your competitor will undercut you.
- Treat agents as a new cost category, not a feature. Codex-class autonomous tasks cost 10–100× more in compute than chat. If you are building agent workflows, you are now exposed to the same inference-cost war OpenAI is — and the same playbook applies: route cheap, monitor per-task spend, and let usage grow into the falling cost curve. Our How to Build a Self-Improving AI Agent Operating System in 2026 walks through the spend discipline.
- Watch the chip layer, because that sets your floor. The reason token prices keep falling is that the providers are now competing on silicon, not just models. Google has TPUs, Meta has MTIA, and now OpenAI has Jalapeño. Every time a hyperscaler vertical-integrates a layer of the stack, the savings eventually flow through to API list prices — even if the hyperscaler keeps most of them. For more on the broader dynamic, our piece on Why US Tech Stocks Are Falling in 2026: The AI Spending Paradox explains the macro view.
The strategic read for builders is simple: in a world where the marginal cost of intelligence is collapsing 10× a year, the winning products are the ones that route to the cheapest-acceptable model and design their unit economics around a falling curve. The providers will keep fighting the inference-cost war at the chip layer; you just get to consume the savings.
Related reading
FAQ
Q: Is OpenAI profitable in 2026? A: No. OpenAI's annualized revenue run rate reached roughly $25 billion in May 2026, but its projected 2026 operating loss is about $14 billion on a non-GAAP basis and roughly $33 billion on a GAAP basis, driven largely by inference cost (FutureSearch, Reuters).
Q: Why does OpenAI lose money on ChatGPT Pro subscribers? A: Because the marginal cost of serving a heavy user exceeds the $200/month fee. Sam Altman confirmed on January 5, 2025 that Pro was unprofitable because people used it more than expected (TechCrunch). Inference cost scales linearly with usage, and inference is the dominant cost for any AI subscription business.
Q: What is the Jalapeño chip and when will it ship? A: Jalapeño is OpenAI's first custom inference-only ASIC, co-designed with Broadcom on TSMC's 3nm node. It was unveiled June 24, 2026; Broadcom's CEO says early testing shows ~50% lower cost per inference token vs. current Nvidia GPUs. Engineering samples are running workloads in OpenAI's lab; initial deployment is targeted for end of 2026, with volume ramp in 2027–2028 (OpenAI, Bloomberg).
Q: How much has the price of an AI token fallen? A: Roughly 1,000× in three years. GPT‑3 at $60 per million tokens (November 2021) versus Llama 3.2 3B at $0.06 per million tokens (late 2024) for equivalent MMLU performance — a roughly 10× per year decline (a16z, LLMflation).
Q: How does GPT‑5.6 Sol compare to Anthropic's Claude Fable 5 on price? A: Sol lists at $5 input / $30 output per million tokens; Claude Fable 5 lists at $10 input / $50 output per million tokens. Sol is roughly half of Fable 5's input price and 60% lower on output, with OpenAI claiming it tops Fable 5 on coding and agent benchmarks (OpenAI, Wired).
Q: Will cheaper inference make AI providers profitable? A: Not automatically. Jevons' paradox means cheaper tokens drive more usage — OpenAI's Codex agent already runs tasks that cost 10–100× more compute than a chat question, and usage is growing 400% year-to-date. Profitability depends on whether the cost reduction from custom silicon and model tiering outpaces the usage explosion from agents.
Q: What should a small business do about falling AI token prices? A: Build a model-routing layer that sends each task to the cheapest acceptable model; do not lock in multi-year per-seat AI pricing; budget for a 10× annual decline in per-token cost; and treat autonomous agents as a distinct cost category that requires per-task spend monitoring.

Discussion
0 comments