Since October 7, OpenAI and Anthropic have charged the same standard list prices at three matched tiers. The difference has moved to cost per task. At equal rates a bill depends on how much text a model reads and writes to finish a job, and Anthropic's models write far more. Artificial Analysis, the one independent benchmark found that prices a task for all three pairs, runs them at maximum effort. On it, Anthropic's cheapest model costs about three times as much per task as OpenAI's. That pair is the cleanest test. At the other two tiers Anthropic's models cost 2.3 and 7.6 times as much, figures that may include charges from other models. Anthropic scores higher at two tiers and ties at the third.
Prices matched at three tiers by October
Both firms price by the million tokens, the units of text a model reads (input) and writes (output). OpenAI launched GPT-6 Astra on 3 September, and it is now listed at $10 and $50, the price of Anthropic's Fable models since June.123 The middle tier matched on 22 September, when OpenAI priced GPT-6 Sol at $2 and $10, Anthropic's Sonnet price since 30 June. GPT-6.1 Sol kept that price on 29 September.13 The bottom tier matched on 7 October, when Anthropic launched Haiku 5.5 at GPT-6 Luna's price.14 Anthropic also sells Opus 5.5 at $4 and $20, a price with no OpenAI match.4
List prices match at three tiers
| Tier | OpenAI | Anthropic | Input, $ per million tokens | Output, $ per million tokens |
|---|---|---|---|---|
| Top | GPT-6 Astra | Claude Fable 5.1 | 10 | 50 |
| Middle | GPT-6.1 Sol | Claude Sonnet 5.5 | 2 | 10 |
| Bottom | GPT-6 Luna | Claude Haiku 5.5 | 0.10 | 0.50 |
OpenAI cut prices on two models to get there. Sol fell from $5 and $30 in July to $2 and $10. Luna's prices fell 80 percent on 30 July and by half or more on 22 September.35 OpenAI says cheaper caching and inference let it pass savings on, and its launch post for Sol argues on cost per task against Anthropic's Fable 5.1.6 Anthropic launched Haiku 5.5 at a tenth of Haiku 4.5's price, and on 10 August it canceled a scheduled rise in Sonnet's.1 Neither firm says it was matching the other.
The match covers standard rates for prompts under 100,000 tokens. Two rates still differ. At the top tier, Artificial Analysis lists cached input at $0.25 per million tokens on Fable 5.1 and $1.00 on Astra.7 OpenAI charges higher long-prompt rates on all three of its models, from 272,000 tokens on Sol. Anthropic holds Fable's and Sonnet's standard rates to a million tokens, and Haiku 5.5's rates rise fivefold above 100,000.24
Anthropic costs more per task at every tier
Artificial Analysis, an independent testing firm, runs one set of evaluations on both firms' models and combines them into a score, its Intelligence Index. It reports the tokens and dollars each model spends on what it counts as an average task in that set, using measured cache use.78 A task here is a benchmark unit and differs from a buyer's workload. The pages give no date for when the data were collected. On 7 October Anthropic cut Sonnet 5.5's cached-input price to $0.10, and the record does not show whether the middle-tier figures include that cut.1
At equal list prices Anthropic scores the same or higher and costs more per task
| Tier | Model | Intelligence Index | Output tokens per task, thousands | Of which reasoning, thousands | Cost per task, $ |
|---|---|---|---|---|---|
| Top | GPT-6 Astra | 53 | 27 | 17 | 3.26 |
| Top | Claude Fable 5.1 | 53 | 78 | 47 | 7.63 |
| Middle | GPT-6.1 Sol | 52 | 38 | 25 | 0.72 |
| Middle | Claude Sonnet 5.5 | 56 | 197 | 147 | 5.46 |
| Bottom | GPT-6 Luna | 38 | 50 | 39 | 0.07 |
| Bottom | Claude Haiku 5.5 | 43 | 162 | 129 | 0.21 |
Source: Artificial Analysis.7
Anthropic's model costs 2.3, 7.6 and about 3.0 times as much per task at the top, middle and bottom tiers. The ratios hold when measured by the cost of running the whole index, at 2.5, 6.7 and 2.7.7 Per 1,000 tasks, the gap is $4,370 at the top, $4,740 in the middle and about $140 at the bottom.
The extra cost buys points at two tiers, at very different prices. At the bottom tier each extra index point costs about 3 cents a task. In the middle each costs $1.19. At the top, 2.9 times the tokens buys no points. Above the bottom tier, list price does not predict score. Sonnet 5.5 scores highest of the six. Sol scores within a point of Astra and Fable 5.1 at a fifth and a tenth of their cost per task. The pages read give no error margins, so gaps of four or five points show direction more reliably than size.7
The token gap also reverses the speed ranking. Anthropic's models write more tokens a second and still take longer to finish. Fable 5.1 takes about 19 minutes to generate a task's output, against Astra's 10. Sonnet 5.5 takes 23 to 26 minutes, on the two speeds its pages have shown, to Sol's 11. Haiku 5.5 takes about 11 minutes to Luna's 7. These times count generation alone and exclude input processing and network delay.7
Output tokens explain part of the gap
Anthropic's models write 2.9, 5.2 and 3.2 times as many output tokens per task as OpenAI's, top to bottom. Most of the excess is reasoning, the text a model writes to itself before it answers, at 59, 77 and 80 percent of each gap.7 At list prices, output tokens account for 58, 34 and 40 percent of each cost gap. The rest is input and cached text, which the benchmark does not itemize. The largest gap, in the middle tier, is therefore two-thirds unexplained by the figures published.
Three explanations for the extra tokens fit the record, and it cannot separate them. The "Max" setting is set per model, and nothing shows the two firms' settings are equivalent. Some requests may be rerun on other models, as the next section explains, though that cannot apply at the bottom tier. And tokens are counted differently. Anthropic says its tokenizer for Claude 4.7 and later produces about 30 percent more tokens for the same text than its earlier one.4 That could account for at most 1.3 times of output gaps of 2.9 to 5.2 times, if the earlier tokenizer matched OpenAI's, which no source tests. A buyer pays for every token as counted, so a counting difference still reaches the bill.
One benchmark explains the difference
Every figure comes from runs at the setting Artificial Analysis labels "Max." The publisher lists no lower-effort runs for the matched pairs.79 Effort moves results. Opus 5.5 scores 58 at "Max" and 54 at "High."9 A buyer who sets lower effort would likely pay less per task and might score lower, by amounts the benchmark does not report. No second independent benchmark of cost per task for these pairs was found. OpenAI's own test, published as an interested party, points the same way. It reports Sol finishing its AutomationBench tasks at 9 percent of Opus 5's cost, though Opus 5 lists at 2.5 times Sol's price.64
Artificial Analysis labels the Fable 5.1 and Sonnet 5.5 runs "Default Fallback" and does not explain the label.78 Anthropic offers a fallback mode in which a request its safety classifiers decline is rerun on another Anthropic model, billed at that model's rates. Haiku 5.5 has no fallback, and its run carries no label.10 If the label marks that mode, part of the top and middle bills belongs to other models. OpenAI reports fallbacks on about 40 percent of tasks in its own test of Fable 5.1, a different setup with Opus 5 as the fallback.6 The bottom tier rules out fallback as the whole story. Haiku 5.5 writes 3.2 times Luna's output tokens with no fallback at all, so Anthropic's heavier token use holds where the doubt is absent.
Buyer data predate the match
A Wall Street Journal headline on 7 October said OpenAI is gaining ground on Anthropic as the AI price war heats up.11 The public buyer data end before the last tier matched. Ramp is a corporate card company with more than 70,000 American business customers. In August it found 43.8 percent of them paying for Anthropic and 39.8 percent paying for OpenAI. Anthropic's share rose faster that month.1213 Ramp's lead economist told TechCrunch in August that OpenAI was growing faster in the third quarter to date.13 That growth came while list prices still differed, so these data cannot show whether buyers now choose on cost per task.
Switching has increased with a record 8 percent of Ramp's firms changed AI provider each month, on a three-month average, to September. Ramp calls this a possible early sign that AI models are becoming commodities.14 If they were, price alone would decide between them. Here list prices are equal, and cost per task differs between 2.3 and 7.6 times. Ramp ties an earlier shift in switching to Anthropic's model launches of May 2025. It ties a second, which it dates to June 2026, to OpenAI's price cuts and launches.14
At equal list prices the model sets the bill
At the bottom tier, the cleanest test found, Anthropic's model costs about three times as much per task as OpenAI's and scores five points higher. At the top and middle tiers the multiples are 2.3 and 7.6, under a label that may fold in other models' charges. In the middle tier Sonnet 5.5's four extra points cost $4,740 more per 1,000 tasks than Sol's.