DOLLARS PER TASK / SEPTEMBER 2026

Intelligence is
cheap. Thinking is not.

The model map ranks intelligence against the memory on your desk. This is the same frontier drawn against a meter: what one Artificial Analysis Intelligence Index task actually costs, measured rather than derived, with the option to re-price it at whatever route the market is offering today.

What the x-axis shows

Measured. Artificial Analysis’ own published cost for one Index task, taken as published. Nothing on this page derives it from prices and token counts.

Assumed input tokens / task

An assumption you are choosing, not a measurement. Artificial Analysis publishes a per-task output volume but no input volume on the same weighting, so the input side of the estimate has to be stated out loud. Only the cheapest-route mode uses it.

50%
Reads per cache write

A cached prefix has to be written before it can be read, and the write is amortized over the reads it earns. At 1× — written once, read once — every Claude model and GPT-5.6 Sol charge more for the cached path than the plain one.

47/97models carry a cost in this mode28 open-weight · the other 50 are listed with their reason, never dropped.

THE COST MAP

Intelligence per dollar.

Open weights Proprietary Pareto frontier Conditional re-rate
Intelligence Index ↑
0
12
24
36
48
60
$0.0010
$0.0030
$0.0100
$0.0300
$0.100
$0.300
$1.00
$3.00
$10.00
Measured USD per Intelligence Index task → log scale
PROPRIETARY · ON THE FRONTIER

Claude Opus 5.5

Anthropic · routed via Claude Opus 5.5

Measured / task
$5.98
Intelligence
57.6
Output tokens
119K
Of which thinking
70%
Full Index run
$8708.20

† CONDITIONAL RE-RATES

  • GPT-6 Astra. re-rates above a 272K-token prompt (output $75.00 / M). Base rate $10.00 in / $50.00 out per M.
  • DeepSeek V4 Pro 0813. off-peak UTC window at $1.98 / M output. Base rate $0.660 in / $1.98 out per M.
  • GPT-5.5. re-rates above a 272K-token prompt (output $45.00 / M). Base rate $5.00 in / $30.00 out per M.
  • GPT-5.4. re-rates above a 272K-token prompt (output $22.50 / M). Base rate $2.50 in / $15.00 out per M.
  • GPT-5.6 Sol. re-rates above a 272K-token prompt (output $15.00 / M). Base rate $2.00 in / $10.00 out per M.
  • Gemini 2.5 Pro. re-rates above a 200K-token prompt (output $15.00 / M). Base rate $1.25 in / $10.00 out per M.
  • Grok 4.6. re-rates above a 200K-token prompt (output $12.00 / M). Base rate $2.00 in / $6.00 out per M.

WHAT YOU PAY TO THINK

Most of the bill never reaches you.

MODELS OVER 60% THINKING29/48reasoning share of output tokens per task
REASONING PRICED SEPARATELY9and all 9 are set to the model’s own output rate anyway

Reasoning tokens are output tokens. 9 models in the catalogue — every Gemini Flash and Flash-Lite generation, plus Gemini 2.5 Pro — publish a distinct reasoning rate, and all 9 of them set it equal to the model’s own output rate. So thinking bills at the output rate everywhere, and the light bar below is the only part of the answer you ever read. The 48 models with published per-task volumes are listed here; the rest have no token breakdown to split.

ModelAnswer vs. thinking · per taskThinkingOutput $ on thinking
gpt-oss 20BOpenAI · open weights89%$0.0029at $0.180 / M out
Nemotron 3 NanoNVIDIA · open weights89%$0.0048at $0.200 / M out
Gemini 3.1 Flash-LiteGoogle · reasoning priced separately82%$0.0148at $1.50 / M out
GLM-5.2Z.ai · open weights79%$0.223at $4.40 / M out
Claude Sonnet 5Anthropic75%$0.884at $10.00 / M out
DeepSeek V4 FlashDeepSeek · open weights74%$0.0114at $0.280 / M out
Claude Opus 4.8Anthropic73%$1.30at $25.00 / M out
DeepSeek V4 Flash 0731DeepSeek · open weights73%$0.0600at $1.32 / M out
Qwen3.8 27BAlibaba · open weights71%$0.143at $3.00 / M out
Claude Opus 5.5Anthropic70%$1.68at $20.00 / M out
K2 Horizon 375B A23BMBZUAI · open weights69%—no output rate published
DeepSeek V4 Pro 0813DeepSeek · open weights69%$0.151at $3.96 / M out
Muse GlimmerMeta · open weights69%$0.0143at $1.50 / M out
GLM-5.3Z.ai · open weights69%$0.214at $4.40 / M out
Qwen3.8 2.4T A95BAlibaba · open weights68%$0.280at $6.00 / M out
GLM-5.3 FlashZ.ai · open weights68%$0.0233at $0.500 / M out
InklingThinking Machines · open weights68%$0.0948at $4.05 / M out
Claude Fable 5Anthropic67%$2.25at $50.00 / M out
Kimi K3Moonshot AI · open weights67%$0.487at $15.00 / M out
Qwen3.8 Flash NextAlibaba · open weights67%$0.0337at $0.470 / M out
Gemini 3.5 FlashGoogle · reasoning priced separately66%$0.327at $9.00 / M out
DeepSeek V4 Flash VisionDeepSeek66%$0.0606at $1.32 / M out
Granite 4.2 8BIBM · open weights65%$0.0055at $0.250 / M out
GPT-5.5OpenAI65%$0.463at $30.00 / M out
Gemini 2.5 ProGoogle · reasoning priced separately64%$0.0677at $10.00 / M out
Kimi K2.7 CodeMoonshot AI · open weights64%$0.0767at $4.00 / M out
GPT-6 AstraOpenAI61%$0.835at $50.00 / M out
Claude Fable 5.1Anthropic60%$2.36at $50.00 / M out
Granite 4.2 3BIBM · open weights60%$0.0013at $0.120 / M out
Gemini 3.5 Flash-LiteGoogle · reasoning priced separately59%$0.0260at $2.50 / M out
Claude Opus 5Anthropic59%$1.07at $25.00 / M out
GPT-5.6 SolOpenAI59%$0.346at $20.00 / M out
Nemotron 3 SuperNVIDIA · open weights59%$0.0604at $0.900 / M out
MiMo-V2.6-ProXiaomi · open weights58%$0.0326at $0.870 / M out
Qwen3.6 27BAlibaba · open weights58%$0.0678at $3.60 / M out
gpt-oss 120BOpenAI · open weights56%$0.0090at $0.595 / M out
Grok 4.6SpaceXAI52%$0.113at $6.00 / M out
Gemini 3.6 FlashGoogle · reasoning priced separately48%$0.0752at $3.75 / M out
Muse Spark 1.2Meta47%$0.0963at $4.25 / M out
Qwen3.5 397B-A17BAlibaba · open weights47%$0.0315at $3.60 / M out
MiniMax M2.7MiniMax · open weights46%$0.0117at $1.20 / M out
Nemotron 3.5 LightningNVIDIA · open weights46%$0.0037at $0.220 / M out
MiniMax M3MiniMax · open weights45%$0.0262at $1.20 / M out
Qwen3.5 122B-A10BAlibaba · open weights42%$0.0233at $3.20 / M out
Claude Haiku 4.5Anthropic41%$0.0383at $5.00 / M out
Gemini 3.7 FlashGoogle · reasoning priced separately36%$0.0796at $3.75 / M out
Qwen3.6 35B-A3BAlibaba · open weights36%$0.0281at $2.25 / M out
Qwen3 Coder NextAlibaba · open weights0%$0.0000at $1.20 / M out

THE CACHE-WRITE PREMIUM

Caching is not automatically a discount.

COSTLIER THAN NOT CACHING15/15at a 50% hit rate with 1 read per write

Eleven models charge a premium to write the cache rather than a discount: every Claude at exactly 1.25× the input rate, GPT-5.6 Sol at 1.25×, and Claude 3 Haiku at 1.20×. Below the break-even reuse count, turning caching on raises the input bill. Move the two controls above and watch the sign flip. Separately, only 61 of the 70 priced models publish a cache-read rate at all; for the rest the slider correctly does nothing.

ModelList inputCache writeBreak-even readsEffective inputΔ
Qwen3.8 Flash NextAlibaba$0.150$0.2001.33× input1.5×$0.183+22%
Claude Fable 5Anthropic$10.00$12.501.25× input1.4×$11.75+18%
Claude Haiku 4.5Anthropic$1.00$1.251.25× input1.4×$1.18+18%
Claude Opus 4.5Anthropic$5.00$6.251.25× input1.4×$5.88+18%
Claude Opus 4.6Anthropic$5.00$6.251.25× input1.4×$5.88+18%
Claude Opus 4.7Anthropic$5.00$6.251.25× input1.4×$5.88+18%
Claude Opus 4.8Anthropic$5.00$6.251.25× input1.4×$5.88+18%
Claude Opus 5Anthropic$5.00$6.251.25× input1.4×$5.88+18%
Claude Sonnet 4.6Anthropic$3.00$3.751.25× input1.4×$3.52+18%
Claude Sonnet 5Anthropic$2.00$2.501.25× input1.4×$2.35+18%
GPT-5.6 SolOpenAI$2.00$2.501.25× input1.4×$2.35+18%
GPT-6 AstraOpenAI$10.00$12.501.25× input1.4×$11.75+18%
Claude 3 HaikuAnthropic$0.250$0.3001.20× input1.4×$0.290+16%
Claude Opus 5.5Anthropic$4.00$5.001.25× input1.3×$4.60+15%
Claude Fable 5.1Anthropic$10.00$12.501.25× input1.3×$11.38+14%

No cache pricing published. GLM-4.7 Flash, GPT-3.5 Turbo, GPT-4, GPT-4 Turbo, gpt-oss 20B, Nemotron 3 Super, Qwen3.5 122B-A10B, Qwen3.5 27B, Qwen3.5 9B. Their input bills at the list rate no matter where the slider sits — an absence of data, not a rate of zero.

SHOPPING AROUND

The headline price is somebody else’s.

OpenRouter’s headline number is the cheapest route it can find, which is often a third-party host rather than the model’s own vendor. That is a real saving you can take, but it is not the list price, and the cheapest-route mode above is quietly built on it. These 14 models diverge by more than 10% from the price Artificial Analysis used when it measured them.

ModelAA’s output rateCheapest routeHostShopping around saves
DeepSeek V4 FlashDeepSeek$0.280 / M$0.098 / MDeepSeek V4 Flash 0423−65%
Qwen3.6 35B-A3BAlibaba$2.25 / M$1.00 / MQwen3.6 35B A3B−56%
GLM-5.2Z.ai$4.40 / M$2.04 / MGLM 5.2−54%
GLM-5.3Z.ai$4.40 / M$2.05 / MGLM 5.3−53%
DeepSeek V4 Flash 0731DeepSeek$1.32 / M$0.640 / MDeepSeek V4 Flash 0731−52%
Nemotron 3 SuperNVIDIA$0.900 / M$0.450 / MNemotron 3 Super−50%
gpt-oss 20BOpenAI$0.180 / M$0.090 / Mgpt-oss-20b−50%
DeepSeek V4 Pro 0813DeepSeek$3.96 / M$1.98 / MDeepSeek V4 Pro 0813−50%
GPT-5.6 SolOpenAI$20.00 / M$10.00 / MGPT-5.6 Sol−50%
Qwen3.5 122B-A10BAlibaba$3.20 / M$2.08 / MQwen3.5-122B-A10B−35%
Qwen3 Coder NextAlibaba$1.20 / M$0.800 / MQwen3 Coder Next−33%
Qwen3.6 27BAlibaba$3.60 / M$2.70 / MQwen3.6 27B−25%
Muse GlimmerMeta$1.50 / M$1.20 / MMuse Glimmer 30B−20%
Kimi K2.7 CodeMoonshot AI$4.00 / M$3.30 / MKimi K2.7 Code−18%

COST PER TOKEN

The cloud price floor keeps falling.

CHEAPEST INPUT$0.14 / MDeepSeek V4 Flash · $0.0028 on cache hits
STEEPEST CUT−80%GPT-5.6 Luna · $6 → $1.20 per M output

First-party API list prices per million tokens, last updated Sep 5, 2026. OpenAI prices were refreshed on that date; other vendors retain the September 1 snapshot. On July 30 OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%; the next day DeepSeek re-released V4 Flash — the model whose April weights sit on the local graph on the model map — with agent-focused post-training and the same $0.14-per-million input price.

ModelIntelligence$ / M tokens · log scaleInputOutput
DeepSeek V4 FlashDeepSeek · open weights24.6$0.14$0.28
DeepSeek V4 ProDeepSeek—$0.435$0.87
GPT-5.6 LunaOpenAI—$0.20$1$1.20$6−80%
Claude Haiku 4.5Anthropic17.6$1$5
Gemini 3.6 FlashGoogle34.3$1.50$7.50
GPT-5.6 TerraOpenAI—$2$2.50$12$15−20%
Claude Sonnet 5Anthropic38.4$3$15
GPT-5.6 SolOpenAI47.1$4$20
Claude Opus 5.5Anthropic57.6$4$20
Claude Opus 5Anthropic50.7$5$25
GPT-6 AstraOpenAI52.8$10$50
Claude Fable 5Anthropic49.7$10$50
Claude Fable 5.1Anthropic53.4$10$50
  • DeepSeek V4 Flash. Off-peak list price; every item bills at 2× during Beijing peak hours.
  • DeepSeek V4 Pro. The V4 Pro API has 1M context; the August 13 weights are also downloadable.
  • GPT-5.6 Luna. OpenAI's high-volume tier for classification, routing, and summarization.
  • Gemini 3.6 Flash. Cut output from $9 to $7.50 versus Gemini 3.5 Flash; batch runs at half price.
  • GPT-5.6 Terra. The mid production tier, pitched at GPT-5.5-class quality.
  • Claude Sonnet 5. Standard pricing after the introductory $2 / $10 offer ended August 31, 2026.
  • GPT-5.6 Sol. Standard rates. Above 272K input tokens, the full request bills at $8 input / $30 output per million tokens.
  • Claude Opus 5.5. Launched Sep 22, 2026 below the Opus line's standing $5 / $25 price. Fast mode runs at 2× these rates.
  • Claude Opus 5. Launched Jul 24, 2026 at the same list price as Opus 4.8. Moved to Anthropic's legacy-model list when Opus 5.5 shipped.
  • GPT-6 Astra. Standard rates; cache writes cost $12.50/M. Above 272K input tokens, the full request uses 2× input and cache rates and 1.5× output rates. Access is rolling out.
  • Claude Fable 5.1. Launched Sep 1, 2026 at the same list price as Fable 5. Cache hits bill at 2.5% of the input rate instead of the usual 10%.
  • Cache reads. All four vendors discount cached input by roughly 90%; DeepSeek goes furthest, reading hits at $0.0028 per million tokens.

A downloaded GGUF has no meter: once weights fit your Mac, the marginal cost of a token is electricity. These list prices are what local inference competes with — and July moved the floor closer to zero.

DEEPSEEK API CHANGELOG

V4 Flash, re-post-trained.

What changed on the API side of the open-weight model this site tracks locally, from DeepSeek’s update log ↗.

  1. V4 Flash official release, re-post-trained for agentsThe official DeepSeek-V4-Flash API enters public beta: the same architecture as the preview, re-post-trained. DeepSeek reports Terminal Bench 2.1 82.7, Cybergym 76.7, Toolathlon 70.3, and DeepSWE 54.4, with native Responses API support adapted for Codex. The update is API-only; the downloadable April weights are unchanged.
  2. Legacy model names retiredThe deepseek-chat and deepseek-reasoner API names were discontinued after a three-month transition in which they mapped to V4 Flash's non-thinking and thinking modes.
  3. DeepSeek-V4 launchdeepseek-v4-pro and deepseek-v4-flash arrive with OpenAI- and Anthropic-compatible interfaces and 1M context. The V4 Flash weights ship openly under MIT — the checkpoint this site tracks locally.

READ THE FINE PRINT

Four claims, four different footings.

01

Cost per task is measured, not derived

The default x-axis is Artificial Analysis’ own published cost for one Intelligence Index task, applied as published. It already carries each vendor’s cache discounts. This page does not reconstruct it from prices and token counts — an earlier attempt to do that missed the published figure by 0.5× to 3×, because the only input-token figure available sits on a different weighting than the per-task output figure.

02

Thinking bills at the output rate

Reasoning tokens are output tokens. Nine models publish a separate reasoning rate and every one of them sets it equal to their own output rate, so no model here bills thinking cheaply. With reasoning running 60–90% of output on the heaviest models, the token volume — not the posted price — is what decides the bill.

03

Cheapest route ≠ list price

OpenRouter publishes the cheapest available route, frequently a third-party host. 14 models diverge from the vendor price by more than 10%, the widest being DeepSeek V4 Flash at roughly 8×. The re-priced mode is an estimate built on that route and on a per-task input volume you choose; it is labelled as an estimate everywhere it appears.

04

Caching can cost more

Anthropic and OpenAI charge 1.20–1.25× the input rate to write the cache. Written once and read once, that is a surcharge, not a saving; it only turns into a discount past each model’s break-even reuse count. Six models also carry conditional re-rates — long-context thresholds or DeepSeek’s UTC off-peak windows — and the chart uses only their base rate, so they are flagged with †.

Per-task token volumes and cost figures © Artificial Analysis, used with attribution; API list prices from OpenRouter’s public catalogue. Pricing snapshot generated Sep 22, 2026; first-party vendor board updated Sep 5, 2026. A missing number here always means it could not be read from a source, never that it is zero.