AI Inference Economics Tracker

GPU cost (input) vs Token price (output) vs Usage volume (demand) | Historical trends and gap analysis

Data refreshed 14 Sep 2026. Where a source's newest published reading is older than today, the note under each figure says so.

1. GPU Rental Cost Tracker (Input Side)

H100 SXM Spot
$2.78
-33.8% since Jan
Ornn OCPI, Sep 13
H200 Spot
$4.38
-17.4% since Jan
Ornn OCPI, Sep 13
B200 Spot
$7.06
-25.7% since Apr
Ornn OCPI, Sep 13
A100 SXM Spot
$1.03
-31.3% since Jan
Ornn OCPI, Sep 13

Sources: Ornn OCPI (data.ornn.com, free API, daily, volume-weighted winsorized average of transacted rentals). H100 SXM is the primary benchmark. B200 data begins Apr 26 (first Blackwell listings). Sep 26 point reflects the Sep 13 settle.

What this tracks: The cost of the INPUT to inference. Falling GPU rental prices mean each dollar of inference revenue costs less compute to serve. This is the supply-side deflation curve. The H100 has fallen from $4.80 (Jul 25) to $2.78 (Sep 13, 2026), a 42% decline in 14 months, despite demand doubling. Supply is catching up.

2. Token Pricing Tracker (Output Side)

Premium Tier (Claude Opus 5)
$11.00
Blended $/M tokens, $5 in / $25 out
Mid Tier (Sonnet 5)
$4.40
Blended $/M tokens, $2 in / $10 out
Commodity Tier (DeepSeek V4 Flash)
$0.06
Blended $/M tokens, $0.035 in / $0.106 out

Sources: claude.com/pricing (Opus 5, Sonnet 5, checked Sep 14), openrouter.ai (DeepSeek V4 Flash Latest, checked Sep 14). Blended = 70:30 input:output ratio. Historical points from Vercel AI Gateway Index avg cost/token (monthly, free blog). Note: the commodity-tier figure was revised up from a prior $0.02 reading - that was a promotional/older-variant price, not the current DeepSeek V4 Flash Latest listing.

Vercel AI Gateway: average cost per token across all models/providers routed through the gateway (~200K teams). MoM change shown. May spike = new frontier model launches at premium pricing. Jul 2026 is the latest full monthly index Vercel has published (checked Sep 14; no Aug index is out yet).

3. The Gap: GPU Cost vs Token Revenue (Inference Margin Proxy)

GPU Cost Trend (H100)
-33.8%
Since Jan 2026, through Sep 13 (Ornn, daily)
Avg Token Price Trend
-13.6%
Jul 2026 MoM - Vercel's latest published reading, checked Sep 14, none newer exists yet

Gap = GPU cost deflation rate minus token price deflation rate. Positive gap = inference margins expanding (GPU costs falling faster than token prices). Negative gap = margins compressing. GPU line runs through Sep 13 (Ornn updates daily). Token price and margin-gap lines stop at Jul 26 and are shown with a visible break - Vercel has not published an Aug or Sep aggregate reading (checked Sep 14 via their blog, index landing page, and third-party trackers).

What the gap tells you: Since Jan 2026, GPU rental costs have fallen ~34% (through Sep 13) while average token prices had fallen ~14% as of the last reading Vercel published (July). The gap was POSITIVE as of July, meaning inference margins were expanding: providers keep more of each dollar of token revenue because their GPU input cost deflates faster. GPU costs have kept falling since then, so if token pricing has held roughly flat the gap has likely widened further - but that is inference, not a confirmed reading, since Vercel's own token-price aggregate has not been updated past July.

3b. Live Proxy: Volume & Blended Price, Computed Directly from OpenRouter (Free, Updates Weekly)

OpenRouter Weekly Volume
126.8T
+1,875% since Jan 5
Week of Sep 7, 2026 (last complete week)
Blended Price, Top Models
$0.42/M
-78.5% since Jan 5
Week of Sep 7, 2026, 70:30 I/O blend

Computed by us directly from two free OpenRouter endpoints, no login or API key: openrouter.ai/api/frontend/v1/rankings/model-rankings-chart (real weekly token volume per top model plus an "Others" residual, since Sep 2025) joined against openrouter.ai/api/v1/models (live per-model $/token pricing). Blended price = 70:30 input:output on each priced model's own rate, volume-weighted across the named top models only (the "Others" bucket, roughly 35-50% of weekly volume depending on the week, has no per-model pricing available and is excluded from the price calculation but included in the volume total). Checked and computed Sep 14, 2026.

Read this one carefully: the price decline shown here (-78.5% since Jan) is far steeper than Vercel's reported market-wide decline (-13.6% MoM as of July) because it captures a genuine compositional shift, not just per-model price cuts. In January the top-volume models on OpenRouter were mostly premium (Claude, GPT-class). By September the top-volume slots are almost entirely ultra-cheap commodity models (DeepSeek, GLM, Hy-series). That mix shift alone drags this number down hard - it is directionally real (cheap models are eating volume share, and volume itself is up nearly 19x) but should not be read as "the same basket of models got 78% cheaper." Use the volume trend with confidence; use the price trend as a directional signal, not a precise index.

4. Token Usage Volume Tracker (Demand Side)

OpenRouter Weekly Volume
~127T
tokens/week (~550T/mo run-rate), week of Sep 7, computed from OpenRouter's own data - see 3b
Vercel Gateway Volume MoM
+37%
Jul 2026 vs Jun
Vercel Gateway Spend MoM
+37%
Volume doubled since May

Vercel AI Gateway monthly token volume growth (rebased to 100 = Oct 2025). Gateway serves 200K+ teams with tens of trillions of tokens. Token growth has been accelerating: volume doubled between May and Jul 2026.

5. OpenRouter Model Rankings (with Historical Shifts)

#ModelProviderOriginDaily Tokens (Sep 13)30-Day TotalStatus
1GPT-5.6 LunaOpenAIUS3.9T45.1TNew #1, up from #3 in Sep 12 reading
2DeepSeek V4 FlashDeepSeekChina1.4T50.7THighest 30-day total overall
3GLM 5.3 FlashZ.ai / ZhipuChina1.4T30.4TSteady top-3
4Hy4 PreviewTencentChina1.4T34.6TNew model, entered top 5
5DeepSeek V4.1 FlashDeepSeekChina1.3T4.9TNewer variant, fast ramp
6MiMo V2.5XiaomiChina985.4B30.2TStable
7DeepSeek V4 Flash (0423)DeepSeekChina519.3B21.5TOlder variant, still active
8Nemotron 3 Ultra 550BNvidiaUS (open)472.4B18.3TLeading open-weight, US
9Hy3TencentChina458.1B24.8TPrior-gen Tencent model
10GLM 5.3Z.ai / ZhipuChina276.9B7.4TNon-flash variant

Source: Tokenmaxxing OpenRouter rankings mirror (tokenmaxxing.com/openrouter-rankings, checked Sep 14, data dated Sep 13). Replaces the Sep 12 snapshot below - note the leader changed (GPT-5.6 Luna overtook DeepSeek V4 Flash 0731 for the daily-volume top spot) and a new Tencent model (Hy4 Preview) entered the top 5. Rankings measure adoption (tokens routed), not quality. WoW % was not available from this reading; 30-day total shown instead.

Chinese-origin model volume share on OpenRouter: from under 2% (Sep 2025) to ~48% (Aug 2026, last confirmed reading as of Sep 14 check). The fastest market share shift in the platform's history, driven by DeepSeek, GLM, Kimi, MiMo, and Tencent Hy3. Volume share, not spend share (Chinese models capture under 10% of spend due to near-zero pricing). The Sep 13 rankings snapshot in section 5 shows a US model (GPT-5.6 Luna) newly on top by daily volume, which may signal this share has since pulled back - not yet reflected here as a confirmed monthly figure.

6. Structural Shift Trackers

Agentic Token Share
58.9%
+27.3pp in 6 months
% of tokens in tool-call requests, Vercel's Apr 26 reading - not updated since
Open-Weight Volume Share
62%
+26pp since Jul
Vol % on Vercel Gateway, spot reading 22 Aug
Anthropic Price Premium
4.4x
vs avg gateway token price
65% of spend on 30% of vol

Vercel AI Gateway monthly data (free blog), checked Sep 14. Agentic share = % of all tokens in requests that end with a tool call - Vercel's own reporting has held this at 58.9% since its Apr 26 report, so the line is flat from Apr onward rather than showing invented growth. Open-weight jumped to 62% volume / under 9% spend by a 22 Aug spot reading (up from 36%/8.6% in the Jul monthly index) - shown as a single dated point since it isn't a full monthly reading. Open-weight = DeepSeek, MiniMax, Moonshot, Z.ai. Feb 6, 2026 was flagged by OpenRouter as "potentially the last day when humans consumed more tokens than agents."

7. Data Source Registry

SourceURLAccessFrequencyWhat it tracksWhat it misses
Ornn OCPIdata.ornn.comFree API (3mo)DailyTransaction-cleared GPU rental pricesLong-term contracts, hyperscaler internal, utilization
Ornn OTPIdata.ornn.comFree API (1mo)DailyRealized token cost by lab, volume-weightedSelf-hosted inference, batch pricing
OpenRouter Rankingsopenrouter.ai/rankingsFree API, CC BY 4.0DailyToken volume by model, 100T+/mo, 427 modelsEnterprise direct API, B2C apps, China domestic
OpenRouter Dataopenrouter.ai/dataFreePeriodicDeep dives: model shifts, pricing, agentic trendsSame as rankings
Vercel AI Gatewayvercel.com/blogFree blogMonthlySpend share, agentic %, open-weight %, cost/tokenEnterprise self-hosted (~90% of inference), China
AIMultiple GPU Indexaimultiple.com/gpu-indexFree webMonthly69 providers, 17 GPUs, 26mo historyContract pricing, intra-month moves
Tokenmaxxingtokenmaxxing.comFree webDailyOpenRouter + company spend estimatesPrivate companies