GPU cost (input) vs Token price (output) vs Usage volume (demand) | Historical trends and gap analysis
Data refreshed 14 Sep 2026. Where a source's newest published reading is older than today, the note under each figure says so.
Sources: Ornn OCPI (data.ornn.com, free API, daily, volume-weighted winsorized average of transacted rentals). H100 SXM is the primary benchmark. B200 data begins Apr 26 (first Blackwell listings). Sep 26 point reflects the Sep 13 settle.
Sources: claude.com/pricing (Opus 5, Sonnet 5, checked Sep 14), openrouter.ai (DeepSeek V4 Flash Latest, checked Sep 14). Blended = 70:30 input:output ratio. Historical points from Vercel AI Gateway Index avg cost/token (monthly, free blog). Note: the commodity-tier figure was revised up from a prior $0.02 reading - that was a promotional/older-variant price, not the current DeepSeek V4 Flash Latest listing.
Vercel AI Gateway: average cost per token across all models/providers routed through the gateway (~200K teams). MoM change shown. May spike = new frontier model launches at premium pricing. Jul 2026 is the latest full monthly index Vercel has published (checked Sep 14; no Aug index is out yet).
Gap = GPU cost deflation rate minus token price deflation rate. Positive gap = inference margins expanding (GPU costs falling faster than token prices). Negative gap = margins compressing. GPU line runs through Sep 13 (Ornn updates daily). Token price and margin-gap lines stop at Jul 26 and are shown with a visible break - Vercel has not published an Aug or Sep aggregate reading (checked Sep 14 via their blog, index landing page, and third-party trackers).
Computed by us directly from two free OpenRouter endpoints, no login or API key: openrouter.ai/api/frontend/v1/rankings/model-rankings-chart (real weekly token volume per top model plus an "Others" residual, since Sep 2025) joined against openrouter.ai/api/v1/models (live per-model $/token pricing). Blended price = 70:30 input:output on each priced model's own rate, volume-weighted across the named top models only (the "Others" bucket, roughly 35-50% of weekly volume depending on the week, has no per-model pricing available and is excluded from the price calculation but included in the volume total). Checked and computed Sep 14, 2026.
Vercel AI Gateway monthly token volume growth (rebased to 100 = Oct 2025). Gateway serves 200K+ teams with tens of trillions of tokens. Token growth has been accelerating: volume doubled between May and Jul 2026.
| # | Model | Provider | Origin | Daily Tokens (Sep 13) | 30-Day Total | Status |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna | OpenAI | US | 3.9T | 45.1T | New #1, up from #3 in Sep 12 reading |
| 2 | DeepSeek V4 Flash | DeepSeek | China | 1.4T | 50.7T | Highest 30-day total overall |
| 3 | GLM 5.3 Flash | Z.ai / Zhipu | China | 1.4T | 30.4T | Steady top-3 |
| 4 | Hy4 Preview | Tencent | China | 1.4T | 34.6T | New model, entered top 5 |
| 5 | DeepSeek V4.1 Flash | DeepSeek | China | 1.3T | 4.9T | Newer variant, fast ramp |
| 6 | MiMo V2.5 | Xiaomi | China | 985.4B | 30.2T | Stable |
| 7 | DeepSeek V4 Flash (0423) | DeepSeek | China | 519.3B | 21.5T | Older variant, still active |
| 8 | Nemotron 3 Ultra 550B | Nvidia | US (open) | 472.4B | 18.3T | Leading open-weight, US |
| 9 | Hy3 | Tencent | China | 458.1B | 24.8T | Prior-gen Tencent model |
| 10 | GLM 5.3 | Z.ai / Zhipu | China | 276.9B | 7.4T | Non-flash variant |
Source: Tokenmaxxing OpenRouter rankings mirror (tokenmaxxing.com/openrouter-rankings, checked Sep 14, data dated Sep 13). Replaces the Sep 12 snapshot below - note the leader changed (GPT-5.6 Luna overtook DeepSeek V4 Flash 0731 for the daily-volume top spot) and a new Tencent model (Hy4 Preview) entered the top 5. Rankings measure adoption (tokens routed), not quality. WoW % was not available from this reading; 30-day total shown instead.
Chinese-origin model volume share on OpenRouter: from under 2% (Sep 2025) to ~48% (Aug 2026, last confirmed reading as of Sep 14 check). The fastest market share shift in the platform's history, driven by DeepSeek, GLM, Kimi, MiMo, and Tencent Hy3. Volume share, not spend share (Chinese models capture under 10% of spend due to near-zero pricing). The Sep 13 rankings snapshot in section 5 shows a US model (GPT-5.6 Luna) newly on top by daily volume, which may signal this share has since pulled back - not yet reflected here as a confirmed monthly figure.
Vercel AI Gateway monthly data (free blog), checked Sep 14. Agentic share = % of all tokens in requests that end with a tool call - Vercel's own reporting has held this at 58.9% since its Apr 26 report, so the line is flat from Apr onward rather than showing invented growth. Open-weight jumped to 62% volume / under 9% spend by a 22 Aug spot reading (up from 36%/8.6% in the Jul monthly index) - shown as a single dated point since it isn't a full monthly reading. Open-weight = DeepSeek, MiniMax, Moonshot, Z.ai. Feb 6, 2026 was flagged by OpenRouter as "potentially the last day when humans consumed more tokens than agents."
| Source | URL | Access | Frequency | What it tracks | What it misses |
|---|---|---|---|---|---|
| Ornn OCPI | data.ornn.com | Free API (3mo) | Daily | Transaction-cleared GPU rental prices | Long-term contracts, hyperscaler internal, utilization |
| Ornn OTPI | data.ornn.com | Free API (1mo) | Daily | Realized token cost by lab, volume-weighted | Self-hosted inference, batch pricing |
| OpenRouter Rankings | openrouter.ai/rankings | Free API, CC BY 4.0 | Daily | Token volume by model, 100T+/mo, 427 models | Enterprise direct API, B2C apps, China domestic |
| OpenRouter Data | openrouter.ai/data | Free | Periodic | Deep dives: model shifts, pricing, agentic trends | Same as rankings |
| Vercel AI Gateway | vercel.com/blog | Free blog | Monthly | Spend share, agentic %, open-weight %, cost/token | Enterprise self-hosted (~90% of inference), China |
| AIMultiple GPU Index | aimultiple.com/gpu-index | Free web | Monthly | 69 providers, 17 GPUs, 26mo history | Contract pricing, intra-month moves |
| Tokenmaxxing | tokenmaxxing.com | Free web | Daily | OpenRouter + company spend estimates | Private companies |