Real token volume · 46% Chinese share · Usage vs quality · Apps layer · August outlook
If you are still picking an LLM from a benchmark chart you saw two months ago, you are already behind. Bottom line: OpenRouter does not score labs—it routes real, paid traffic. Through July 25, 2026, Xiaomi's Mimo V2.5 leads at ~1.4T tokens/day, Chinese-origin labs hold ~46% of volume, and usage rank is not a quality signal—the market is splitting into a barbell of cheap open-weight volume and closed frontier pricing power on hard work. This guide covers: Top 12 models, provider share, the Apps leaderboard (Hermes / Kilo Code / OpenClaw), August outlook, role-based routing advice, and a five-step Mac acceptance checklist. See also our June rankings recap and OpenRouter API guide.
As of July 25, the top three models by daily token volume are Mimo V2.5 (1.4T/day), DeepSeek V4 Flash (943.9B/day), and Tencent Hy3 (590B/day). Seven of the top ten models are Chinese; NVIDIA Nemotron 3 Ultra, Claude, and Gemini still anchor the US side.
| Rank | Model | Lab | Daily tokens | 30-day total |
|---|---|---|---|---|
| 1 | Mimo V2.5 | Xiaomi | 1.4T | 31.2T |
| 2 | DeepSeek V4 Flash | DeepSeek | 943.9B | 23.6T |
| 3 | Hy3 | Tencent | 590B | 23.4T |
| 4 | Nemotron 3 Ultra 550B (free) | NVIDIA | 428.6B | 9T |
| 5 | DeepSeek V4 Pro | DeepSeek | 413.7B | 11.6T |
| 6 | GLM 5.2 | Z.ai | 316.7B | 13.3T |
| 7 | MiniMax M3 | MiniMax | 262.5B | 15.1T |
| 8 | Step 3.7 Flash | StepFun | 204.8B | 5.9T |
| 9 | Kimi K3 | Moonshot | 157.6B | 1.6T (new) |
| 10 | Ling 3.0 Flash | InclusionAI | 128.3B | 417.3B |
| 11 | Gemini 3 Flash Preview | 106.3B | 4T | |
| 12 | Claude Sonnet 5 | Anthropic | 99.5B | 3.6T |
At provider level, Chinese labs now account for ~46% of identified token volume—up from under 2% a year ago. US labs (OpenAI, Anthropic, Google) fell from ~70% in mid-2025 to 30–36%. This is pricing math: DeepSeek V4 Flash lists around $0.05–$0.14/M input vs GPT-5.5 near $5/M—roughly a 35× gap.
Citable: DeepSeek remains the most stable #1 provider (~16–18%), but the "model of the month" crown keeps rotating—MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July.
A cheap, fast model wired into one high-traffic consumer app can outrank a more capable model reserved for the hardest 10% of work.
Confusing volume with capability: Roleplay, chat, and light coding dominate open-model traffic—not enterprise hard reasoning.
Ignoring spend mix: By category, chat is 35.7%, agentic workflows 30.4%, code 26.5%—but in Classification, Claude Sonnet 4.6 and Opus 4.7 tie at 13.5% spend each.
Daily volatility: Claude Opus 4.8 was top 10 on 7/24; Ling 3.0 Flash pushed it out of the top 12 by 7/25. Always cite a cutoff date.
Rising security weight: OpenAI's sandbox-escape disclosure this week will push "vendor safety track record" into enterprise scorecards.
Single-model lock-in: Claude Opus 5 (July 24) tops FrontierBench v0.1 at 43.3% while holding Opus-tier $5/$25 pricing—frontier labs still command hard-task pricing power.
| Segment | Typical workloads | Volume chart | Hard-task spend |
|---|---|---|---|
| High-volume, tolerant | Chat, creative, roleplay, routine coding | Chinese open Flash models | Low share |
| Low-tolerance, high value | Complex reasoning, agent planning, classification | Rarely top 10 | Claude / GPT frontier |
Model rankings show which brain is popular. The Apps leaderboard shows what that brain is doing:
| Rank | App | Type | Share (~) |
|---|---|---|---|
| 1 | Hermes Agent | Personal / CLI agent | ~45% |
| 2 | Kilo Code | Coding agent | ~13% |
| 3 | OpenClaw | General agent | ~9% |
| 4 | Claude Code | Coding agent | ~6% |
| 5 | Descript | Content | ~4.5% |
| 6–10 | pi, Lemonade, ISEKAI ZERO, Janitor AI, Cline | Agent / companion / roleplay | ~1.7%–3.3% each |
Cline → Roo Code → Kilo Code are three forks of the same lineage; the youngest fork, Kilo Code, has overtaken both ancestors. First-mover advantage in open-source dev tooling clearly does not last.
OpenRouter × a16z's State of AI report found creative roleplay accounts for more than half of open-model usage. If your view of AI comes only from enterprise headlines, you are missing half the market.
Chinese open-weight share likely climbs toward or past 50% unless a major US provider cuts price.
The monthly #1 will keep rotating across Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot.
Anthropic may ship a cheaper volume tier—Opus 5 is their fourth flagship in under two months.
Kimi K3's 1.4TB weights should see community quantization within 2–4 weeks.
Security and governance become selection criteria—US "AI Kill Switch" legislation and a White House pre-release review framework are expected before August.
| Model | Input/M | Output/M | Positioning |
|---|---|---|---|
| DeepSeek V4 Flash | ~$0.05–0.14 | ~$0.24–0.28 | Best value; agentic coding |
| MiniMax M3 | $0.10 | $1.21 | Long context on a budget |
| GLM 5.2 | $0.45 | $3.31 | Open-weight Opus-style planning |
| Kimi K3 | ~$3 | ~$15 | 1.4TB open weights |
| Claude Opus 5 | $5 (fast $10) | $25 (fast $50) | Closed frontier; hard tasks |
| Role | Recommended approach |
|---|---|
| Indie / small team | Use OpenRouter as a sandbox; start coding on DeepSeek V4 Flash and GLM 5.2, reserve Opus 5 / GPT-5.6 for steps that actually fail |
| Infra lead | Do not select by usage rank alone; tier routes by task risk and add vendor safety to scorecards |
| Agent / devtool builder | Study the Kilo Code fork story; companion/entertainment volume is real even if invisible in enterprise press |
On Mac with Claude Code, OpenClaw, or Kilo Code, implement barbell routing in order:
Tier tasks: S-tier (hardest 5%) → frontier closed models; A-tier → DeepSeek / GLM; B-tier → Flash class.
Single router: OpenRouter or OpenClaw models config—no hard-coded provider in app code.
Spend alerts: Daily/weekly caps per model so agents do not loop on Opus by mistake.
Quality probes: 10–20 fixed regression prompts before shifting traffic to a new model.
VNC acceptance: OAuth, Gateway, browser MCP, and permission dialogs verified in a graphical Mac session co-located with the agent.
One key, 400+ models.
Read →61% Chinese share and Q3 outlook.
Read →Late-July dual headlines.
Read →No—it sorts by paid token volume. On hard-task spend, Claude Sonnet 4.6 and Opus 4.7 still tie for the lead.
As of July 25, Xiaomi Mimo V2.5 at ~1.4T tokens/day, then DeepSeek V4 Flash and Tencent Hy3.
Roughly 46% combined, up from under 2% a year ago; US labs sit around 30–36%.
Nous Research's open self-improving agent holds ~45% of app token share—high-frequency personal automation, not a quality vote.
The line to remember from July: capability and popularity are diverging. Chinese open-weight models bought half the market with price; US closed-frontier labs defend the other half on hard tasks and safety credibility.
If your daily driver is Windows or Linux but you need macOS for Claude Code, OpenClaw, or Kilo Code, buying a Mac adds depreciation and sleep-policy overhead; SSH-only setups miss OAuth and permission dialogs. Renting a VNCMac remote Mac with VNC lets you validate barbell routing against the OpenRouter dashboard on the same machine as your Gateway.
Data as of July 25, 2026—verify live figures at openrouter.ai/rankings before citing.