LLM Trends July 27, 2026 ~16 min read OpenRouter DeepSeek

OpenRouter Rankings July 2026
Who's Actually Winning the AI Model Race

Real token volume · 46% Chinese share · Usage vs quality · Apps layer · August outlook

OpenRouter July 2026 rankings chart showing Chinese model market share

If you are still picking an LLM from a benchmark chart you saw two months ago, you are already behind. Bottom line: OpenRouter does not score labs—it routes real, paid traffic. Through July 25, 2026, Xiaomi's Mimo V2.5 leads at ~1.4T tokens/day, Chinese-origin labs hold ~46% of volume, and usage rank is not a quality signal—the market is splitting into a barbell of cheap open-weight volume and closed frontier pricing power on hard work. This guide covers: Top 12 models, provider share, the Apps leaderboard (Hermes / Kilo Code / OpenClaw), August outlook, role-based routing advice, and a five-step Mac acceptance checklist. See also our June rankings recap and OpenRouter API guide.

01

July leaderboard: Xiaomi takes #1, Chinese labs cross 46%

As of July 25, the top three models by daily token volume are Mimo V2.5 (1.4T/day), DeepSeek V4 Flash (943.9B/day), and Tencent Hy3 (590B/day). Seven of the top ten models are Chinese; NVIDIA Nemotron 3 Ultra, Claude, and Gemini still anchor the US side.

RankModelLabDaily tokens30-day total
1Mimo V2.5Xiaomi1.4T31.2T
2DeepSeek V4 FlashDeepSeek943.9B23.6T
3Hy3Tencent590B23.4T
4Nemotron 3 Ultra 550B (free)NVIDIA428.6B9T
5DeepSeek V4 ProDeepSeek413.7B11.6T
6GLM 5.2Z.ai316.7B13.3T
7MiniMax M3MiniMax262.5B15.1T
8Step 3.7 FlashStepFun204.8B5.9T
9Kimi K3Moonshot157.6B1.6T (new)
10Ling 3.0 FlashInclusionAI128.3B417.3B
11Gemini 3 Flash PreviewGoogle106.3B4T
12Claude Sonnet 5Anthropic99.5B3.6T

At provider level, Chinese labs now account for ~46% of identified token volume—up from under 2% a year ago. US labs (OpenAI, Anthropic, Google) fell from ~70% in mid-2025 to 30–36%. This is pricing math: DeepSeek V4 Flash lists around $0.05–$0.14/M input vs GPT-5.5 near $5/M—roughly a 35× gap.

Citable: DeepSeek remains the most stable #1 provider (~16–18%), but the "model of the month" crown keeps rotating—MiniMax M2.5 in February, MiMo-V2-Pro in April, Mimo V2.5 in July.

02

What most coverage misses: usage rank is not a quality signal

A cheap, fast model wired into one high-traffic consumer app can outrank a more capable model reserved for the hardest 10% of work.

Five costs of misreading the leaderboard

  1. 01

    Confusing volume with capability: Roleplay, chat, and light coding dominate open-model traffic—not enterprise hard reasoning.

  2. 02

    Ignoring spend mix: By category, chat is 35.7%, agentic workflows 30.4%, code 26.5%—but in Classification, Claude Sonnet 4.6 and Opus 4.7 tie at 13.5% spend each.

  3. 03

    Daily volatility: Claude Opus 4.8 was top 10 on 7/24; Ling 3.0 Flash pushed it out of the top 12 by 7/25. Always cite a cutoff date.

  4. 04

    Rising security weight: OpenAI's sandbox-escape disclosure this week will push "vendor safety track record" into enterprise scorecards.

  5. 05

    Single-model lock-in: Claude Opus 5 (July 24) tops FrontierBench v0.1 at 43.3% while holding Opus-tier $5/$25 pricing—frontier labs still command hard-task pricing power.

SegmentTypical workloadsVolume chartHard-task spend
High-volume, tolerantChat, creative, roleplay, routine codingChinese open Flash modelsLow share
Low-tolerance, high valueComplex reasoning, agent planning, classificationRarely top 10Claude / GPT frontier
03

Apps layer: coding agents dominate; roleplay is the invisible half

Model rankings show which brain is popular. The Apps leaderboard shows what that brain is doing:

RankAppTypeShare (~)
1Hermes AgentPersonal / CLI agent~45%
2Kilo CodeCoding agent~13%
3OpenClawGeneral agent~9%
4Claude CodeCoding agent~6%
5DescriptContent~4.5%
6–10pi, Lemonade, ISEKAI ZERO, Janitor AI, ClineAgent / companion / roleplay~1.7%–3.3% each

Cline → Roo Code → Kilo Code are three forks of the same lineage; the youngest fork, Kilo Code, has overtaken both ancestors. First-mover advantage in open-source dev tooling clearly does not last.

OpenRouter × a16z's State of AI report found creative roleplay accounts for more than half of open-model usage. If your view of AI comes only from enterprise headlines, you are missing half the market.

04

August outlook and pricing snapshot

  1. 01

    Chinese open-weight share likely climbs toward or past 50% unless a major US provider cuts price.

  2. 02

    The monthly #1 will keep rotating across Xiaomi, DeepSeek, Tencent, Z.ai, MiniMax, and Moonshot.

  3. 03

    Anthropic may ship a cheaper volume tier—Opus 5 is their fourth flagship in under two months.

  4. 04

    Kimi K3's 1.4TB weights should see community quantization within 2–4 weeks.

  5. 05

    Security and governance become selection criteria—US "AI Kill Switch" legislation and a White House pre-release review framework are expected before August.

Representative pricing (July 2026)

ModelInput/MOutput/MPositioning
DeepSeek V4 Flash~$0.05–0.14~$0.24–0.28Best value; agentic coding
MiniMax M3$0.10$1.21Long context on a budget
GLM 5.2$0.45$3.31Open-weight Opus-style planning
Kimi K3~$3~$151.4TB open weights
Claude Opus 5$5 (fast $10)$25 (fast $50)Closed frontier; hard tasks
05

Role-based takeaways + five-step Mac routing checklist

RoleRecommended approach
Indie / small teamUse OpenRouter as a sandbox; start coding on DeepSeek V4 Flash and GLM 5.2, reserve Opus 5 / GPT-5.6 for steps that actually fail
Infra leadDo not select by usage rank alone; tier routes by task risk and add vendor safety to scorecards
Agent / devtool builderStudy the Kilo Code fork story; companion/entertainment volume is real even if invisible in enterprise press

On Mac with Claude Code, OpenClaw, or Kilo Code, implement barbell routing in order:

  1. 01

    Tier tasks: S-tier (hardest 5%) → frontier closed models; A-tier → DeepSeek / GLM; B-tier → Flash class.

  2. 02

    Single router: OpenRouter or OpenClaw models config—no hard-coded provider in app code.

  3. 03

    Spend alerts: Daily/weekly caps per model so agents do not loop on Opus by mistake.

  4. 04

    Quality probes: 10–20 fixed regression prompts before shifting traffic to a new model.

  5. 05

    VNC acceptance: OAuth, Gateway, browser MCP, and permission dialogs verified in a graphical Mac session co-located with the agent.

Further reading
FAQ

FAQ

No—it sorts by paid token volume. On hard-task spend, Claude Sonnet 4.6 and Opus 4.7 still tie for the lead.

As of July 25, Xiaomi Mimo V2.5 at ~1.4T tokens/day, then DeepSeek V4 Flash and Tencent Hy3.

Roughly 46% combined, up from under 2% a year ago; US labs sit around 30–36%.

Nous Research's open self-improving agent holds ~45% of app token share—high-frequency personal automation, not a quality vote.

Closing

The line to remember from July: capability and popularity are diverging. Chinese open-weight models bought half the market with price; US closed-frontier labs defend the other half on hard tasks and safety credibility.

If your daily driver is Windows or Linux but you need macOS for Claude Code, OpenClaw, or Kilo Code, buying a Mac adds depreciation and sleep-policy overhead; SSH-only setups miss OAuth and permission dialogs. Renting a VNCMac remote Mac with VNC lets you validate barbell routing against the OpenRouter dashboard on the same machine as your Gateway.

Data as of July 25, 2026—verify live figures at openrouter.ai/rankings before citing.