Opus 5 becomes Claude Max default · White House distillation claim · Greenblatt evidence
Two unrelated-looking stories this week both orbit the same question: where does frontier capability actually come from, and who pays for it? On July 24, 2026, Anthropic shipped Claude Opus 5 as the new Claude Max default. Days earlier, Moonshot AI's freshly launched Kimi K3 landed in a Kimi K3 distillation controversy after White House officials accused industrial-scale copying—and independent researcher Ryan Greenblatt showed the model sometimes self-identifies as Claude and leaks internal deployment IDs. This piece covers Opus 5 benchmarks, pricing, safety, and data retention; K3 specs; the accusation timeline; expert pushback; community reaction; a combined cost thesis, and FAQ.
Anthropic released Claude Opus 5 (model ID: claude-opus-5) on July 24, 2026 (US Pacific). It immediately became the Claude Max default and the strongest tier available to Claude Pro users. The positioning is deliberate: daily-driver intelligence close to flagship Claude Fable 5, at roughly half the token bill.
| Spec | Detail |
|---|---|
| Pricing | Input $5 / million tokens, output $25 / million tokens (unchanged from Opus 4.8) |
| Context window | 1 million tokens (default and only tier) |
| Max output | 128K tokens |
| Default reasoning | Thinking enabled by default; Effort parameter controls depth |
| Platforms | Claude API / Platform / AWS Bedrock / Google Vertex AI / Microsoft Foundry |
| Fast mode | ~2.5× speed at 2× base price (same as Opus 4.8) |
| Data retention | No mandatory retention on default access (Fable 5 / Mythos 5 require opt-in to a 30-day retention policy) |
Early customer quotes: Cursor described Opus 5 as "near Fable 5 intelligence at Opus speed and cost." Zapier reported AutomationBench leadership without higher token burn. Box cited +11% on analytics workflows, +17% on diligence tasks, and +8% overall accuracy.
Automated behavioral audits position Opus 5 as Anthropic's most aligned model to date: lowest deception rate, hardest to steer into abuse, and safest on irreversible reckless actions. On dual-use frontier skills (offensive cyber, biosecurity), Anthropic intentionally did not push the frontier—restricted Mythos 5 still holds that lane. The cyber classifier is roughly 85% less restrictive than Fable 5 (fewer false blocks), enabling source-level vulnerability discovery while still blocking binary scanning, penetration testing, and exploit generation. Biosecurity prompts blocked on Fable 5 now route to Opus 5 instead of falling back to Opus 4.8.
| Date | Event |
|---|---|
| 2026-06-09 | Claude Fable 5 / Mythos 5 announced |
| 2026-07-01 | Fable 5 generally available to the public |
| 2026-07-24 | Claude Opus 5 ships; becomes Claude Max default |
Fable 5 price floor: roughly $10/$50 per million input/output tokens doubles daily agent spend
30-day retention policy: Fable 5 / Mythos 5 require opt-in retention—problematic for strict enterprise compliance
Overkill capacity: most coding and automation workloads do not need Mythos-class dual-use headroom
Default-model lag: Pro/Max users still on legacy Opus miss roughly 2× software-engineering uplift
Validation environment: side-by-side Opus 5 vs Fable 5 in Cursor / Claude Code often needs macOS GUI sessions and OAuth callbacks
| Dimension | Claude Opus 5 | Claude Fable 5 |
|---|---|---|
| Relative pricing | ~half ($5/$25) | ~$10/$50 |
| CursorBench 3.2 (max) | 0.5% below peak | Peak reference |
| Default data retention | Not mandatory | 30-day opt-in required |
| Claude Max default | Yes (from 7/24) | No |
| Dual-use frontier (cyber/bio) | Intentionally not maxed | Mythos 5 restricted tier leads |
Moonshot AI launched Kimi K3 on July 16, 2026: 2.8 trillion total parameters—the first open-weight model to enter 3T-class territory. Architecture is sparse MoE with 896 experts and 16 active per token (~50B activated parameters). Training stack adds Kimi Delta Attention (KDA), Attention Residuals, and Stable LatentMoE for ~2.5× efficiency vs K2. Context is 1M tokens with native vision. Full weights were promised for July 27 (not yet public at time of writing).
| Benchmark | Score | Note |
|---|---|---|
| GPQA-Diamond | 93.5% | Highest open model at launch |
| Terminal-Bench 2.1 | 88.3% | 0.5 pt behind GPT-5.6 Sol |
| BrowseComp | 91.2% | Top score at launch |
| Humanity's Last Exam (tools) | 56.0% | — |
| MCP Atlas | 84.2% | — |
| Program Bench | 77.8% | Best composite |
| SWE Marathon | 42.0% | Best composite |
| DeepSearchQA (F1) | 95.0% | — |
Moonshot positions K3 just below Claude Fable 5 and GPT-5.6 Sol at materially lower API prices. For architecture depth and the July 27 weight timeline, see our open-weights countdown and deep review.
On July 22–23, 2026, White House OSTP director Michael Kratsios posted on X that Moonshot AI engaged in "large-scale, covert industrial distillation" aimed at stealing Anthropic Fable capabilities, and alleged use of export-controlled Nvidia GB300 (Grace Blackwell 300) chips—possibly via servers in Thailand. Verbatim: "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable."
Treasury Secretary Scott Bessent echoed that "watermarks" from US models appear in many Chinese systems—but Treasury did not define what "watermark" means in this context. Kratsios published no supporting evidence; Moonshot did not respond to training-process inquiries.
This is not new ground: in February 2026, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of "industrial-scale distillation attacks," claiming it identified more than 3.4 million anomalous Moonshot-linked interactions via fake accounts—"clearly deviating from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use"—and tied some activity to Moonshot leadership through API request metadata. Moonshot never confirmed or denied.
| Date | Event |
|---|---|
| 2026-02 | Anthropic first public distillation accusation vs Moonshot / DeepSeek / MiniMax |
| 2026-07-01 | Claude Fable 5 public availability |
| 2026-07-16 | Kimi K3 API / product launch |
| 2026-07-22/23 | White House OSTP director Kratsios public distillation + chip allegation |
| 2026-07-23 | TechCrunch publishes expert skepticism piece |
| ~2026-07-24 | Ryan Greenblatt publishes "K3 self-identifies as Claude" analysis |
| 2026-07-27 (planned) | Kimi K3 full open weights for independent verification |
TechCrunch (July 23) interviewed several independent researchers who doubt the distillation narrative on timing alone—Fable 5 was publicly available only from July 1 to K3's July 16 launch: two weeks. Deep distillation, especially RL-style teacher scoring at scale, implies enormous API spend, compute, and pipeline time that skeptics call unrealistic in that window.
Braden Hancock (Laude Institute / Snorkel AI co-founder): "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation... Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks."
Nathan Lambert (Allen Institute for AI): as Chinese models approach the frontier, marginal returns from distillation shrink; the gap is increasingly bought with expensive RL. If distillation were that effective, everyone would have distilled their way to GLM or K3 parity—but they have not.
Others noted a double standard: Elon Musk testified that xAI distilled OpenAI models while training Grok and called the practice common—raising the harder policy question of where "normal technical borrowing" ends and "covert industrial extraction" begins.
Redwood Research chief scientist Ryan Greenblatt published statistical work (~July 24; GitHub: rgreenblatt/which_claude_is_k3) comparing identity behavior across models. This is the clearest technical hook behind searches like Why does Kimi K3 say it's Claude and Kimi K3 self-identifies as Claude—queries with almost no English competition at research time.
Important caveat: Greenblatt stresses these findings do not prove distillation. Identity confusion could also come from data pollution, leaked system prompts, or public dataset synthesis—but combined with Anthropic's February telemetry and the statistical regularity, this is the most grounded technical signal in the Kimi K3 distillation controversy, not just political rhetoric.
Reddit r/LocalLLaMA threads split three ways: celebration that open and closed gaps may shrink to days instead of months; jokes that almost nobody can run 2.8T locally; and pragmatists who say K3's real sell is low price plus lighter moderation, not "beat Fable 5 on every chart."
Full weights were not downloadable when the controversy peaked, so external researchers could not yet independently verify parameter count, architecture details, or benchmark reproduction—a major reason discourse stayed speculative. Policy layers matter too: US officials are debating restrictions on Chinese open-weight releases, and export-control talk resurfaced (echoing the May 2026 Supermicro founder advanced-chip smuggling indictment).
Opus 5 answers the market with "half price, near-flagship." Kimi K3 answers with "open-weight narrative plus aggressive API pricing." The distillation fight is really a public accounting of whether your discount came from engineering—or from someone else's model.
Practical guidance for builders: if compliance, retention, and supply-chain risk matter, Opus 5's unchanged price with a capability jump is compelling. If you optimize for floor cost and self-host control, wait for July 27 weights and independent reruns before committing—until then, K3 remains "vendor scorecard plus speculation."
Compliance first: Opus 5 + no mandatory retention + Anthropic official API
Floor cost: defer self-host vs API decisions until K3 weights and third-party benchmarks land
Agent acceptance: compare Cursor / Claude Code / Kimi Code on CursorBench-class tasks
GUI session: OAuth, Gateway, and model switching are easier to validate on a VNC remote Mac desktop
After 7/27: watch for independent benchmark reruns once weights drop, to see whether the distillation claims hold up
At the same token tier, Opus 5 runs about half the price ($5/$25 vs roughly $10/$50 per million input/output tokens). CursorBench 3.2 max effort trails Fable 5's peak by less than 1%.
Yes. From July 24, 2026, Opus 5 is the Claude Max default and the top tier available to Claude Pro subscribers.
Still disputed and not finally proven. The White House claim lacks public evidence; timeline skeptics say two weeks is too short for deep distillation after Fable 5 went public; Greenblatt's Claude self-ID and internal deployment strings remain the strongest technical indirect signal.
Ryan Greenblatt's analysis shows K3 disproportionately answers identity prompts with Claude branding and internal IDs like claude-opus-4-5-20250929—often more precisely than real Claude models. That pattern suggests training data mixed with Claude API logs or labeled synthetic samples, not mere conversational mimicry.
K3 is open-weight (downloadable weights), not strictly fully open-source in the OSI sense. Expect Moonshot's modified-MIT-style license; treat the LICENSE file published on the July 27, 2026 release as authoritative.
Opus 5 compresses flagship-grade agent capability into a daily token budget. K3 pushes frontier performance into an open-weight story—but distillation accusations plus Claude self-identification stats remind builders that model selection is not just benchmark shopping. Retention policy, compliance, and provenance matter.
If you want to validate Opus 5 vs K3 routing inside Cursor, Claude Code, or Kimi Code, Windows and Linux primary machines often break on OAuth callbacks, macOS permission dialogs, and Gateway GUI comparisons. Buying a Mac adds sleep policies, OS updates, and depreciation for episodic agent work. Renting a VNCMac remote Mac with VNC desktop access lets you switch models, capture logs, and reconcile identity quirks on the same host as your Gateway—turning this week's headlines into reproducible agent validation instead of headline scrolling.