AI Weekly July 25, 2026 ~24 min read Claude Opus 5 Kimi K3

Claude Opus 5 Cuts the Price in Half
Meanwhile Kimi K3 Gets Caught Calling Itself Claude

Opus 5 becomes Claude Max default · White House distillation claim · Greenblatt evidence

Concept art for Claude Opus 5 launch and Kimi K3 distillation controversy AI weekly roundup

Two unrelated-looking stories this week both orbit the same question: where does frontier capability actually come from, and who pays for it? On July 24, 2026, Anthropic shipped Claude Opus 5 as the new Claude Max default. Days earlier, Moonshot AI's freshly launched Kimi K3 landed in a Kimi K3 distillation controversy after White House officials accused industrial-scale copying—and independent researcher Ryan Greenblatt showed the model sometimes self-identifies as Claude and leaks internal deployment IDs. This piece covers Opus 5 benchmarks, pricing, safety, and data retention; K3 specs; the accusation timeline; expert pushback; community reaction; a combined cost thesis, and FAQ.

01

Claude Opus 5: Not the Flagship, Maybe the One You Actually Use

Anthropic released Claude Opus 5 (model ID: claude-opus-5) on July 24, 2026 (US Pacific). It immediately became the Claude Max default and the strongest tier available to Claude Pro users. The positioning is deliberate: daily-driver intelligence close to flagship Claude Fable 5, at roughly half the token bill.

SpecDetail
PricingInput $5 / million tokens, output $25 / million tokens (unchanged from Opus 4.8)
Context window1 million tokens (default and only tier)
Max output128K tokens
Default reasoningThinking enabled by default; Effort parameter controls depth
PlatformsClaude API / Platform / AWS Bedrock / Google Vertex AI / Microsoft Foundry
Fast mode~2.5× speed at 2× base price (same as Opus 4.8)
Data retentionNo mandatory retention on default access (Fable 5 / Mythos 5 require opt-in to a 30-day retention policy)

Benchmark Highlights (Anthropic Official)

  • Frontier-Bench v0.1 (software engineering): beats all listed models, more than Opus 4.8, lower per-task cost
  • CursorBench 3.2: max effort within 0.5% of Fable 5 peak at half the cost; high/xhigh/max tiers lead on price-performance across the field
  • ARC-AGI 3: roughly the next-best score
  • Zapier AutomationBench: ~1.5× pass rate vs runner-up; lowest effort still beats every other model; 100% on end-to-end account-health workflows (prior models failed)
  • OSWorld 2.0: best at any cost tier; exceeds Fable 5's best run using slightly more than one-third of Fable 5 spend
  • Science / life sciences: broad gains over Opus 4.8; internal evals show +10.2 points on spectral structure inference and +7.7 points on protein variant function prediction

Early customer quotes: Cursor described Opus 5 as "near Fable 5 intelligence at Opus speed and cost." Zapier reported AutomationBench leadership without higher token burn. Box cited +11% on analytics workflows, +17% on diligence tasks, and +8% overall accuracy.

Alignment and Safety

Automated behavioral audits position Opus 5 as Anthropic's most aligned model to date: lowest deception rate, hardest to steer into abuse, and safest on irreversible reckless actions. On dual-use frontier skills (offensive cyber, biosecurity), Anthropic intentionally did not push the frontier—restricted Mythos 5 still holds that lane. The cyber classifier is roughly 85% less restrictive than Fable 5 (fewer false blocks), enabling source-level vulnerability discovery while still blocking binary scanning, penetration testing, and exploit generation. Biosecurity prompts blocked on Fable 5 now route to Opus 5 instead of falling back to Opus 4.8.

DateEvent
2026-06-09Claude Fable 5 / Mythos 5 announced
2026-07-01Fable 5 generally available to the public
2026-07-24Claude Opus 5 ships; becomes Claude Max default
02

Before You Choose: Claude Opus 5 vs Fable 5 Decision Matrix

  1. 01

    Fable 5 price floor: roughly $10/$50 per million input/output tokens doubles daily agent spend

  2. 02

    30-day retention policy: Fable 5 / Mythos 5 require opt-in retention—problematic for strict enterprise compliance

  3. 03

    Overkill capacity: most coding and automation workloads do not need Mythos-class dual-use headroom

  4. 04

    Default-model lag: Pro/Max users still on legacy Opus miss roughly 2× software-engineering uplift

  5. 05

    Validation environment: side-by-side Opus 5 vs Fable 5 in Cursor / Claude Code often needs macOS GUI sessions and OAuth callbacks

DimensionClaude Opus 5Claude Fable 5
Relative pricing~half ($5/$25)~$10/$50
CursorBench 3.2 (max)0.5% below peakPeak reference
Default data retentionNot mandatory30-day opt-in required
Claude Max defaultYes (from 7/24)No
Dual-use frontier (cyber/bio)Intentionally not maxedMythos 5 restricted tier leads
03

Kimi K3: First 3T-Class Open-Weight Model and Its Benchmarks

Moonshot AI launched Kimi K3 on July 16, 2026: 2.8 trillion total parameters—the first open-weight model to enter 3T-class territory. Architecture is sparse MoE with 896 experts and 16 active per token (~50B activated parameters). Training stack adds Kimi Delta Attention (KDA), Attention Residuals, and Stable LatentMoE for ~2.5× efficiency vs K2. Context is 1M tokens with native vision. Full weights were promised for July 27 (not yet public at time of writing).

BenchmarkScoreNote
GPQA-Diamond93.5%Highest open model at launch
Terminal-Bench 2.188.3%0.5 pt behind GPT-5.6 Sol
BrowseComp91.2%Top score at launch
Humanity's Last Exam (tools)56.0%
MCP Atlas84.2%
Program Bench77.8%Best composite
SWE Marathon42.0%Best composite
DeepSearchQA (F1)95.0%

Moonshot positions K3 just below Claude Fable 5 and GPT-5.6 Sol at materially lower API prices. For architecture depth and the July 27 weight timeline, see our open-weights countdown and deep review.

04

The Distillation Row: White House Accusation and Anthropic's February Claims

On July 22–23, 2026, White House OSTP director Michael Kratsios posted on X that Moonshot AI engaged in "large-scale, covert industrial distillation" aimed at stealing Anthropic Fable capabilities, and alleged use of export-controlled Nvidia GB300 (Grace Blackwell 300) chips—possibly via servers in Thailand. Verbatim: "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable."

Treasury Secretary Scott Bessent echoed that "watermarks" from US models appear in many Chinese systems—but Treasury did not define what "watermark" means in this context. Kratsios published no supporting evidence; Moonshot did not respond to training-process inquiries.

This is not new ground: in February 2026, Anthropic publicly accused Moonshot, DeepSeek, and MiniMax of "industrial-scale distillation attacks," claiming it identified more than 3.4 million anomalous Moonshot-linked interactions via fake accounts—"clearly deviating from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use"—and tied some activity to Moonshot leadership through API request metadata. Moonshot never confirmed or denied.

DateEvent
2026-02Anthropic first public distillation accusation vs Moonshot / DeepSeek / MiniMax
2026-07-01Claude Fable 5 public availability
2026-07-16Kimi K3 API / product launch
2026-07-22/23White House OSTP director Kratsios public distillation + chip allegation
2026-07-23TechCrunch publishes expert skepticism piece
~2026-07-24Ryan Greenblatt publishes "K3 self-identifies as Claude" analysis
2026-07-27 (planned)Kimi K3 full open weights for independent verification
05

Independent Experts Push Back: The Timeline Does Not Add Up

TechCrunch (July 23) interviewed several independent researchers who doubt the distillation narrative on timing alone—Fable 5 was publicly available only from July 1 to K3's July 16 launch: two weeks. Deep distillation, especially RL-style teacher scoring at scale, implies enormous API spend, compute, and pipeline time that skeptics call unrealistic in that window.

"

Braden Hancock (Laude Institute / Snorkel AI co-founder): "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation... Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks."

"

Nathan Lambert (Allen Institute for AI): as Chinese models approach the frontier, marginal returns from distillation shrink; the gap is increasingly bought with expensive RL. If distillation were that effective, everyone would have distilled their way to GLM or K3 parity—but they have not.

Others noted a double standard: Elon Musk testified that xAI distilled OpenAI models while training Grok and called the practice common—raising the harder policy question of where "normal technical borrowing" ends and "covert industrial extraction" begins.

06

The Hard Evidence: Why Does Kimi K3 Say It's Claude?

Redwood Research chief scientist Ryan Greenblatt published statistical work (~July 24; GitHub: rgreenblatt/which_claude_is_k3) comparing identity behavior across models. This is the clearest technical hook behind searches like Why does Kimi K3 say it's Claude and Kimi K3 self-identifies as Claude—queries with almost no English competition at research time.

  • Kimi K3 abnormally often claims to be Claude, emitting internal deployment IDs such as claude-opus-4-5-20250929 and claude-sonnet-4-5-20250929
  • Real Claude Sonnet 4.5 says "I am Claude Sonnet 4.5"; Opus 4.5 often omits or misstates internal version strings
  • Greenblatt's read: K3 reproduces teacher deployment metadata more accurately than the teacher itself—hard to explain as casual style mimicry; more consistent with training on Claude data labeled with deployment metadata (API logs or synthetic labeling pipelines)
  • Version claims cluster on the Claude 4.5 era (late 2025), not current Fable/Mythos; K2 signals point to earlier Claude Sonnet 4 (mid-2025)—a "chasing the latest generation" pattern

Important caveat: Greenblatt stresses these findings do not prove distillation. Identity confusion could also come from data pollution, leaked system prompts, or public dataset synthesis—but combined with Anthropic's February telemetry and the statistical regularity, this is the most grounded technical signal in the Kimi K3 distillation controversy, not just political rhetoric.

07

Community Reaction and the July 27 Weight Cliffhanger

Reddit r/LocalLLaMA threads split three ways: celebration that open and closed gaps may shrink to days instead of months; jokes that almost nobody can run 2.8T locally; and pragmatists who say K3's real sell is low price plus lighter moderation, not "beat Fable 5 on every chart."

Full weights were not downloadable when the controversy peaked, so external researchers could not yet independently verify parameter count, architecture details, or benchmark reproduction—a major reason discourse stayed speculative. Policy layers matter too: US officials are debating restrictions on Chinese open-weight releases, and export-control talk resurfaced (echoing the May 2026 Supermicro founder advanced-chip smuggling indictment).

08

Read Both Stories Together: Cost Efficiency Is the Real Battlefield

Opus 5 answers the market with "half price, near-flagship." Kimi K3 answers with "open-weight narrative plus aggressive API pricing." The distillation fight is really a public accounting of whether your discount came from engineering—or from someone else's model.

Practical guidance for builders: if compliance, retention, and supply-chain risk matter, Opus 5's unchanged price with a capability jump is compelling. If you optimize for floor cost and self-host control, wait for July 27 weights and independent reruns before committing—until then, K3 remains "vendor scorecard plus speculation."

  1. 01

    Compliance first: Opus 5 + no mandatory retention + Anthropic official API

  2. 02

    Floor cost: defer self-host vs API decisions until K3 weights and third-party benchmarks land

  3. 03

    Agent acceptance: compare Cursor / Claude Code / Kimi Code on CursorBench-class tasks

  4. 04

    GUI session: OAuth, Gateway, and model switching are easier to validate on a VNC remote Mac desktop

  5. 05

    After 7/27: watch for independent benchmark reruns once weights drop, to see whether the distillation claims hold up

FAQ

Frequently Asked Questions

At the same token tier, Opus 5 runs about half the price ($5/$25 vs roughly $10/$50 per million input/output tokens). CursorBench 3.2 max effort trails Fable 5's peak by less than 1%.

Yes. From July 24, 2026, Opus 5 is the Claude Max default and the top tier available to Claude Pro subscribers.

Still disputed and not finally proven. The White House claim lacks public evidence; timeline skeptics say two weeks is too short for deep distillation after Fable 5 went public; Greenblatt's Claude self-ID and internal deployment strings remain the strongest technical indirect signal.

Ryan Greenblatt's analysis shows K3 disproportionately answers identity prompts with Claude branding and internal IDs like claude-opus-4-5-20250929—often more precisely than real Claude models. That pattern suggests training data mixed with Claude API logs or labeled synthetic samples, not mere conversational mimicry.

K3 is open-weight (downloadable weights), not strictly fully open-source in the OSI sense. Expect Moonshot's modified-MIT-style license; treat the LICENSE file published on the July 27, 2026 release as authoritative.

Closing Thoughts

Opus 5 compresses flagship-grade agent capability into a daily token budget. K3 pushes frontier performance into an open-weight story—but distillation accusations plus Claude self-identification stats remind builders that model selection is not just benchmark shopping. Retention policy, compliance, and provenance matter.

If you want to validate Opus 5 vs K3 routing inside Cursor, Claude Code, or Kimi Code, Windows and Linux primary machines often break on OAuth callbacks, macOS permission dialogs, and Gateway GUI comparisons. Buying a Mac adds sleep policies, OS updates, and depreciation for episodic agent work. Renting a VNCMac remote Mac with VNC desktop access lets you switch models, capture logs, and reconcile identity quirks on the same host as your Gateway—turning this week's headlines into reproducible agent validation instead of headline scrolling.