July 30 repricing · Luna $0.20/$1.20 · Terra $2/$12 · Sol Fast $10/$60 · Kimi K3 pressure · Self-optimized GPU story
Summary: On July 30, 2026, OpenAI cut API prices for two of its three GPT-5.6 models: Luna dropped 80% to $0.20/$1.20 per million input/output tokens, and Terra fell 20% to $2/$12. Sol kept its price but gained a Fast mode that costs twice as much for up to 2.5x the speed. OpenAI says part of the savings came from Sol autonomously rewriting its own production GPU code—against a backdrop of Kimi K3 and DeepSeek pricing pressure.
Billing math shifted: ChatGPT Work and Codex subscriptions now burn fewer credits on Luna/Terra workloads
Sol did not get cheaper: Standard stays $5/$30; Fast mode at $10/$60 replaces Priority Processing
Competitive pressure is real: Kimi K3 launched July 16; weights dropped ~July 27; DeepSeek V4 Pro cut 75% in May
Efficiency claims need scrutiny: 20% serving-cost reduction is self-reported with no independent audit
Sticker price misleads: Artificial Analysis puts cost-per-completed-task for Sol and Kimi K3 close—about $0.94 vs $1.04
| Date | Event |
|---|---|
| July 9, 2026 | OpenAI launches GPT-5.6: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens |
| July 16, 2026 | Moonshot AI releases Kimi K3 (2.8T open-weight MoE) at $3/$15 ($0.30 cache hits) |
| ~July 27, 2026 | Kimi K3 open weights become downloadable |
| July 29, 2026 | Engineering blog: Sol in Codex rewrote production GPU kernels (Triton/Gluon) and speculative-decoding draft model |
| July 30, 2026 | Luna/Terra price cuts and Sol Fast mode ship—roughly three weeks after launch |
| July 31, 2026 | Coverage across CNBC, Reuters-sourced reports, IT Home, 36Kr, VentureBeat |
This is not a one-off promo—it is a fast "launch → reveal efficiency story → reprice" sequence rare in OpenAI's history, signaling pricing power eroding faster than labs expected.
| Model | Old (in/out per 1M) | New | Change |
|---|---|---|---|
| GPT-5.6 Luna | $1.00 / $6.00 | $0.20 / $1.20 | -80% |
| GPT-5.6 Terra | $2.50 / $15.00 | $2.00 / $12.00 | -20% |
| GPT-5.6 Sol (Standard) | $5.00 / $30.00 | $5.00 / $30.00 | No change |
| GPT-5.6 Sol (Fast) | N/A | $10.00 / $60.00 | 2x price, up to 2.5x speed |
Subscription prices for ChatGPT Work and Codex are unchanged, but Luna/Terra usage now consumes fewer credits against those plans.
Per OpenAI's engineering post, Sol inside Codex rewrote production GPU kernels in Triton and Gluon, redesigned the speculative-decoding draft model, and tuned KV-cache handling and GPU scheduling. Claimed results: 20% lower end-to-end serving cost and 15%+ better token throughput, checked partly with open-source FpSan.
Luna gets the deepest cut for high-volume agent workloads; Terra gets a modest trim; Sol holds rate and monetizes speed via Fast—a barbell strategy: compete on price at the bottom, capability (and latency) at the top.
Kimi K3 landed ahead of WAIC Shanghai; enterprise buyers are cautious on AI ROI; Sam Altman has called cost "a huge issue." Repricing, efficiency disclosure, and Fast mode in one week reads as reactive competitive positioning as much as pure cost pass-through.
| Model | Vendor | Input $/1M | Output $/1M | Note |
|---|---|---|---|---|
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | Post-cut |
| GPT-5.6 Terra | OpenAI | $2.00 | $12.00 | Post-cut |
| GPT-5.6 Sol | OpenAI | $5.00 | $30.00 | Unchanged |
| Kimi K3 | Moonshot AI | $3.00 ($0.30 cache) | $15.00 | Open weights, 2.8T MoE |
| DeepSeek V4 Pro | DeepSeek | $0.435 ($0.0036 cache) | $0.87 | Permanent 75% cut since May |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | Lightweight tier |
| Claude Sonnet 5 | Anthropic | $3.00 (promo $2.00 thru Aug 31) | $15.00 (promo $10.00) | Matches K3 list rate |
| Gemini 3.5 Flash-Lite | ~$2.80 combined | Lightweight tier | ||
| MAI-Code-1-Flash | Microsoft | $0.75 | $4.50 | GitHub Copilot only |
Luna's combined $1.40/million undercuts Gemini 3.5 Flash-Lite but DeepSeek V4 Flash/Pro remain far cheaper on raw per-token price.
DeepSeek made V4 Pro discounts permanent in May; Moonshot shipped K3 open weights three weeks before this cut. Microsoft is pushing in-house MAI models to reduce OpenAI dependence for everyday coding. Enterprise caution on unproven ROI plus Altman's cost comments suggest frontier pricing power is eroding—with implications for infrastructure spending across the industry.
Baseline pre-cut spend: Export 30-day token use and cost per task by tier
Re-route agents: High-volume tool loops → Luna; hard reasoning → Sol standard or Fast
Run Kimi K3 / DeepSeek controls: Compare list price and task-completion quality
Check subscription credit burn: Confirm Work/Codex plans reflect lower Luna/Terra consumption
Acceptance in a macOS VNC session: Switch models in Codex/OpenClaw, read Gateway logs, handle Keychain prompts—SSH alone often fails
Citable figures: Luna -80% · Terra -20% · Sol Fast 2x price / 2.5x speed · claimed 20% serving savings · Sol vs K3 ~$1.04 vs $0.94 per completed task (Artificial Analysis).
$0.20 per million input tokens and $1.20 per million output tokens, down 80% from $1.00/$6.00 at launch.
No. Standard stayed $5/$30. Fast mode is $10/$60 for up to 2.5x speed with unchanged model intelligence.
On raw per-token pricing, DeepSeek remains cheaper and Kimi K3 list rates sit above Luna's old price. Cost-per-completed-task benchmarks show Sol and K3 much closer than sticker prices suggest.
Kimi K3's July 16 launch, enterprise ROI caution, and Microsoft's cheaper in-house MAI push all landed in the same window—a fast turnaround suggesting reactive positioning.
This repricing is layered strategy—Luna for agent volume, Terra for everyday work, Sol selling speed at a premium—not a uniform markdown. If your daily driver is Windows or Linux but you need Codex, OpenClaw, or ChatGPT Work to switch tiers and benchmark agents quickly, buying a Mac is expensive and SSH cannot click through Keychain or permission dialogs. Renting a VNCMac remote Mac gives you an hourly VNC desktop to validate new pricing in isolation. See Mac Mini M4 plans; background in our GPT-5.6 launch review and Kimi K3 open weights posts.
Sources: OpenAI blogs (July 29–30, 2026), VentureBeat, The Decoder, Model Price Watch, METR reporting. Verify current pricing before production use.