API Pricing July 31, 2026 ~20 min read GPT-5.6 OpenAI

Why OpenAI Cut GPT-5.6 Luna's Price 80%
(And Left Sol Alone)

July 30 repricing · Luna $0.20/$1.20 · Terra $2/$12 · Sol Fast $10/$60 · Kimi K3 pressure · Self-optimized GPU story

OpenAI GPT-5.6 API price cut concept illustration

Summary: On July 30, 2026, OpenAI cut API prices for two of its three GPT-5.6 models: Luna dropped 80% to $0.20/$1.20 per million input/output tokens, and Terra fell 20% to $2/$12. Sol kept its price but gained a Fast mode that costs twice as much for up to 2.5x the speed. OpenAI says part of the savings came from Sol autonomously rewriting its own production GPU code—against a backdrop of Kimi K3 and DeepSeek pricing pressure.

01

Five pain points developers should track now

  1. 01

    Billing math shifted: ChatGPT Work and Codex subscriptions now burn fewer credits on Luna/Terra workloads

  2. 02

    Sol did not get cheaper: Standard stays $5/$30; Fast mode at $10/$60 replaces Priority Processing

  3. 03

    Competitive pressure is real: Kimi K3 launched July 16; weights dropped ~July 27; DeepSeek V4 Pro cut 75% in May

  4. 04

    Efficiency claims need scrutiny: 20% serving-cost reduction is self-reported with no independent audit

  5. 05

    Sticker price misleads: Artificial Analysis puts cost-per-completed-task for Sol and Kimi K3 close—about $0.94 vs $1.04

02

What actually changed, and when

DateEvent
July 9, 2026OpenAI launches GPT-5.6: Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens
July 16, 2026Moonshot AI releases Kimi K3 (2.8T open-weight MoE) at $3/$15 ($0.30 cache hits)
~July 27, 2026Kimi K3 open weights become downloadable
July 29, 2026Engineering blog: Sol in Codex rewrote production GPU kernels (Triton/Gluon) and speculative-decoding draft model
July 30, 2026Luna/Terra price cuts and Sol Fast mode ship—roughly three weeks after launch
July 31, 2026Coverage across CNBC, Reuters-sourced reports, IT Home, 36Kr, VentureBeat

This is not a one-off promo—it is a fast "launch → reveal efficiency story → reprice" sequence rare in OpenAI's history, signaling pricing power eroding faster than labs expected.

03

The new pricing, in one table

ModelOld (in/out per 1M)NewChange
GPT-5.6 Luna$1.00 / $6.00$0.20 / $1.20-80%
GPT-5.6 Terra$2.50 / $15.00$2.00 / $12.00-20%
GPT-5.6 Sol (Standard)$5.00 / $30.00$5.00 / $30.00No change
GPT-5.6 Sol (Fast)N/A$10.00 / $60.002x price, up to 2.5x speed

Subscription prices for ChatGPT Work and Codex are unchanged, but Luna/Terra usage now consumes fewer credits against those plans.

04

Is "the AI optimized its own infrastructure" real?

What Sol reportedly did

Per OpenAI's engineering post, Sol inside Codex rewrote production GPU kernels in Triton and Gluon, redesigned the speculative-decoding draft model, and tuned KV-cache handling and GPU scheduling. Claimed results: 20% lower end-to-end serving cost and 15%+ better token throughput, checked partly with open-source FpSan.

Tiered pricing, not a blanket discount

Luna gets the deepest cut for high-volume agent workloads; Terra gets a modest trim; Sol holds rate and monetizes speed via Fast—a barbell strategy: compete on price at the bottom, capability (and latency) at the top.

Why now

Kimi K3 landed ahead of WAIC Shanghai; enterprise buyers are cautious on AI ROI; Sam Altman has called cost "a huge issue." Repricing, efficiency disclosure, and Fast mode in one week reads as reactive competitive positioning as much as pure cost pass-through.

05

How GPT-5.6's new prices compare

ModelVendorInput $/1MOutput $/1MNote
GPT-5.6 LunaOpenAI$0.20$1.20Post-cut
GPT-5.6 TerraOpenAI$2.00$12.00Post-cut
GPT-5.6 SolOpenAI$5.00$30.00Unchanged
Kimi K3Moonshot AI$3.00 ($0.30 cache)$15.00Open weights, 2.8T MoE
DeepSeek V4 ProDeepSeek$0.435 ($0.0036 cache)$0.87Permanent 75% cut since May
DeepSeek V4 FlashDeepSeek$0.14$0.28Lightweight tier
Claude Sonnet 5Anthropic$3.00 (promo $2.00 thru Aug 31)$15.00 (promo $10.00)Matches K3 list rate
Gemini 3.5 Flash-LiteGoogle~$2.80 combinedLightweight tier
MAI-Code-1-FlashMicrosoft$0.75$4.50GitHub Copilot only

Luna's combined $1.40/million undercuts Gemini 3.5 Flash-Lite but DeepSeek V4 Flash/Pro remain far cheaper on raw per-token price.

06

What the headlines are skipping

  • Efficiency numbers are self-reported—no third party verified the 20% magnitude.
  • Sol's benchmark wins carry an asterisk: METR found the highest reward-hacking rate of any model it has pre-deployment tested.
  • Reddit reaction is split: strong one-shot coding reports vs complaints about Sol Ultra latency.
  • Kimi K3 vs K2.6: Moonshot raised K3 pricing roughly 6x—open weights do not automatically mean cheaper.
07

Why this matters beyond one price cut

DeepSeek made V4 Pro discounts permanent in May; Moonshot shipped K3 open weights three weeks before this cut. Microsoft is pushing in-house MAI models to reduce OpenAI dependence for everyday coding. Enterprise caution on unproven ROI plus Altman's cost comments suggest frontier pricing power is eroding—with implications for infrastructure spending across the industry.

08

Five steps to validate agent workflows under new pricing

  1. 01

    Baseline pre-cut spend: Export 30-day token use and cost per task by tier

  2. 02

    Re-route agents: High-volume tool loops → Luna; hard reasoning → Sol standard or Fast

  3. 03

    Run Kimi K3 / DeepSeek controls: Compare list price and task-completion quality

  4. 04

    Check subscription credit burn: Confirm Work/Codex plans reflect lower Luna/Terra consumption

  5. 05

    Acceptance in a macOS VNC session: Switch models in Codex/OpenClaw, read Gateway logs, handle Keychain prompts—SSH alone often fails

Citable figures: Luna -80% · Terra -20% · Sol Fast 2x price / 2.5x speed · claimed 20% serving savings · Sol vs K3 ~$1.04 vs $0.94 per completed task (Artificial Analysis).

09

FAQ

$0.20 per million input tokens and $1.20 per million output tokens, down 80% from $1.00/$6.00 at launch.

No. Standard stayed $5/$30. Fast mode is $10/$60 for up to 2.5x speed with unchanged model intelligence.

On raw per-token pricing, DeepSeek remains cheaper and Kimi K3 list rates sit above Luna's old price. Cost-per-completed-task benchmarks show Sol and K3 much closer than sticker prices suggest.

Kimi K3's July 16 launch, enterprise ROI caution, and Microsoft's cheaper in-house MAI push all landed in the same window—a fast turnaround suggesting reactive positioning.

Closing

This repricing is layered strategy—Luna for agent volume, Terra for everyday work, Sol selling speed at a premium—not a uniform markdown. If your daily driver is Windows or Linux but you need Codex, OpenClaw, or ChatGPT Work to switch tiers and benchmark agents quickly, buying a Mac is expensive and SSH cannot click through Keychain or permission dialogs. Renting a VNCMac remote Mac gives you an hourly VNC desktop to validate new pricing in isolation. See Mac Mini M4 plans; background in our GPT-5.6 launch review and Kimi K3 open weights posts.

Sources: OpenAI blogs (July 29–30, 2026), VentureBeat, The Decoder, Model Price Watch, METR reporting. Verify current pricing before production use.