GPT-5.6 Luna price cut [What It Changes]
TL;DR
The GPT-5.6 Luna price cut is OpenAI’s July 30, 2026, 80% drop on its cheapest GPT-5.6 API model, from $1/$6 to $0.20/$1.20 per million input/output tokens, three weeks after the July 9 launch.
Terra, the mid-tier GPT-5.6 model, fell 20% to $2/$12 per million tokens; flagship Sol’s list rate stayed $5/$30, with a separate promotional $4/$20 card on OpenAI’s site.
For a shop that pays by the token—HVAC dispatch notes, agency lead classify, clinic intake extract—the cheapest usable OpenAI model is now one-fifth the launch-week unit cost on that mix.
ChatGPT plan prices did not move; API usage is still billed separately, and Fast mode, Batch, cache hits, and long-context multipliers still change the real invoice.
Key Takeaways
Treat Luna as the default for high-volume, tool-using text jobs; keep Terra or Sol for work where a miss costs more than the token savings.
Recalculate bills from input and output tokens, not from “chat seats,” because API spend is metered and ChatGPT subscriptions do not cover it.
Stack the cut with cache hits (Luna cached input is $0.02 per million tokens) and Batch’s 50% discount on jobs that can wait up to 24 hours.
Do not send medical, credit, or hiring decisions through an unattended Luna call; OpenAI’s usage rules still require a licensed human in those loops.
If you already route documents through a workflow layer, this is a model-id swap, not a rebuild of forms, CRM writes, or human review.
What the GPT-5.6 Luna price cut is
The GPT-5.6 Luna price cut is OpenAI’s July 30, 2026 reduction of GPT-5.6 Luna, its fastest and lowest-cost GPT-5.6 API model, from $1 per million input tokens and $6 per million output tokens down to $0.20 and $1.20.
A two-truck HVAC shop that turns voicemails into job tickets, a ten-person marketing agency that tags inbound forms, and a solo clinic that copies intake PDFs into a chart all hit the same meter: they pay for tokens, not for a seat. When the cheapest model that can call tools and finish a multi-step job drops 80% three weeks after launch, the monthly cost of those features stops looking like a science-fair line item and starts looking like postage.
That is why this hub exists. The vendor move is a frontier-lab rate card. The operational question is whether you can afford to run the same extract-classify-draft loop for every customer instead of only for the ones you have time to type.
If you already map those loops in small-business automation or form-to-CRM tools, the cut changes the model you pick inside the step, not the step itself.
What happened, and when
OpenAI shipped the GPT-5.6 family on July 9, 2026, as a three-rung ladder: Sol as flagship, Terra as the balanced middle, and Luna as the cheap, fast rung, according to Yahoo Finance, which also reports Luna at about 85% of Sol’s quality.
Twenty-one days later, as of July 30, 2026, the company cut Luna 80% and Terra 20% and renamed Priority processing to Fast mode. According to EdTech Innovation Hub, the new API rates took effect on July 30, 2026, with Luna at $0.20 per million input tokens and $1.20 per million output tokens.
Luna now costs $0.20 per million input tokens. That figure matches OpenAI’s live business pricing card and the API pricing table for short-context gpt-5.6-luna.
ChatGPT Work and Codex subscribers did not get a cheaper monthly plan. EdTech Innovation Hub reports that subscription prices and quota budgets stayed put, while Terra and Luna now consume fewer credits inside those products. The same piece says Free and Go users can reach Terra in ChatGPT Work and Codex, while Plus, Pro, Business, and Enterprise can pick Terra or Luna, and that the API change began rolling out through AWS on July 30, 2026.
Scott Rosecrans, OpenAI’s Vice President of Strategic Pursuits, wrote on LinkedIn, as quoted by EdTech Innovation Hub: “As we get more efficient, we are passing the savings on to you!”
| Date | Model | Input $/1M | Output $/1M |
|---|---|---|---|
| 2026-07-09 | GPT-5.6 Luna (launch) | 1.00 | 6.00 |
| 2026-07-09 | GPT-5.6 Terra (launch) | 2.50 | 15.00 |
| 2026-07-09 | GPT-5.6 Sol (list) | 5.00 | 30.00 |
| 2026-07-30 | GPT-5.6 Luna (cut) | 0.20 | 1.20 |
| 2026-07-30 | GPT-5.6 Terra (cut) | 2.00 | 12.00 |
| 2026-07-30 | GPT-5.6 Sol (list, unchanged) | 5.00 | 30.00 |
Sources: Yahoo Finance for launch and cut rates; OpenAI API pricing for the post-cut Luna and Terra standard short-context card.
The rate card after the cut
According to Yahoo Finance, Terra moved 20%, from $2.50/$15 to $2/$12 per million tokens, while Sol held at $5/$30.
OpenAI’s public API pricing page lists short-context Luna at $0.20 input, $0.02 cached input, and $1.20 output per million tokens, Terra at $2.00 / $0.20 / $12.00, and Sol at a promotional $4.00 / $0.40 / $20.00, with that Sol promo available at least through November 21, 2026. Long-context columns on the same page list Luna at $0.40 / $0.04 / $1.80 and Terra at $4.00 / $0.40 / $18.00. Regional processing adds a 10% uplift for eligible models released on or after March 5, 2026.
The GPT-5.6 Luna model card states a 1,050,000-token context window, 128,000 max output tokens, a February 16, 2026 knowledge cutoff, text-and-image input, text output, and a 2× input / 1.5× output multiplier on the full request when prompts exceed 272K input tokens. Cache writes are billed at 1.25× the uncached input rate. Fine-tuning is not supported on Luna.
Terra's output rate fell to $12 per million. That is the post-cut standard output price on both the Yahoo recap and OpenAI’s pricing table.
| Model | Input $/1M | Cached input $/1M | Output $/1M | Context tokens |
|---|---|---|---|---|
| GPT-5.6 Luna | 0.20 | 0.02 | 1.20 | 1,050,000 |
| GPT-5.6 Terra | 2.00 | 0.20 | 12.00 | 1,050,000 |
| GPT-5.6 Sol (promo) | 4.00 | 0.40 | 20.00 | 1,050,000 |
| GPT-5.6 Sol (Yahoo list) | 5.00 | — | 30.00 | 1,050,000 |
Sources: OpenAI API pricing and OpenAI models for Luna, Terra, Sol promo, and context; Yahoo Finance for Sol list $5/$30. The em dash is not a figure; Yahoo did not publish a Sol cached-input list rate.
OpenAI’s model catalog assigns Luna the id gpt-5.6-luna and describes it as the GPT-5.6 model for cost-sensitive, high-volume work, with the same tool set as Sol and Terra: functions, web search, file search, and computer use. The Luna card also lists Chat Completions, Responses, and Batch among supported endpoints.
A token is not a word. OpenAI’s help article on counting tokens gives a rough English rule of 1 token ≈ 4 characters, or about three-quarters of a word, so 100 tokens ≈ 75 words. The on-site tokenizer repeats the same ~4-character / ~¾-word rule of thumb. For a full Responses payload, including tools and images, use the input-token counting API rather than a character guess. OpenAI’s open-source tiktoken library is the local tokenizer for plain text; its README notes that, on average in practice, each token corresponds to about 4 bytes.
Why the floor moved now
The constraint that broke was unit cost, not a missing model. Yahoo Finance frames the 80% Luna cut, three weeks after launch, as a defense of the utility tier while Sol’s list price stayed put.
EdTech Innovation Hub reports OpenAI’s own efficiency story: kernel work helped cut the end-to-end cost of serving the model by 20%, and draft-model experiments increased token-generation efficiency by more than 15%. The same report says OpenAI’s agentic serving stack caps tool output at 10,000 tokens by default and preserves prompt prefixes so earlier computation can be reused.
According to EdTech Innovation Hub, OpenAI claims Luna delivers performance comparable to models that were considered frontier-class a year earlier at approximately six cents per dollar per task and nearly nine times the speed, and that Luna outperforms Fable 5 on professional work measured by Agents’ Last Exam at an estimated cost per task almost 99% lower. Those are vendor claims, not independent lab scores; they belong on the “signal” side only as what OpenAI asserted.
Competitors already publish cheap floors. According to Anthropic’s Claude API pricing, Claude Fable 5 costs $10 per million input tokens and $50 per million output tokens, while Claude Sonnet 5 costs $2 / $10. DeepSeek’s pricing page lists deepseek-v4-pro peak cache-miss input at $1.32 per million tokens and peak output at $3.96, with off-peak rates at half of peak; peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday. Google’s Gemini Developer API pricing lists Gemini 3.8 Flash paid input at $0.75 per million tokens and output at $3.75 through December 31, 2026.
| Provider | Model | Input $/1M | Output $/1M |
|---|---|---|---|
| OpenAI | GPT-5.6 Luna (standard, short) | 0.20 | 1.20 |
| OpenAI | GPT-5.6 Terra (standard, short) | 2.00 | 12.00 |
| Anthropic | Claude Sonnet 5 | 2.00 | 10.00 |
| Anthropic | Claude Fable 5 | 10.00 | 50.00 |
| DeepSeek | V4-Pro (peak, cache miss) | 1.32 | 3.96 |
| Gemini 3.8 Flash (through 2026-12-31) | 0.75 | 3.75 |
Sources: OpenAI API pricing; Anthropic pricing; DeepSeek models and pricing; Gemini Developer API pricing. DeepSeek off-peak is half of peak on that page; this row uses peak cache-miss.
Anthropic’s consumer Claude plans are a different product: Pro is $17 per month with annual prepay ($20 billed monthly), and Enterprise is listed at $20 per seat per month plus usage at API rates. Do not mix those seat prices into an API token budget.
Fast mode, Batch, Flex, and cache
The price cut is not the only lever. OpenAI split speed from intelligence on the same day.
According to EdTech Innovation Hub, Fast mode can deliver speeds up to 2.5 times faster than Standard processing without changing the model’s intelligence, and it costs twice the Standard processing rate. OpenAI’s Fast mode docs confirm the July 30, 2026 rename from Priority processing, the 2.5× Sol speed claim, and that service_tier: "fast" or "priority" both work. For GPT-5.6 Sol, that page prices Fast short-context at $8 per million input tokens and $40 per million output tokens.
Sol Fast mode bills at 2× Standard rates. OpenAI’s Fast mode customer page publishes the doubled Luna Fast short-context card at $0.40 input and $2.40 output per million tokens, versus Luna Standard $0.20 / $1.20.
The same Fast mode page lists Luna Fast latency as 99% > 100 tokens per second for Enterprise SLA customers, with 99.9% uptime on that table, and it warns that Fast mode shares rate limits with Standard and can fall back if traffic jumps more than 50% TPM in 15 minutes after you are already at 1 million TPM.
Batch API is the overnight lane: 50% lower cost than synchronous APIs, a separate higher rate-limit pool, and a 24-hour completion window. Flex processing is priced at Batch rates for Responses or Chat Completions in exchange for slower replies and possible 429 Resource Unavailable (you are not charged when that error hits). Flex is in beta with limited model availability.
Prompt caching is on by default for supported models. For GPT-5.6 and later, cache writes cost 1.25× uncached input and cache reads cost 0.1×; the minimum cacheable prompt length is 1,024 tokens. Luna’s cached input rate of $0.02 per million tokens on the pricing page is that 0.1× read.
Web search, if you turn it on, is a separate $10.00 per 1,000 calls on OpenAI’s pricing page, plus search-content tokens at model rates for several tools.
Rate limits still cap a cheap model
A cheaper token is useless if the API returns 429. OpenAI’s rate limit guide meters RPM, RPD, TPM, and TPD at the organization and project level, not the user. Usage tiers step from Free ($100 / month) through Tier 1 ($5 paid, $100 / month) up to Tier 5 ($1,000 paid, $200,000 / month).
The Luna model card publishes per-tier Luna caps:
| Tier | RPM | TPM | Batch queue tokens |
|---|---|---|---|
| Free | — | — | — |
| 1 | 500 | 500,000 | 5,000,000 |
| 2 | 5,000 | 2,000,000 | 20,000,000 |
| 3 | 5,000 | 4,000,000 | 40,000,000 |
| 4 | 10,000 | 10,000,000 | 1,000,000,000 |
| 5 | 30,000 | 180,000,000 | 15,000,000,000 |
Source: GPT-5.6 Luna model. Free is listed as not supported. Em dashes are not figures.
Ramp-rate advice on the rate limit page: once traffic reaches 1 million input TPM, increase it by no more than 50% every 15 minutes.
USTA analysis: unit cost after the cut
USTA analysis (inputs only from the Yahoo launch card and the post-cut OpenAI/Yahoo rates above; the token counts are a calculator quantity, not a claimed “average SMB usage”).
Take one mixed unit of 1,000,000 input tokens plus 1,000,000 output tokens on short-context Standard processing.
| Mix (1M in + 1M out) | Launch $/unit | After Jul 30 $/unit | $ saved | Share of launch $ |
|---|---|---|---|---|
| Luna | 7.00 | 1.40 | 5.60 | 0.20 |
| Terra | 17.50 | 14.00 | 3.50 | 0.80 |
| Sol list | 35.00 | 35.00 | 0.00 | 1.00 |
Inputs: Luna launch $1+$6, Terra launch $2.50+$15, Sol list $5+$30 from Yahoo Finance; Luna after $0.20+$1.20 and Terra after $2+$12 from Yahoo and OpenAI pricing. Arithmetic: 1.00+6.00=7.00; 0.20+1.20=1.40; 7.00−1.40=5.60; 1.40/7.00=0.20. Terra: 2.50+15.00=17.50; 2.00+12.00=14.00; 14.00/17.50=0.80. Sol list is unchanged on Yahoo’s card.
On that mixed unit, Luna after the cut costs $1.40 versus Terra’s $14.00, a 10× gap ($14.00 / $1.40). Luna versus Sol list is $35.00 / $1.40 ≈ 25×. Batch at 50% off would price the same Luna unit at $0.70 if the job can wait 24 hours, using the Batch discount on those post-cut rates.
If 80% of the input tokens are cache hits at Luna’s $0.02 cached rate, the input line on that 1M+1M unit becomes (0.80 × $0.02) + (0.20 × $0.20) = $0.016 + $0.040 = $0.056, plus $1.20 output, for $1.256 instead of $1.40. That cache share is an assumption you must measure in your own prompt-caching dashboard; it is not a published industry average.
The U.S. Small Business Administration tells owners to put money-in and money-out on a balance sheet and to run a simple cost-benefit check before a recurring spend. Token APIs belong on that sheet as a variable operating cost, next to payroll software, not as a one-time tool purchase.
What the cut does not change
ChatGPT and the API are separate bills. OpenAI’s business pricing FAQ states that OpenAI APIs are billed separately from ChatGPT Plus, Business, Enterprise, and Edu, and that Playground usage is billed at the same per-token rates. The ChatGPT pricing page lists GPT-5.6 Luna on Free, Go, Plus, and Pro for in-product chat; that is not an API allowance.
Data handling did not get cheaper by getting looser. OpenAI’s data controls still say that, as of March 1, 2023, API data is not used to train models unless you opt in, and that abuse-monitoring logs are kept up to 30 days by default. Zero Data Retention is an approval path, not the default for a new API key.
OpenAI usage policies (effective October 29, 2025) still bar unattended high-stakes automation in medical, legal, employment, credit, insurance, and essential government services without human review, and they still bar tailored licensed advice without a licensed professional. A cheaper Luna call does not make an unlicensed medical chatbot compliant.
NIST’s AI Risk Management Framework remains voluntary guidance: AI RMF 1.0 shipped January 26, 2023, the generative-AI profile (NIST-AI-600-1) shipped July 26, 2024, and a concept note for a critical-infrastructure profile went out April 7, 2026. Price is not a substitute for that risk work.
If you buy through AWS, OpenAI’s Bedrock guide says Amazon Bedrock bills OpenAI models through AWS, feature coverage is a subset (no computer use, no hosted file search, on-demand inference only as of the July 13, 2026 table), and GPT-5.6 Sol, Terra, and Luna have a 1,050,000-token context window there. Confirm current dollars on Amazon Bedrock pricing before you assume the OpenAI.com card.
How a small team should route work
Put Luna on the jobs that repeat all day: classify the inbound form, draft the job note, summarize the call, extract the invoice fields. Keep Terra when the draft will be sent to a customer with little editing. Keep Sol, and Fast mode, for the few requests where a person is waiting and a wrong answer is expensive.
Teams already routing documents through US Tech Automations workflows will plug this in as a model swap, not a rebuild: same form trigger, same CRM write, new gpt-5.6-luna id on the generate step.
A clinic that files intake through US Tech Automations can keep the same data-extraction step and the same staff review, then point only the model field at Luna for the first-pass parse.
Agencies that classify leads through US Tech Automations can send overnight batches to Luna instead of paying Terra on every record, then keep a person on the shortlist—the same pattern as marketing-agency workflow tools and executive-assistant task automation.
Set a monthly budget in the API limits dashboard described on OpenAI’s pricing FAQ, because enforcement can lag and you still own overage. For customer-facing support agents, log model id, token counts, and cache-hit share per conversation so a Fast-mode spike is visible.
Home base for that wiring is US Tech Automations.
Signal vs Speculation
Demonstrated (sourced): As of July 30, 2026, OpenAI cut Luna 80% to $0.20/$1.20 and Terra 20% to $2/$12, left Sol’s list at $5/$30, renamed Priority to Fast mode at 2× Standard, and published Luna cached input at $0.02. ChatGPT subscription prices did not change. API data-training default, 30-day abuse logs, usage-policy high-stakes rules, Batch’s 50% discount, and Luna’s 1,050,000-token window are documented on OpenAI properties cited above. Anthropic, DeepSeek, and Google publish the competitor floors in the table.
Our read (12–36 months, not a sourced forecast): If X holds—if utility-tier models keep getting repriced inside a month of launch—then small firms will treat “which model” as a weekly config, not a contract. The likely failure mode is not the $1.40 mixed unit; it is sending every customer-facing email through Luna Fast, missing cache, adding web search at $10 per 1,000 calls, and wondering why the invoice did not fall 80%. Another failure mode is putting medical or credit language on Luna because the tokens are cheap. We expect Sol to stay the scarce, pricey rung while Luna absorbs classify-and-draft volume, and we expect Batch/Flex to become the default for anything that can wait overnight. We do not know whether OpenAI will cut Luna again, raise Sol, or let the November 21, 2026 Sol promo expire; those are open.
Glossary
GPT-5.6 Luna price cut: OpenAI’s July 30, 2026, 80% reduction of Luna API rates to $0.20/$1.20 per million input/output tokens.
Token: The billed unit of text (and some other modalities); roughly ~4 characters or ~¾ of an English word, not a word count.
Cached input: Reused prompt-prefix tokens billed at the model’s cached rate (Luna $0.02 per million on the standard short-context card).
Fast mode: Pay-as-you-go low-latency tier, formerly Priority processing; 2× Standard rates; Sol claimed up to 2.5× faster.
Batch API: Asynchronous 24-hour job queue at 50% off Standard token rates.
Flex processing: Beta service tier priced like Batch, with slower replies and possible unpaid resource-unavailable errors.
GPT-5.6 Terra / Sol: Mid-tier ($2/$12) and flagship (list $5/$30, promo $4/$20) siblings in the same family.
TPM / RPM: Tokens per minute and requests per minute, the caps that still bind after a price cut.
FAQ
What is the GPT-5.6 Luna price cut?
It is OpenAI’s July 30, 2026, 80% reduction of GPT-5.6 Luna API prices from $1/$6 to $0.20/$1.20 per million input/output tokens, as reported by Yahoo Finance and EdTech Innovation Hub. Terra fell 20% the same day; Sol’s list rate did not.
Did GPT-5.6 Sol get cheaper too?
No on the list card Yahoo published at $5/$30. OpenAI’s API pricing page separately lists promotional Sol at $4/$20 per million tokens, available at least through November 21, 2026. Fast mode on Sol is 2× Standard, not a discount.
How much cheaper is Luna than Terra after the cut?
On a 1 million input plus 1 million output mixed unit, Luna is $1.40 and Terra is $14.00 after July 30, using the post-cut rates. That 10× gap is the USTA analysis arithmetic above, not a vendor slogan.
Does a ChatGPT plan include the API discount?
No. OpenAI’s business pricing FAQ says APIs are billed separately from ChatGPT Plus, Business, Enterprise, and Edu. ChatGPT pricing covers in-product Luna chat, not API tokens.
Should a small clinic put medical advice on Luna because it is cheap?
No. OpenAI usage policies still require appropriate involvement by a licensed professional for tailored medical advice and human review for high-stakes medical automation. Use Luna to draft or extract, then have a clinician sign off.
What is Fast mode versus Batch?
Fast mode is the low-latency, 2×-price synchronous path described in OpenAI’s Fast mode guide. Batch is the 50%-off, 24-hour asynchronous path. Do not run ETL on Fast mode; the ramp-rate notes say those jobs often do not need it.
How do I keep a Luna bill from surprising me?
Count tokens with the counting API or tiktoken before you scale, set a monthly budget in API limits, watch cache-hit rate, and keep web search and Fast mode off unless the step needs them. Prompt caching only pays off if the prefix actually matches.
What to do this week
Read the live OpenAI pricing card, not last month’s spreadsheet. Point high-volume classify and extract jobs at gpt-5.6-luna, send overnight jobs through Batch, and leave Sol Fast for the queue a person is staring at.
If those jobs already sit in a workflow, swap the model on the generate step and keep the rest of the path—trigger, tools, review—unchanged.
About the Author

Helping businesses leverage automation for operational efficiency.
Related Articles
See how AI agents fit your team
US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.
View pricing & plans