Skip to content
Frontier Tech

Grok 4.6 [What It Changes]

Sep 2, 2026

TL;DR

  • Grok 4.6 is SpaceXAI's flagship model for coding agents and long-running multi-step work, announced as of August 12, 2026, at $2 per million input tokens and $6 per million output tokens.

  • It shipped the same day in Cursor and Grok Build, with a first-week 2x usage promo that was not a standing discount.

  • Artificial Analysis scores Grok 4.6 (high) at 61 on the Intelligence Index, matching SpaceXAI's launch composite.

  • A 2-truck shop, 10-person agency, or solo clinic should treat this as a model swap with spend caps and human review, not as a new stack.

Key Takeaways

  • List price is the operational headline: $2/$6 per million tokens on short context, with long-context and fast/priority paths at 2x.

  • The capability pitch is persistence: research, codebase work, and visual first passes that keep going across many steps.

  • Benchmarks are mixed versus GPT-5.6 Sol Max; Grok 4.6 leads some knowledge-work scores and trails on DeepSWE and Terminal-Bench.

  • Distribution is already broad: Cursor, Grok Build, the SpaceXAI API, OpenRouter, and later cloud and Copilot listings.

  • Small teams win by scoping one workflow, comparing diffs on real jobs, and keeping personal data out of prompts.

What Grok 4.6 means for a small shop

Grok 4.6 is SpaceXAI's flagship large language model for agentic coding and long-running tasks: it is built to stay with a multi-step job such as researching an unfamiliar market, working through a large codebase, or refining a design across several rounds of feedback, and it is sold through the SpaceXAI API at $2.00 per million input tokens and $6.00 per million output tokens.

A 2-truck HVAC company does not need a research lab. It needs an agent that can hold a messy job: read last week's notes, draft a quote from a nameplate photo, and not drop the thread after the third follow-up. A 10-person marketing agency needs the same persistence for a campaign teardown, a visual first pass, and a CRM update without paying frontier output rates that now sit at $20 to $50 per million tokens at OpenAI and Anthropic. A solo-run clinic needs document summaries and form-to-record drafts that a human still signs.

If you already move forms into a CRM or run agency dispatch, Grok 4.6 is a model you can point at those same steps. Teams already routing intake through US Tech Automations workflows can treat it as a model swap, not a rebuild. The constraint that broke is that a frontier-tier coding agent is now cheap enough, and persistent enough, that a small team can leave it on a real job.

The SBA's manage-your-business guide already tells owners to weigh AI risks and benefits next to bookkeeping, payroll, and cyber hygiene. This release is one more line on that same cost-benefit sheet.

What shipped on August 12, 2026

On August 12, 2026, SpaceXAI published "Introducing Grok 4.6". The post says Grok 4.6 builds on Grok 4.5 with a focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps: research, analysis, codebase work, or turning an idea into a polished application.

The same day, Cursor's research post repeated that launch copy and said Grok 4.6 was available immediately in Cursor and Grok Build. According to Cursor's Grok 4.6 post, SpaceXAI offered 2x included usage inside Cursor and Grok Build for the first week. That window was a launch promo; do not budget as if it is still on.

SpaceXAI's news index later listed Grok 4.6 in GitHub Copilot (August 19, 2026), on Amazon Bedrock (August 19, 2026), on Google's Gemini Enterprise Agent Platform (August 21, 2026), and on Microsoft Foundry (August 26, 2026). Those are availability notes, not extra benchmark claims.

Grok Build now markets itself as a coding agent powered by Grok 4.6, with plan mode, skills, plugins, MCP servers, and subagents. Cursor's homepage shows Grok 4.6 as a selectable agent model. Neither surface is a reason to skip a human review of the diff.

How the model is trained, in plain language

The launch post says Grok 4.6 got a longer supplemental training run than Grok 4.5, using curated model-generated data for reasoning and technical concepts, high-quality engineering data, and an improved optimizer. That run fed supervised fine-tuning and reinforcement learning.

The SFT step used Grok 4.5 to regenerate trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, then filtered bad traces with model-based checks. Reinforcement learning then covered knowledge work, general coding, and environments for kernel optimization, web development, and computer-aided design.

What you feel in a shop is three behavior shifts SpaceXAI and Cursor both describe: the model stays on a long trajectory; it lays down structure and visual language for an app in one pass more often than Grok 4.5; and on longer runs it checks its own work before moving on. That last habit reduces obvious errors. It does not replace a person reading the pull request.

SpaceXAI's model docs call Grok 4.6 the flagship for code and everything else, with agentic tool calling, configurable reasoning, and a 500k-token context window. The knowledge cut-off is February 1, 2026. Without web or X search tools enabled, it has no knowledge of events after that date. Any job that needs this week's vendor price or this week's regulation must turn search on and pay tool-call fees.

Pricing, context, and what "half the price" actually means

According to SpaceXAI's API pricing docs, grok-4.6 lists $2.00 input and $6.00 output per 1M tokens at short context, with cached input at $0.50. Once a prompt hits the long-context threshold of 200k tokens, the same model bills $4.00 / $1.00 cached / $12.00 output. Grok 4.6 lists $2 input and $6 output per million tokens.

The launch post also says a fast variant is twice the price. Cursor's models page posts Grok 4.6 (Fast) at $4 input and $12 output, and SpaceXAI bills priority processing at a 2x premium when the response confirms the priority tier.

Compare that with the other frontier stickers small teams actually see. According to OpenAI's API pricing, gpt-5.6-sol lists $4.00 per 1M input tokens and $20.00 per 1M output tokens at short context (promotional pricing stated as available at least through November 21, 2026). According to Anthropic's platform pricing, Claude Fable 5.1 lists $10 per million input tokens and $50 per million output tokens. Claude Sonnet 5 on the same table is $2 / $10.

So "half the price of other frontier models" is true for Grok 4.6 input versus GPT-5.6 Sol's $4.00. It is not true for every rival and every token type. Output versus Sol is $6 vs $20. Output versus Fable 5.1 is $6 vs $50. Output versus Sonnet 5 is $6 vs $10. Always compare the pair you would actually run.

According to OpenRouter's Grok 4.6 listing, list pricing is $2 / $6 per 1M tokens with a 500K context window. OpenRouter also posts Amazon Bedrock (US) at $2.20 / $6.60, and SpaceXAI Priority at $4 / $12. Weighted average input paid on that page sat under list because of cache hits; do not budget the discounted average as a guaranteed rate.

Server-side tools on the SpaceXAI pricing page add invocation fees on top of tokens: web search and X search at $5 per 1k calls, code execution at $5 per 1k, file-attachment search at $10 per 1k. Long-running agents that browse and execute will spend on tools even when the model sticker looks cheap.

ModelInput / 1MOutput / 1MContext / notes
grok-4.6 short$2.00$6.00500k; cache $0.50
grok-4.6 long (≥200k)$4.00$12.00same 500k window
grok-4.6 Fast / priority$4.00$12.002x standard
gpt-5.6-sol short$4.00$20.00OpenAI list
Claude Fable 5.1$10.00$50.00Anthropic list
Claude Sonnet 5$2.00$10.00Anthropic list

Sources: SpaceXAI pricing; OpenAI API pricing; Anthropic platform pricing; Cursor models.

Inside Cursor, Grok 4.6 sits in the Cursor Models pool with Grok 4.5 and Composer 2.5, separate from the Other Models pool that bills third-party APIs. Cursor plan prices are $20/mo Pro, $60/mo Pro Plus, and $200/mo Ultra; Teams Standard is $40/user/mo and Premium is $120/user/mo.

PlanPriceCursor ModelsOther Models
Pro$20/moIncludedIncluded
Pro Plus$60/moIncludedIncluded
Ultra$200/moIncludedIncluded
Teams Standard$40/user/moIncludedIncluded
Teams Premium$120/user/mo5x Standard Agent limitsIncluded

Source: Cursor models and pricing.

Benchmarks you can check

According to SpaceXAI's Grok 4.6 announcement, Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol Max and sitting one point behind Fable 5 Max at 62. Grok 4.6 High scores 61 on the AA Intelligence Index. The same table reports GDPVal-AA v2 at 1753 for Grok 4.6 High versus 1526 for Grok 4.5 High and 1728 for GPT-5.6 Sol Max. That 1753 figure is a knowledge-work eval, not an LMSYS Chatbot Arena ELO.

EvaluationGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
AA-Briefcase1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%

Source: SpaceXAI, Introducing Grok 4.6. Competitor figures on that page are drawn from developers' system cards or public leaderboards.

The generation-over-generation jumps are the cleanest signal. DeepSWE v1.1 moves from 54% to 65.9%. Terminal-Bench v3.0 moves from 15.7% to 26%. APEX-Agents moves from 47.1% to 57.5%. GDPVal-AA v2 moves from 1526 to 1753. Those are real gains at the same $2/$6 list price Grok 4.5 already had.

Versus GPT-5.6 Sol Max the picture is mixed. Grok 4.6 leads on GDPVal-AA v2, CursorBench, FrontierCode, APEX-Agents, AA-Briefcase, and Harvey LAB. It trails on DeepSWE (65.9% vs 73%) and Terminal-Bench (26% vs 34.6%). If your work is CLI-heavy DevOps agents, Sol still looks stronger on the published table. If your work is in-IDE coding plus knowledge work, Grok 4.6 is in the same band at a lower output sticker.

According to Artificial Analysis, Grok 4.6 (high) scores 61 on the Intelligence Index and ranks 9 of 196 models in its class. The same page lists $2.00 / $6.00 pricing, $0.94 cost per Intelligence Index task, 54.1 output tokens per second, a 500k context window, and a release date of August 12, 2026. Speed is below the class median; time-to-first-token on SpaceXAI's API is reported at 39.79s because reasoning time is included. Independent scoring matched the launch composite of 61.

Launch-week recaps at Basenor and Mool Studio repeated the $2/$6 price and the eval table. Treat SpaceXAI and Artificial Analysis as the carriers of the figures.

Where a small team can actually run it

You do not need a new vendor relationship to try Grok 4.6. SpaceXAI shows model="grok-4.6" in Python, TypeScript, OpenAI-SDK, and curl examples against https://api.x.ai/v1. The API landing page repeats $2 / $6, 500k context, prepaid credits, enterprise invoicing, SOC 2 Type II, HIPAA eligibility with a BAA, data residency, and Zero Data Retention as a console toggle.

Grok remains the consumer assistant on web, iOS, and Android. That is a different surface from the coding agent. Do not confuse a grok.com chat quota with API spend.

Cursor is the IDE path most agencies already have. GitHub Copilot is another editor path SpaceXAI listed on its news index; confirm the model picker in your plan before you assume Grok 4.6 is in the included set.

OpenRouter, Vercel, and Cloudflare were named as launch partners in both the SpaceXAI and Cursor posts. If you already route models through a marketplace, the swap is a slug change plus a spend cap.

For agencies comparing Plutio-class ops tools or marketing-agency automation, the useful test is one real client artifact: a proposal draft, a landing-page first pass, or a research brief. For shops already automating executive-assistant style tasks, the useful test is a long thread that used to die.

A 10-person agency that already automates client intake can plug the same model into US Tech Automations agentic workflows for research and draft loops. That is a routing change, not a new product category.

Honest limits

Grok 4.6 is not a license to skip review. SpaceXAI's Acceptable Use Policy (effective August 14, 2026) bars high-stakes automated decisions that affect a person's safety, legal or material rights, or well-being, including financial credit, educational, employment, housing, insurance, legal, and medical decisions.

The SpaceXAI privacy policy (effective August 24, 2026) asks users not to put personal information in prompts and says the consumer policy does not govern API processing on behalf of business customers. A solo clinic that pastes a chart note into grok.com is not on the enterprise BAA track.

According to the FTC's Protecting Personal Information guide, a sound data security plan is built on 5 key principles: take stock, scale down, lock it, pitch it, and plan ahead. Sending customer files to a coding agent is a "take stock" event.

CISA's Secure Our World still reduces the same risk to 4 daily habits: recognize phishing, use strong passwords, turn on MFA, and update software. An agent with repo and inbox plugins expands the blast radius of a stolen laptop.

According to NIST, the AI RMF 1.0 was released on January 26, 2023 for voluntary use. The Generative AI Profile (NIST AI 600-1) followed in July 2024 and maps extra GAI risks onto Govern / Map / Measure / Manage. Neither document certifies Grok 4.6.

Technical limits sit next to the policy limits. Knowledge is frozen at February 1, 2026 unless search tools are on. Long prompts flip to 2x token rates at 200k. Independent speed is below average. Terminal-Bench and DeepSWE still trail Sol Max. Self-testing on long trajectories is a reported behavior, not a guarantee that the test suite is the right suite.

Implementation path for small teams

Name one workflow with an acceptance test. "Refactor auth" is not a test. "All login tests green and the session cookie gone" is a test. Run the same prompt on your current default and on Grok 4.6. Save the diffs, the number of follow-ups, and the token bill. Cap iterations, because persistence raises quality and also raises output tokens. Keep search off unless the job needs a date after February 1, 2026. Keep secrets, patient files, and raw card data out of the prompt. Put a human on the merge button.

The state of small-business automation is still about routing work you already do. If the workflow is already on https://ustechautomations.com/, swap the model string, watch the bill, and keep the rest of the path.

USTA analysis: token-cost delta on a 1:1 million-token pair

USTA analysis (derived only from the list prices cited above): for a simplified job that burns 1 million input tokens and 1 million output tokens at short-context list rates, Grok 4.6 costs $2 + $6 = $8. GPT-5.6 Sol costs $4 + $20 = $24. Claude Fable 5.1 costs $10 + $50 = $60. Claude Sonnet 5 costs $2 + $10 = $12. Grok 4.6 Fast / long-context / priority on the 2x path costs $4 + $12 = $16.

Inputs: Grok 4.6 short from SpaceXAI pricing; Sol short from OpenAI; Fable 5.1 and Sonnet 5 from Anthropic; 2x fast/priority from SpaceXAI and Cursor.

Arithmetic: Sol minus Grok = $16 more per that 1:1 pair (Sol is 3x the Grok blended sticker). Fable minus Grok = $52. Sonnet minus Grok = $4. Grok 4.5 short is also $2/$6, so the price delta versus the previous Grok flagship is $0; you are buying the eval jumps (GDPVal-AA 1753 vs 1526, AA Index 61 vs 56) at the same list. Real jobs are not 1:1 in/out and will add reasoning tokens, cache hits, and tool calls. Use the $8 / $24 / $60 row as a checkable floor, then meter your own traces.

GPT-5.6 Sol short-context output is $20.00 per 1M tokens. That output gap is why agent loops, which emit a lot of code, feel the Grok 4.6 sticker first.

Signal vs Speculation

Sourced fact: SpaceXAI released Grok 4.6 on August 12, 2026 at $2/$6, with 500k context, a 2x fast path, and same-day Cursor and Grok Build distribution, including a one-week 2x usage promo.

Sourced fact: SpaceXAI's published evals show a tie with GPT-5.6 Sol Max at 61 on the AA Intelligence Index, a GDPVal-AA v2 score of 1753, and mixed coding results that trail Sol on DeepSWE and Terminal-Bench.

Sourced fact: Artificial Analysis independently scores Grok 4.6 (high) at 61, ranks it 9 of 196, and measures 54.1 tok/s and $0.94 per Intelligence Index task.

Sourced fact: OpenAI and Anthropic still post higher output stickers on their frontier SKUs; Cursor bills Grok 4.6 from the first-party pool; OpenRouter lists Bedrock at a 10% premium.

Sourced fact: SpaceXAI's AUP forbids high-stakes automated decisions; NIST's AI RMF and GAI profile remain voluntary checklists; FTC and CISA still describe basic data-security and account-hygiene steps.

Our read: If the $2/$6 list holds and Cursor keeps Grok 4.6 in the included pool, small teams will default to it for long coding and research loops within 12 months because the output token math is the bill they feel. If independent speed stays near 54 tok/s and Terminal-Bench stays in the mid-20s, DevOps-heavy shops will keep a second model for CLI agents. If rivals cut output prices toward $6, the "half the price" pitch fades and Grok 4.6 becomes one more interchangeable frontier SKU. Over 12–36 months, the durable change for SMBs is not the version number. It is that multi-hour agent jobs are now cheap enough to leave on, which makes prompt scoping, spend caps, and human merge control the actual operating system. That is also where a workflow layer earns its keep: the model will churn; the routing, logs, and approval steps should not.

Glossary

  • Grok 4.6: SpaceXAI flagship model for coding and long-running agents, API id grok-4.6, $2/$6 per 1M tokens, 500k context.

  • Grok Build: SpaceXAI's coding-agent harness that launched with Grok 4.6 as its default model.

  • Artificial Analysis Intelligence Index: Composite of nine evals; SpaceXAI and AA both publish 61 for Grok 4.6 (high).

  • GDPVal-AA v2: Knowledge-work eval on which SpaceXAI reports 1753 for Grok 4.6 High.

  • Long-context pricing: SpaceXAI bills the entire request at the higher $4/$12 rates once the prompt reaches 200k tokens.

  • Priority / Fast variant: 2x token prices for higher scheduling priority or Cursor's Fast mode.

  • Zero Data Retention (ZDR): Console setting so API prompts and outputs are not persisted.

  • AI RMF: NIST's voluntary AI Risk Management Framework (AI 100-1), with a generative-AI profile in AI 600-1.

FAQ

What is Grok 4.6?

Grok 4.6 is SpaceXAI's flagship coding and long-running-agent model, released August 12, 2026. It is meant to stay with multi-step jobs such as research, large-codebase work, and visual product first passes, and it is available in Cursor, Grok Build, and the SpaceXAI API.

How much does Grok 4.6 cost?

Short-context API list is $2 per million input tokens and $6 per million output tokens, with $0.50 cached input, per SpaceXAI's pricing docs. Long context (≥200k) and Fast/priority paths are $4 / $12. Tool calls such as web search add $5 per 1k invocations.

Is Grok 4.6 better than GPT-5.6 Sol?

Not on every eval. SpaceXAI's table ties Sol at 61 on the AA Intelligence Index and leads several knowledge-work scores, but trails on DeepSWE (65.9% vs 73%) and Terminal-Bench (26% vs 34.6%). Price is the clearer gap: Sol output is $20 per 1M tokens versus $6.

Where can a small team use Grok 4.6 today?

Cursor, Grok Build, the SpaceXAI API, and OpenRouter were live at launch; SpaceXAI's news index later listed Copilot, Bedrock, Gemini Enterprise, and Microsoft Foundry. Confirm the model id in your editor or marketplace before you assume it is on your plan.

Should a clinic send patient files to Grok 4.6?

Not on the consumer app, and not on the API unless you have a BAA and a written retention setting that matches HIPAA. SpaceXAI's AUP also bars high-stakes medical decisions. Summaries that a clinician still signs are a different job from letting the model decide coverage.

Did the 2x Cursor usage promo last?

No. Cursor and SpaceXAI described doubled included usage for the first week after the August 12, 2026 launch. Budget current Cursor pool rates from the models page, not the launch promo.

Grok 4.6 is a cheaper, more persistent frontier coding model with a mixed but checkable eval sheet. US Tech Automations maps the swap onto existing agentic workflows so the routing, logs, and approval steps stay put while the model string changes. If you want that path in one place, open the agentic workflows platform and plug Grok 4.6 into a job you already run.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.

See how AI agents fit your team

US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.

View pricing & plans