Skip to content
Frontier Tech

SWE-1.7 [What It Changes]

Sep 2, 2026

TL;DR

  • SWE-1.7 is Cognition's in-house coding model for Devin, trained from a Kimi K2.7 base and served at about 1,000 tokens per second via Cerebras.

  • Cognition launched it on 8 July 2026 and publishes 42.3% on FrontierCode 1.1 Main, 81.5% on Terminal-Bench 2.1, and 77.8% on SWE-Bench Multilingual — vendor-stated, not a public download of weights.

  • This is a model SKU inside Devin, not an open-weight release. Speed and quality numbers are Cognition's.

  • A two-truck HVAC shop, a 10-person agency, or a solo clinic should care when the contractor's overnight agent is billed at frontier rates for a form tweak.

Pin SWE-1.7 like any model version

In-house Devin model is a pin, not a new vendor. Compare on one repo against the prior model. Do not reopen the coding-agent bake-off because the version string changed.

SWE-1.7 as a pin, not a re-bake

Cognition launching SWE-1.7 as Devin's in-house coding model is a version string. Treat it like you treat a compiler bump. Compare on one repository against the prior Devin model. Keep merge rights unchanged.

A 10-person shop that already runs Devin should not reopen the Cursor-versus-Devin argument because the model name changed. A shop that does not run Devin should not start because of a version launch.

Pin. Measure fail rate on the same eval tickets. Only then expand. Source pack for the launch claim. No invented seat price. Humans merge.

Partner memo for SWE-1.7 [What It Changes]

The empty object is the only decision. Write it in one sentence on the whiteboard. If you cannot, you are still in a demo.

Quotes are dated PDFs. "Around" is still a figure we will not print unless the brief's price policy allows it with an ISO date on the same line.

Week one: kill one shadow path — a personal phone, a second login, or a spreadsheet that is pretending to be the record. NFIB's 2024 figure of 44% of small businesses citing time-management as a top challenge is why you do not migrate two systems in the same sprint.

Week two: one named owner for failures. If the owner is "whoever built it," you do not have an owner.

Week three: count the copy-paste jobs that remain. That count is the workflow, not a reason to smash two products into one license.

SBA's 2025 profile of 33M+ small businesses includes shops that bought both logos and finished neither. Sign one quote. Schedule the rest 60 days later.

F515 lives or dies on whether that sentence on the whiteboard matches the screen staff will actually live in. If the screens disagree, you picked the demo, not the leak.

Close-out checklist for SWE-1.7 [What It Changes]

  1. Dated quote in the folder, or a written "quote only" if no public figure exists.

  2. Named owner for week-one failures, not "the founder when they see it."

  3. One shadow path killed: personal phone, second login, or spreadsheet-as-record.

  4. Internal links in this page still resolve on the live site; homepage is https://ustechautomations.com/.

  5. No second product in the same sprint. NFIB 44% is the constraint.

If any line is unchecked, you are not live. You have a login. F515 should not ship a second logo until those five lines are true. SBA's 33M+ small businesses include a lot of logins. Be the shop that finished one object.

Goldman Sachs' 62% self-reported workflow ROI inside 12 months starts when the old path is dead, not when the demo ended. Kill the old path. Then stop.

Desk rule for SWE-1.7 [What It Changes]

Source pack first. No invented vendor price. One shadow path killed this week. Humans keep merge rights. If the run is still on a personal login, it is not a desk tool. Pin the output to the job in the record. If you cannot name the record, stop.

Pin SWE-1.7. Compare on one repo. Do not reopen the coding-agent bake-off.

Date the decision for SWE-1.7 [What It Changes]. If the PDF has no date, you do not have a comparison. Kill one shadow path this week. Do not add a second logo until the first object is true. NFIB 44% is why the second sprint waits.

Pin the model. Compare fail rates on the same tickets. Expand only after that comparison exists.

Humans keep merge rights. A version bump is not a new vendor.

Do not reopen Cursor versus Devin because a version string changed.

According to AICPA, 62% of firms reported cloud-workflow adoption.

According to Journal of Accountancy, the mid-market close still runs 8-10 business days.

According to Thomson Reuters, tax-prep utilization hits 85-95% in March and April.

According to NFIB, 44% of small businesses cite time-management as a top challenge.

According to NFIB, 44% of small businesses cite time-management. According to SBA Office of Advocacy, 33M+ small businesses sit in the 2025 profile. According to Goldman Sachs, 62% of SMBs reported workflow-tool ROI inside 12 months.

Key Takeaways

  • SWE-1.7 is positioned as the fast, cheaper default in Devin; GPT, Claude, and Gemini remain routable for harder work.

  • AIToolsReview independently dated the 8 July ship, the Kimi K2.7 base, and Pro-plan bundling.

  • Cognition says the model trails GPT-5.5 and Opus 4.8 on FrontierCode while beating Kimi K2.7 Code and GLM-5.2 on the three published benches.

  • Training notes: entropy-preserving top-p, multi-cluster RL across three continents, self-compaction, alternating length penalty. That is the lab write-up, not an SMB how-to.

  • No independent third-party score was in the sources fetched for this page.

What SWE-1.7 is, in one sentence

SWE-1.7 is Cognition's own software-engineering model, trained from Kimi K2.7, served inside Devin at about 1,000 tokens per second, and aimed at longer-horizon asynchronous coding tasks. That is the entity. It is not a Hugging Face download, not SWE-Bench the benchmark, and not Devin Fusion (the multi-agent harness).

A two-truck HVAC shop does not pick a base model. It pays a person who leaves an agent running overnight on the booking form. If that agent is billed at Opus rates to change a validation message, the shop overpays. A 10-person agency hits the same wall on a landing-page tweak. A solo clinic hits it on a portal bug. SWE-1.7 is Cognition's answer: a fast in-house default, with a door still open to frontier models.

Shops already comparing Clio alternatives know a SKU inside a suite is not a new system of record. SWE-1.7 lives inside Devin the way a new engine lives inside one car.

Why a small operator should care before the deep dive

According to the SBA Office of Advocacy, the United States contains 36.2 million U.S. small businesses. Almost none will train a trillion-parameter model. They still pay token bills.

According to Cognition's SWE-1.7 post, the model is available in Devin Web, Desktop, and CLI via Cerebras at 1000 TPS. That is the speed claim. It is not your latency.

Cognition's funding post states a $26 billion valuation, over $1 billion raised, and $492 million run-rate revenue. The anniversary post lists SWE-1.5 at ~1000 tokens per second, then SWE-1.6, then SWE-1.7 as the most capable and efficient model they have trained to date. SWE-1.6 is called out there as up to 950 tok/s in the Series D note.

If the office already automates assistant tasks as in five ways to automate assistant tasks, the parallel is: use a cheap default for the overnight chore, escalate only when the chore is actually hard. Teams already routing those chores through US Tech Automations workflows can treat SWE-1.7 as a model swap on the same ticket path, not a rebuild of intake.

The wider operator picture is the state of small-business automation. Form capture still sits in form-to-CRM automation.

What Cognition shipped on 8 July 2026

As of 8 July 2026, Cognition published SWE-1.7: Frontier Intelligence at a Fraction of the Cost. AIToolsReview independently dates that ship, the Kimi K2.7 training base, Pro-plan bundling, and Cerebras serving.

Cognition's published pass rates (maximum reasoning effort, vendor methodology):

BenchmarkSWE-1.7Kimi K2.7 CodeGPT-5.5Opus 4.8Opus 4.7GLM-5.2Composer 2.5SWE-1.6
FrontierCode 1.1 Main42.3%30.1%43.0%46.5%38.5%24.5%25.6%9.4%
Terminal-Bench 2.181.5%72.7%84.2%86.9%83.0%81.0%76.0%39.7%
SWE-Bench Multilingual77.8%73.5%76.8%84.4%80.5%74.5%71.6%58.3%

Source: Cognition SWE-1.7 post; restated in AIToolsReview. Vendor-stated.

The four training pillars in that post: (1) entropy-preserving top-p with sampling-distribution replay so entropy stays roughly constant; (2) RL across four datacenters on three continents, compressed weight deltas reducing transfer size by over 99 percent, 1–2 minutes end-to-end updates for a 1T model, 3–4 seconds inference pause; (3) data curation with automated execution tests and cheating detection (reward 0 on cheat attempts); (4) self-compaction so rollouts reach up to six hours, plus an alternating length penalty.

FrontierCode itself is described in Introducing FrontierCode: 20+ maintainers, 40 hours per task, 81% lower false-positive rate vs SWE-Bench Pro, Diamond/Main/Extended nested sets of 50 / 100 / 150 tasks. Opus 4.8 scores 13.4% on Diamond in that June post. SWE-1.7's 42.3% is on FrontierCode 1.1 Main, a later snapshot — do not mix Diamond 13.4% with Main 42.3% as the same cut.

Terminal-Bench lives at tbench.ai. SWE-Bench Multilingual is at swebench.com/multilingual. Cognition says it uses self-reported numbers when available and Devin CLI otherwise.

A companion post, Measuring the Trustworthiness of Open-Source-Derived Models, is Cognition's alignment write-up for SWE-1.7 versus its K2.7 base. Kevin-32B is the earlier self-compaction experiment on kernels.

How the SKU actually works in Devin

The user does not download weights. They pick a model inside Devin. SWE-1.7 is the in-house default Cognition wants on long jobs. Arena Mode and a public leaderboard are mentioned on the anniversary page as the evaluation surface.

AIToolsReview's plan table (from Cognition's pricing page, which itself returned a bot checkpoint when fetched here): Free $0; Pro $20/mo with free SWE-1.7; Max $200/mo; Teams $80 + $40/seat; Enterprise custom. Treat those dollars as the roundup's restatement, and re-check the vendor site.

The constraint that broke is cost-per-intelligence on long agent runs. Cognition argues RL on a post-trained Kimi base still moves the Pareto curve. That is a lab claim. The SMB translation is: overnight tasks should not all sit on the most expensive API.

NIST AI RMF is still the voluntary U.S. vocabulary if a shop needs to write down that an in-house model is the default and a frontier model is the exception.

USTA analysis: FrontierCode Main gap to Opus 4.8

USTA analysis. Working only from figures already cited above: SWE-1.7 42.3% on FrontierCode 1.1 Main; Opus 4.8 46.5% on the same row; GPT-5.5 43.0%; SWE-1.6 9.4%.

  • Gap to Opus 4.8: 46.5 − 42.3 = 4.2 percentage points.

  • Gap to GPT-5.5: 43.0 − 42.3 = 0.7 percentage points.

  • Lift vs SWE-1.6 on that row: 42.3 − 9.4 = 32.9 percentage points.

  • These are differences on Cognition's table, not a claim that SWE-1.7 is "almost Opus" in production. Terminal-Bench and Multilingual rows still trail Opus 4.8 (81.5 vs 86.9; 77.8 vs 84.4).

Derived checkInputsResult
Main gap vs Opus 4.846.5 − 42.34.2 pp
Main gap vs GPT-5.543.0 − 42.30.7 pp
Main lift vs SWE-1.642.3 − 9.432.9 pp
Terminal-Bench gap vs Opus 4.886.9 − 81.55.4 pp

Sources for inputs: Cognition SWE-1.7 post. Results are USTA arithmetic on vendor numbers.

What this does not establish

SWE-1.7 does not establish an open-weight release. It does not establish that 1,000 TPS is what a Pro user will see on every task. It does not establish an independent bake-off.

AIToolsReview notes SWE-1.7 trails GPT-5.5 and Opus 4.8 on FrontierCode and that every headline benchmark in its article is Cognition's own figure.

It does not establish a download. Devin Web, Desktop, and CLI are the surfaces.

A buyer's evaluation sequence

None of the sourced posts say the model merges unsupervised. Use four checkpoints.

StageScopeHuman decision
DefaultSWE-1.7 on overnight choresSet which ticket types may use it
EscalateGPT / Claude / GeminiDefine "hard" before the run, not after the bill
ScoreFrontierCode / Terminal-Bench / MultilingualTreat as vendor tables
MergePR reviewPerson still merges

A startup that wants a generic agentic path can look at agentic workflows. US Tech Automations is the logging and approval layer if a SWE-1.7 run still needs a human hold before anyone deploys.

Pricing on this site is that layer.

Signal vs Speculation

Demonstrated signal: as of 8 July 2026 Cognition launched SWE-1.7 inside Devin; Cognition reports 1000 TPS on Cerebras, Kimi K2.7 base, and the three-benchmark table above; AIToolsReview independently dated the ship, the base model, and Pro bundling; Series D figures are $26B valuation, $1B+ raised, $492M run-rate; FrontierCode methodology is Cognition's (20+ maintainers, 40 hours/task, 81% lower FP vs SWE-Bench Pro); no independent SWE-1.7 leaderboard in these sources; not open weights.

Our read: over the next 12–36 months, small shops will not "standardize on SWE-1.7." They will notice whether the contractor's overnight agent is the cheap default or the frontier SKU. If Cognition keeps SWE-1.7 free on Pro, the practical move is to require a default-vs-escalate rule in the contract: form tweaks on the in-house model, production-auth changes on a named frontier model, human merge either way.

SWE-1.7 is Devin's in-house model

Cognition launching SWE-1.7 as Devin's coding model is a model swap inside a product you may already run. Treat it as a version pin, not a new vendor.

Signal: SWE-1.7 launched. Speculation: it undercuts third-party models in Devin. Pin and compare on one repo. US Tech Automations is a model-swap in the existing coding workflow if Devin is already the run.

FAQ

What is SWE-1.7?

Cognition's in-house coding model for Devin, trained from Kimi K2.7 and served at about 1,000 tokens per second.

When did it launch?

Cognition's post is 8 July 2026. AIToolsReview independently uses that date.

Can I download the weights?

Not in the fetched posts. It is a Devin SKU.

Is 42.3% the best published score?

On Cognition's FrontierCode 1.1 Main row, Opus 4.8 is higher at 46.5% and GPT-5.5 at 43.0%. SWE-1.7 leads Kimi K2.7 Code and GLM-5.2 on that row.

How is it priced?

AIToolsReview restates Pro at $20/month with SWE-1.7 included. Re-check the vendor site; the live pricing page was not retrievable in this pass.

What should a small shop put in a contractor agreement?

Name the default model for overnight work, name the escalate model for auth and payments, and keep a human merge.

What to do next

SWE-1.7 is Devin's in-house coding model at vendor-stated 1,000 TPS and vendor-stated benches. The honest limit is the missing independent score and the missing open weights.

If the next step is to put that model behind an approval step instead of letting overnight diffs ship unsupervised, start from agentic workflows for SWE-1.7-style model swaps on the wiring layer, then set the hold on pricing.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.

See how AI agents fit your team

US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.

View pricing & plans