Skip to content
Frontier Tech

Devin Security Swarm [What It Changes]

Sep 2, 2026

TL;DR

  • Devin Security Swarm is Cognition's multi-agent vulnerability hunt on Devin: parallel agents search a codebase, confirm exploits in a sandbox, and open patch pull requests.

  • Cognition launched it on 1 July 2026 and reports 72% recall on a 50-CVE set at $90.23 per run; that eval is vendor-stated, not a public ranking.

  • This is a product mode on Devin, not an open standard. Fixes still need a human merge.

  • A two-truck HVAC shop, a 10-person agency, or a solo clinic should care when a contractor ships code faster than anyone can prove the findings are real.

Swarm multiplies findings

Security Swarm on MapReduce is parallel review. Staff triage first. One service, not the estate. Confirmed findings go to the escalate queue. Unconfirmed stay out of the board.

Partner memo for Devin Security Swarm [What It Changes]

The empty object is the only decision. Write it in one sentence on the whiteboard. If you cannot, you are still in a demo.

Quotes are dated PDFs. "Around" is still a figure we will not print unless the brief's price policy allows it with an ISO date on the same line.

Week one: kill one shadow path — a personal phone, a second login, or a spreadsheet that is pretending to be the record. NFIB's 2024 figure of 44% of small businesses citing time-management as a top challenge is why you do not migrate two systems in the same sprint.

Week two: one named owner for failures. If the owner is "whoever built it," you do not have an owner.

Week three: count the copy-paste jobs that remain. That count is the workflow, not a reason to smash two products into one license.

SBA's 2025 profile of 33M+ small businesses includes shops that bought both logos and finished neither. Sign one quote. Schedule the rest 60 days later.

F514 lives or dies on whether that sentence on the whiteboard matches the screen staff will actually live in. If the screens disagree, you picked the demo, not the leak.

According to AICPA, 62% of firms reported cloud-workflow adoption.

According to Journal of Accountancy, the mid-market close still runs 8-10 business days.

According to Thomson Reuters, tax-prep utilization hits 85-95% in March and April.

According to NFIB, 44% of small businesses cite time-management. According to SBA Office of Advocacy, 33M+ small businesses sit in the 2025 profile. According to Goldman Sachs, 62% of SMBs reported workflow-tool ROI inside 12 months.

Key Takeaways

  • Agentic MapReduce is Cognition's name for splitting the codebase across parallel agents, then composing attack paths.

  • The 50-CVE comparison against Claude Security, Codex Security, and Cursor Security is Cognition's own harness.

  • PR Newswire carried the 1 July launch, the 50-advisory set, 14 languages, and a "30% lower cost per finding" line.

  • Subsequent scans process only changed code after a full baseline; cadence is daily, weekly, or custom.

  • A six-week Security Vulnerability Remediation Program is a separate enterprise engagement. No public list price for Swarm itself.

What Devin Security Swarm is, in one sentence

Devin Security Swarm is a Devin product mode that runs many agents in parallel over a codebase, reproduces suspected vulnerabilities in an isolated sandbox, and opens pull requests with patches for a human to merge. That is the entity. It is not a replacement for a scanner license inventory, not a FedRAMP authorization, and not an open MapReduce spec.

A two-truck HVAC company does not run a security team. It still pays a developer who pushed a customer portal last quarter. When a cheap scanner dumps 200 "findings," nobody knows which three would actually take a card number. A 10-person agency hits the same wall on a WordPress stack. A solo clinic hits it on a patient-intake form. Swarm's pitch is: prove exploitability, then file the fix, still with a human merge.

Law shops comparing Clio alternatives already know a finding without a next action is noise. Family-law and transactional firms comparing MyCase vs Clio Manage or Smokeball vs Clio Manage are picking a system of record. Swarm does not replace that pick. It hunts the code those systems sit on.

Why a small operator should care before the deep dive

According to the SBA Office of Advocacy, the United States contains 36.2 million U.S. small businesses. Almost none will buy an enterprise security swarm. They still ship code.

According to Cognition's Security Swarm post, some teams are seeing 10–100× more security findings, many of them false positives. That is Cognition's problem statement, not an independent census.

According to PR Newswire, Devin Security Swarm found 36 of 50 real-world vulnerabilities in Cognition's test, at 30 percent lower cost per finding than the next most accurate alternative. That is the wire restatement of Cognition's eval, not a second lab.

According to AIToolsReview, Cognition reports 72 percent recall at $90.23 per run on that 50-CVE set. The same roundup dates FedRAMP High In-Process as of 13 July 2026 and restates Devin Pro at $20 per month as the platform plan, not Swarm's SKU.

According to the same SBA release, those firms account for almost 46 percent of private-sector employment. The Advocacy homepage restates 62.3 million small-business employees. That is the labor pool whose contractor-pushed portals still need a human merge.

PR Newswire's 1 July launch cites a CloudBees-linked figure that 42 percent of code is now AI-generated or AI-assisted, and says monthly findings in some enterprises climbed from about 1,000 to more than 10,000 in six months. Those are the launch's scale claims. They are not your shop's ticket count.

If the office already automates assistant triage as in five ways to automate assistant tasks, Swarm is that triage for CVEs: prove it, patch it, queue a person. Teams already routing form-to-CRM work, as in form-to-CRM automation, still need the portal that captures the form to be patchable. Teams already routing those patch tickets through US Tech Automations workflows can treat Swarm PRs as another human-merge step, not a rebuild of intake.

What Cognition shipped on 1 July 2026

As of 1 July 2026, Cognition published Introducing Devin Security Swarm and a matching PR Newswire item. AIToolsReview's July roundup independently dated the 1 July launch, the MapReduce framing, and the self-reported eval.

The mechanism in plain language: a swarm of agents each take a region of the repo. They reason across files for business-logic flaws and chained attacks, not only signature matches. Devin composes findings into attack paths, reproduces each one in a sandbox, then writes a patch and opens a PR. Supported languages listed by Cognition include Go, Python, JavaScript, Rust, Ruby, C#, Java, Swift, PHP, Elixir, Erlang, C, Kotlin, and Dart — 14 languages, matching the GHSA eval set.

Cognition points to a technical post on Agentic MapReduce and an eval methodology page. Scan profiles can be generated from existing threat-model docs. Batch size is configurable. First scan is a full baseline; later scans process only changed code.

The anniversary shipping list puts Security Swarm and the Vulnerability Remediation Program in July 2026, after June's FrontierCode and Devin Fusion. The remediation program is a six-week embed: burn down the CVE backlog, then leave Swarm on a schedule. devin.ai/security and devin.ai/security-program are the product URLs in the PR.

Harness (Cognition's 50-CVE test)Recall$ / run
Devin Security72%$90.23
Claude Security68%$131.87
Codex Security48%$118.20
Cursor Security26%$4.60

Sources: Cognition Security Swarm post; AIToolsReview July 2026. Vendor-stated, not independently reproduced.

Cognition says only Devin found three critical issues the others missed: a PHP sandbox bypass via template injection, an argument injection through metadata parsing, and a broad deserialization surface in Spring Kafka. PR Newswire restates 36 of 50 found and 30 percent lower cost per finding than the next most accurate alternative.

How Agentic MapReduce actually works

MapReduce, in the original data-processing sense, splits work, maps a function over chunks, then reduces the results. Cognition's "Agentic MapReduce" is that shape with LLM agents: map = many agents on many code regions; reduce = compose attack paths and sandbox-confirm. It is a product architecture, not a published open standard.

The constraint that broke is not "scanners exist." Scanners scale. Triage does not. Cognition's 10–100× findings line is why a sandbox gate matters: the security team is supposed to see confirmed exploit paths, not another CSV.

The honest limit is the merge. The PR is a proposal. A person still reviews. Swarm does not claim to push to production unattended in the fetched posts.

Cognition's Series D / $26 billion post (over $1 billion raised, run-rate $492 million) is the company scale behind the product, not a Swarm benchmark. The anniversary post says the combined company grew from 44 to 350 people and revenue run rate from $73 million to $500 million+ since merging brands. Itaú is quoted there as fixing 70 percent of security vulnerabilities automatically with Devin — a named-customer line on the anniversary page, not the 50-CVE table.

NIST's AI RMF remains the voluntary U.S. risk frame if a shop needs a governance vocabulary around an agent that writes patches. NIST does not evaluate Swarm.

USTA analysis: cost delta on Cognition's own table

USTA analysis. Working only from figures already cited above: Devin $90.23 per run at 72% recall; Claude Security $131.87 at 68%; 50 CVEs; PR Newswire's 36 found and 30 percent lower cost-per-finding claim.

  • Dollar delta vs Claude Security per run: $131.87 − $90.23 = $41.64. $41.64 / $131.87 ≈ 31.6% cheaper per run, which is in the same band as Cognition's "30% lower cost" line (that line is per finding, not per run).

  • Findings at 72% of 50: 0.72 × 50 = 36, matching PR Newswire's "found 36."

  • Cost per found CVE at those vendor numbers: $90.23 / 36 ≈ $2.51 per found item on this 50-set, vs Claude $131.87 / (0.68 × 50) = $131.87 / 34 ≈ $3.88. $2.51 / $3.88 ≈ 0.65, i.e. about 35% lower cost per found item on this table. Close to, not identical to, the 30% marketing line — different denominator (per finding vs per run) explains the gap.

Derived checkInputsResult
Devin vs Claude $/run delta$131.87 − $90.23$41.64
Devin findings at 72% of 500.72 × 5036
Devin $/found on this set$90.23 / 36$2.51
Claude $/found on this set$131.87 / 34$3.88

Sources for inputs: Cognition Swarm post; PR Newswire. Results are USTA arithmetic on vendor numbers, not an independent eval.

Cursor at $4.60 and 26% recall is a different trade: cheap first pass, fewer confirmed finds. Cognition flags that trade-off; this analysis does not rank Cursor "worse" on cost.

What this does not establish

Swarm does not establish an independent leaderboard. AIToolsReview says no third-party evaluator such as Artificial Analysis had published a competing security-agent benchmark at the time of its roundup.

It does not establish a public list price for Swarm. Plans on AIToolsReview's pricing table are Devin Free / Pro $20/mo / Max $200/mo / Teams $80 + $40/seat; Swarm is described as a separate enterprise offering. Check current pricing on the vendor site. (The Devin pricing page itself returned a bot checkpoint when fetched for this article.)

It does not establish FedRAMP High authorization. AIToolsReview dates FedRAMP High In-Process as of 13 July 2026, which is not complete.

It does not establish that a two-person shop should run a 50-CVE enterprise eval. The 14-language GHSA set is Cognition's test kit.

A buyer's evaluation sequence

None of the sourced posts say the agent merges to main. Use four checkpoints.

StageScopeHuman decision
BaselineFirst full scanAccept the cost of the full pass
ConfirmSandbox exploitabilityReject findings that cannot be reproduced
PatchPR from DevinEngineer reviews and merges
CadenceChanged-code scansSet daily vs weekly vs custom

A startup that wants a generic agentic path rather than a Devin-only swarm can look at agentic workflows. US Tech Automations is the logging and approval layer if a Swarm PR still needs a conflicts-style hold before anyone deploys.

Pricing on this site is that layer. The wider operator picture is the state of small-business automation.

Adjacent Cognition surfaces that are not Swarm

FrontierCode measures mergeable code quality, not CVE recall. SWE-1.7 is the in-house model. Devin Fusion is the sidekick harness. Cognition for Government is the public-sector surface next to the in-process FedRAMP listing. None of those pages are Security Swarm.

The Cognition blog index is where the July 2026 blitz lived: Swarm on the 1st, remediation program on the 2nd, FrontierCode 1.1 on the 7th, SWE-1.7 on the 8th. AIToolsReview counted nine substantive posts in twenty days. That cadence is context, not a vulnerability score.

GitHub Security Advisories are the GHSA records Cognition says the 50-set was drawn from. This page does not re-score those advisories. NIST's AI RMF PDF and the AI Resource Center are the voluntary U.S. risk vocabulary if a shop needs to write down that an agent opens patch PRs.

Nick Wong's quote in the PR — security teams can validate and fix instead of waiting on engineering — is a vendor line. The six-week program at devin.ai/security-program is how Cognition sells the embed. A two-person shop should not treat that as a list price.

Mercedes-Benz, Infosys, Cognizant, and Itaú anecdotes live on the Series D and anniversary pages (eight months to eight days; 70 percent of vulns auto-fixed at Itaú). They are named-customer stories, not the 50-CVE table.

Signal vs Speculation

Demonstrated signal: as of 1 July 2026 Cognition launched Devin Security Swarm; Cognition reports 72% recall at $90.23/run on a 50-GHSA set across 14 languages, vs Claude 68%/$131.87, Codex 48%/$118.20, Cursor 26%/$4.60; PR Newswire independently carried the launch, 36/50, and 30% lower cost-per-finding claim; AIToolsReview independently dated the launch and labeled the eval self-reported; architecture is Agentic MapReduce on Devin, not an open standard; human merge still required; six-week remediation program is a separate offer.

Our read: over the next 12–36 months, small shops will not "deploy a security swarm." They will notice whether the contractor who ships AI-written code also files proven fixes. If Cognition's sandbox gate holds up outside its own 50-set, the constraint that broke is triage, not detection. Until an independent eval exists, treat 72% as Cognition's number. Put every PR behind a human merge, on Devin or not.

Security Swarm is more agents, not less review

Cognition shipping Devin Security Swarm on Agentic MapReduce is parallel review. Findings will multiply. Staff triage before you staff more agents.

Signal: Security Swarm. Speculation: it becomes a checkbox in enterprise RFPs. Start on one service. US Tech Automations can route confirmed findings into the escalate queue.

FAQ

What is Devin Security Swarm?

A Devin mode that hunts vulnerabilities with parallel agents, confirms them in a sandbox, and opens patch PRs for humans to merge.

When did it launch?

Cognition and PR Newswire date it 1 July 2026. AIToolsReview independently used that date.

Are the 72% and $90.23 figures independent?

No. They are Cognition's eval on Cognition's 50-CVE set. AIToolsReview flags that.

Does it replace a human security review?

No. The product writes PRs. A person still merges.

How much does it cost?

No standalone public list price appeared in the fetched Cognition posts. Devin plan prices in the July roundup are separate. Check the vendor site.

What should a small shop do this quarter?

Inventory the repos a contractor can push, require reproducible exploits before anyone treats a finding as real, and keep a human merge on every patch.

What to do next

Devin Security Swarm is a Devin-only hunt-and-patch mode with a vendor 50-CVE scorecard. The honest limit is the self-reported eval and the missing public SKU price.

If the next step is to put those PRs behind an approval step instead of merging from the agent unsupervised, start from agentic workflows for Swarm-style patch queues on the wiring layer, then set the hold on pricing.

About the Author

Garrett Mullins
Garrett Mullins
Workflow Specialist

Helping businesses leverage automation for operational efficiency.

See how AI agents fit your team

US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.

View pricing & plans