Gemini Robotics 2 [What Whole-Body Control Changes]
Gemini Robotics 2 is Google's three-model physical-AI suite: one model controls robot motion, ER 2 plans embodied tasks and coordinates robots, and On-Device 2 adapts local control to supported hardware.
It is not a robot, a universal autonomy layer, or a generally available fleet product. The launch advances the connection between high-level reasoning and whole-body action, but Google also publishes uneven task results and explicit limits around movement speed and multi-finger dexterity. As of July 30, 2026, access depends on which of the three models you mean.
TL;DR: Gemini Robotics 2 matters because a shared physical-AI stack can reason about a task, allocate work across robots, and drive more of a humanoid or bi-arm body's motion. ER 2 has an AI Studio route and private preview; the action and on-device models remain early-access offerings. Before buying hardware, evaluate the exact task, embodiment, gripper demands, safety boundary, exception path, evidence, and system integration—not the humanoid label.
Key Takeaways
Gemini Robotics 2 is a vision-language-action model for whole-body control; Gemini Robotics ER 2 is the high-level embodied reasoner; On-Device 2 runs locally and adapts to supported bodies.
Google reports strong results on some gripper and whole-body tasks, but considerably weaker results on several multi-finger tasks.
A demonstration of two robots collaborating does not establish production fleet orchestration, throughput, uptime, or ROI.
The model that is easiest to access is not the same model that directly controls a humanoid body.
Businesses need structured task states, safety interlocks, human takeover, completion evidence, and a system-of-record handoff before model capability becomes an operating workflow.
No verified public price, deployment count, labor-replacement rate, or customer ROI appears in the launch material.
The Three Models in One Answer
The naming is easier when organized by altitude. ER 2 decides what physical subtask comes next. Gemini Robotics 2 converts perception and instruction into coordinated whole-body action. On-Device 2 moves part of that action capability onto the robot for lower-latency or disconnected operation and adaptation.
| Model | Main job | Typical output | Task horizon | Deployment location | Current access |
|---|---|---|---|---|---|
| Gemini Robotics ER 2 | Embodied reasoning and planning | Multistep plan or robot coordination | Several minutes | Cloud-facing development path | AI Studio and private preview |
| Gemini Robotics 2 | Vision-language-action control | Full-body or bi-arm motor action | Action sequence | Supported robot deployment | Early-access partners |
| Gemini Robotics On-Device 2 | Local action and embodiment adaptation | On-robot motor action | Local task sequence | Robot hardware | Early-access partners |
According to Google DeepMind, On-Device 2 typically adapts with fewer than 200 examples gathered over a few hours for a new bi-arm embodiment. That is an adaptation method reported by Google, not zero-shot support for any robot a buyer owns.
According to Google DeepMind, ER 2 plans tasks lasting several minutes and coordinates multiple robots from the same checkpoint. The page demonstrates the pattern; it does not publish a warehouse throughput study, a fleet service level, or a general-purpose autonomy result.
What the Task Results Show
Google reports task-specific internal evaluations. The numbers are useful because they expose variation that a highlight reel can hide. They are not independently reproduced field reliability, uptime, productivity, or ROI.
| Google evaluation task | Success rate | Miss rate |
|---|---|---|
| Whole-body table pick | 68.4% | 31.6% |
| Whole-body floor pick | 45.7% | 54.3% |
| Whole-body shelf pick | 76.3% | 23.7% |
| Gripper pick and place | 74.2% | 25.8% |
| Gripper kitting | 78.9% | 21.1% |
| Gripper precise insertion | 89.6% | 10.4% |
Source: Google DeepMind's Gemini Robotics 2 launch. “Miss rate” is transparent arithmetic from Google's success rates, not an additional Google metric.
The floor-pick result is a caution against treating “whole-body” as “solved.” Moving the base, torso, arms, and hands together expands the reachable task space, but it also introduces balance, collision, perception, and timing dependencies. A task that succeeds on a table may not transfer unchanged to a floor, shelf, pallet, or shifting jobsite.
Multi-finger manipulation is more uneven still.
| Multi-finger task | Success rate | Unsuccessful share |
|---|---|---|
| Unscrew light bulb | 92% | 8% |
| Screw light bulb | 36% | 64% |
| Tie bag | 44% | 56% |
| Use dustpan | 32% | 68% |
| Close ziplock | 40% | 60% |
Source: Google DeepMind. The unsuccessful share is calculated as 100% minus Google's reported success rate.
Google's precise-insertion result reached 89.6% in its evaluation. The Google DeepMind launch page pairs stronger gripper results with an explicit statement that movement speed and human-level multi-finger dexterity still need work. A buyer should keep the task label attached to every percentage.
What Changed: From Arm Skills to Coordinated Bodies
Earlier robotic systems often separated high-level planning, navigation, arm motion, and tool control into distinct components tuned to a fixed environment. Gemini Robotics 2 brings more of the perception-to-action chain into a shared model family and extends control to full humanoid bodies and bi-arm systems.
That matters when a task requires repositioning the body to reach, handing off an object between arms, opening access before manipulating an item, or coordinating two robots around one objective. The gain is not that integration disappears. The gain is that the control model can express a richer physical sequence while the surrounding operation supplies task state, authority, safety, and evidence.
This is distinct from a world model or navigation model. NVIDIA Cosmos 3 focuses on physical-AI world modeling; Qwen-RobotNav focuses on navigation. Gemini Robotics 2's frontier is coordinated embodied reasoning and action. Those layers may eventually work together, but the launch does not establish an interchangeable full stack.
Demonstrated, Plausible, and Still Unproven
The launch demonstrates whole-body object work, gripper manipulation, multi-finger tasks, embodied plans, on-device adaptation, and cooperation between robots. “Demonstrated” means Google showed or evaluated the behavior under its conditions. It does not mean a customer can reproduce the behavior on different hardware, objects, layouts, or safety controls.
Near-term plausibility is broader but conditional. A bounded kitting task may be a candidate when the supported gripper already handles the object set, the work cell controls human access, and the MES supplies a released order. A warehouse handoff may be a candidate when locations and inventory are accurate and the exception path can pause both robots. Those are evaluation hypotheses, not launch claims.
Several commercial conclusions remain unproven. The release does not establish sustained production uptime, a cycle-time gain, a labor-replacement ratio, total deployment cost, mean time to recovery, or customer ROI. It also does not show that an On-Device adaptation transfers among arbitrary humanoids or that a multi-robot demonstration replaces a fleet manager.
Keep a claim ledger with three columns: what the primary source reports, what the pilot must reproduce, and what the business case assumes. Every assumption should have an owner and an expiry date. If a hardware partner, integrator, or model provider supplies a new figure, attach the task, embodiment, environment, sample, and test protocol. This prevents a result for one gripper or surface from quietly becoming the forecast for an entire facility.
A disciplined claim ledger also improves procurement. It lets legal and operations separate warranty language from demonstrations, safety documentation from model benchmarks, and service-level commitments from aspirational roadmaps. When evidence is missing, the right entry is “not yet established,” not a fabricated target.
The Safety Context Is Operational, Not Abstract
Physical AI enters workplaces where the consequence of a bad action differs from a bad text answer. Current injury data does not measure these models, but it identifies the environments in which robot pilots need strong boundaries.
According to the U.S. Bureau of Labor Statistics, construction recorded 1,034 fatal work injuries in 2024, transportation and warehousing recorded 865, and manufacturing recorded 353. The respective rates were 9.2, 12.2, and 2.4 per 100,000 full-time-equivalent workers; NIST's robotics program separately places safety and performance measurement inside robot evaluation.
According to the U.S. Bureau of Labor Statistics, manufacturing recorded 332,600 nonfatal cases at a 2.7 rate in 2024; transportation and warehousing recorded 261,500 at 4.4, and construction recorded 167,100 at 2.2 per 100 full-time workers. These are industry baselines, not robot-caused incident counts; NIST's HRI work provides the separate measurement context.
| 2024 industry | Fatal cases | Fatal rate | Nonfatal cases | Nonfatal rate |
|---|---|---|---|---|
| Manufacturing | 353 | 2.4 | 332,600 | 2.7 |
| Transportation and warehousing | 865 | 12.2 | 261,500 | 4.4 |
| Construction | 1,034 | 9.2 | 167,100 | 2.2 |
Sources: BLS fatal work injuries and BLS nonfatal injuries and illnesses; NIST robotics supplies evaluation context. Fatal rates are per 100,000 full-time-equivalent workers; nonfatal rates are per 100 full-time workers.
The design response is a bounded task: known zone, known object range, validated embodiment, explicit stop conditions, independent interlock, human authorization, safe state, and retained evidence. “The model understood the instruction” is not a safety case.
What NIST's Robotics Work Adds to the Evaluation
According to NIST, its robotics program lists 8 active projects spanning safety, perception, mobility, and interaction. The program focuses on measurement science and performance evaluation, which is the missing bridge between a model demo and an operating claim.
According to NIST, its human-robot-interaction research plan organizes evaluation around 4 principal capabilities and develops test methods, metrics, and protocols for close-proximity performance. That supports a buyer discipline: specify the task and interaction conditions before comparing systems.
For each pilot, define the object set, starting states, success condition, maximum time, allowed interventions, stop triggers, people present, lighting, floor or surface condition, network status, and recovery procedure. Publish both successes and non-completions. A percentage without those conditions cannot travel safely from a lab task to an operating decision.
Google also introduced ASIMOV-Agentic as a safety benchmark. Treat it as a Google-introduced evaluation, not an established industry standard. Ask what hazards it covers, how the target deployment differs, and which independent functional-safety practices remain outside the model test.
The Workflow Layer a Robot Still Needs
A production task begins before motion. A work order, warehouse assignment, inspection request, or site plan names the desired outcome. A policy checks asset, zone, prerequisites, and authority. The robot receives a bounded instruction. An independent controller enforces safe motion. Exceptions go to an operator. Completion evidence returns to the system of record.
US Tech Automations fits at those handoffs: a released work order triggers task preparation, a missing prerequisite routes to the owner, a human approves execution, and the completion payload updates the MES, WMS, or project record. It does not supply Gemini Robotics 2, robot hardware, or a functional-safety controller.
That separation preserves optionality. The physical model can change while work-order state, approvals, exception ownership, and evidence remain stable. Teams can compare manufacturing implications, logistics implications, and construction implications without pretending one embodiment or task fits all three.
A Buyer-Ready Pilot Design
Choose a repetitive task with controlled inputs and a reversible failure. Favor stable object geometry, an enclosed or geofenced zone, clear completion evidence, and a human already accountable for the process. Avoid an open-ended “help the team” objective.
Create three test sets: normal cases, boundary cases, and stop cases. Normal cases measure task completion. Boundary cases vary lighting, object position, congestion, or connectivity within the authorized envelope. Stop cases confirm that the system refuses unsafe or underspecified work and reaches a safe state.
Track task success, intervention, cycle time, false completion, safety stop, recovery, evidence completeness, and system-of-record reconciliation separately. Do not invent target thresholds from an unsourced blog. Set them from the current manual process, risk assessment, equipment specification, and operational requirement.
If a task completes but the WMS still shows the old location, the business process failed. If the robot reports success without a quality check, the evidence is incomplete. If an operator must physically reset it after every blocked aisle, the exception path is not yet operational.
The workflow around that pilot can use agentic orchestration from US Tech Automations: the approved task becomes a controlled run, exceptions pause and reach the named operator, and the final state returns with photos, sensor evidence, or quality results. The automation layer remains subordinate to the robot's safety system.
Signal vs Speculation
Sourced signal: Google published a three-model suite, whole-body and bi-arm control, several-minute embodied planning, multi-robot coordination, on-device adaptation with a limited example set, internal task results, and explicit dexterity and speed limitations. ER 2 has the broadest development access; the action models remain early-access offerings.
Our read: in the next 12–36 months, bounded physical workflows will advance faster than general-purpose humanoid labor. The best early tasks will have stable inputs, machine-verifiable completion, low fine-finger demand, and a safe exception path. That is a deployment forecast, not a Google claim.
Our read: multi-robot coordination will create more value when tied to an external source of work than when treated as a self-contained demonstration. A WMS, MES, or project system must still allocate priority, resolve inventory or schedule conflicts, and record completion. Fleet intelligence without authoritative workflow state becomes another island.
Frequently Asked Questions
Is Gemini Robotics 2 a robot?
No. It is a physical-AI model suite designed to reason about and control supported robots. A deployment still requires compatible hardware, controllers, sensors, safety systems, task integration, and operating ownership.
Which Gemini Robotics 2 model can developers access?
Gemini Robotics ER 2 has an AI Studio path and private preview. Gemini Robotics 2 and On-Device 2 were described for early-access partners, so buyers should not assume general access to the motor-control models.
Can Gemini Robotics On-Device 2 work offline?
It is designed to run locally on the robot, which supports low-latency and disconnected operation. That does not mean every task, model update, fleet decision, or business-system handoff works without a network.
Can multiple robots collaborate?
Google demonstrates ER 2 coordinating robots on multistep work. The demonstration establishes a capability signal, not a published production fleet service level, throughput gain, or warehouse ROI.
Does whole-body control mean human-level dexterity?
No. Google explicitly identifies movement speed and multi-finger dexterity as unfinished areas, and its published task results vary widely. Whole-body coordination expands capability without making every manual task reliable.
Can a business deploy Gemini Robotics 2 now?
Only through the access channel available for the specific model and compatible hardware. A business can evaluate ER 2 more broadly, but action-model deployment requires early-access eligibility plus a task, safety, integration, and evidence plan.
Conclusion
Gemini Robotics 2 is a meaningful control-stack advance, not permission to skip deployment engineering. Keep the three models distinct, attach every success rate to its task, acknowledge the access limits, and prove the workflow under normal, boundary, and stop conditions.
The strongest pilot starts with a bounded operating record and ends with verified evidence in the same system. Use governed physical-work orchestration from US Tech Automations to route that work order, approval, exception, and completion record around the robot—while independent safety controls retain authority over motion. Get benchmarks.
About the Author

Helping businesses leverage automation for operational efficiency.
Related Articles
See how AI agents fit your team
US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.
View pricing & plans