Skip to main content
AI Development

Gemini Robotics 2 is three products, and only one is public

Google's Gemini Robotics 2 family includes a public reasoning API and two limited-access action models. Choose the layer that matches what you can test.

Sean McLellan profile photo

Sean McLellan

Lead Architect & Founder

6 min read
A barista's hand sorts roasted coffee beans beside tasting cups, a portafilter, and a tamper while a robotic arm sits behind a transparent divider and a compact brewer stands on a separate platform.
A public reasoning layer and two limited-access action layers call for different experiments. Generated editorial image, not Google product UI or a robotics benchmark.

Google DeepMind announced Gemini Robotics 2 on July 30, 2026, but the family name covers three products with different outputs and access. Gemini Robotics ER 2 is a public reasoning model that returns text. Gemini Robotics 2 and Gemini Robotics On-Device 2 are robot-action models with limited access. A team that treats the announcement as one available robotics stack can choose the wrong interface before its experiment reaches a robot.

For software, operations, automation, and product leaders, the decision starts with the output the experiment needs. This article explains which product can plan and call tools, which products can issue robot actions, and where Google’s safety evidence stops. It shows when to test a public reasoning API, pursue an action model, or stop because the work is safety-critical.

One family name covers one reasoning model and two action models

Google’s family announcement separates three products into two model classes. Their output paths are reasoning text and tool calls from ER 2, robot actions from Gemini Robotics 2, and numerical robot actions generated locally by On-Device 2. ER 2 is a vision-language model, or VLM, which processes language with visual inputs and returns text. The other products are vision-language-action models, or VLAs, which convert language, visual information, and robot state into actions.

ER 2 interprets an instruction, examines the environment, prepares a plan, and selects a tool. A VLA converts that instruction and the current robot state into actions. The broader physical-AI field includes other ways to combine temporal awareness, spatial reasoning, and control. This release puts those functions in products with separate access terms.

Gemini Robotics ER 2 is public, but its output is text

The ER 2 model card describes a VLM based on Gemini 3.5 Flash. It accepts interleaved text, images, video, and audio, then returns text. Google distributes it through the Gemini API and Google AI Studio, while Gemini Enterprise Agent Platform remains a private preview.

ER 2 can reason about physical scenes and time, plan multi-step tasks, track progress, detect success or failure, and communicate with people. Google’s developer announcement also describes declared tool use. A declared tool is an external function that a developer makes available to the model, such as a navigation API or lower-level VLA interface.

ER 2 can select that function and send the call that the developer defined. It does not produce the numerical motor actions that move a robot. The application, action model, controller, and hardware remain separate parts of the control path.

Robot actions require a limited-access VLA model

Gemini Robotics 2 is the general VLA in the family. Google says it converts text and image input into motor control for several robot forms. Its product page lists the output as actions, labels the status as private preview, and directs developers to an early-access waitlist.

Gemini Robotics On-Device 2 is a local VLA based on Google’s on-device Gemma technology. It accepts text, images, and proprioception, which is numerical sensor data about robot position and state. It outputs robot actions as numerical values and is distributed only to trusted testers. Its model card lists out-of-distribution tasks and robots with many degrees of freedom as limitations.

In a constructed coffee workflow, ER 2 could call a developer-defined tamping tool after deciding that the step should run. A VLA would use the current visual input and robot state to produce action output for a specific arm. On-Device 2 represents its robot actions as numerical values. This path depends on the robot, sensor definitions, action representation, runtime, controller, and stop behavior.

A barista sorts roasted coffee beans beside a ceramic cup and a filled portafilter while a separate robotic arm presses a tamper into another portafilter held in a metal cradle.
Generated editorial image, not Google product UI or a robotics benchmark.

A team can test ER 2’s reasoning and tool selection now without receiving either VLA. The public API does not include a public motor-control model or a supported path to the team’s hardware. Action experiments depend on model access and robot integration.

The demos and evaluations cited here come from Google. They do not establish production reliability, supported hardware, commercial terms, or success on a team’s tasks.

Semantic-safety results do not replace functional-safety proof

Google’s Gemini Robotics 2 safety report evaluates high-level semantic reasoning and orchestration. Semantic safety here means that the model interprets an instruction or scene, then chooses a VLA call, a request for human clarification, or a safety-tool call. The report tests unsafe-task refusal, human-proximity monitoring, VLA feasibility checks, and responses to ambiguity.

Functional safety uses system and hardware mechanisms to keep equipment in a safe state when a component fails or a hazard occurs. The report states that it did not evaluate the underlying functional-safety architecture, including certified hardware components, redundancy mechanisms, and real-time system guarantees needed for a compliant physical deployment. The report is not a functional-safety certification.

The human-proximity results reinforce that boundary. A false negative means that the model fails to call for a stop when a person is inside the safety perimeter. A false positive is an unnecessary stop. No tested model achieved near-zero false-negative and false-positive rates in that evaluation. The report says deterministic, low-level safeguards remain necessary.

The ER 2 model card also tells users to apply discretion in production, commercial, or public environments. It says not to use the Robotics Models for safety-critical applications or work, including settings where a malfunction could lead to death, injury, or property damage.

The required output and safety boundary decide what to test

A bounded ER 2 evaluation can use recorded or streamed input with declared test tools. The team can check planning, progress judgments, clarification, safe refusal, and exact tool calls without treating the result as motor control.

A physical VLA pilot starts after a team receives limited access. It also needs a supported robot, sensor and action definitions, a compatible runtime, low-level controllers, and hardware safeguards. The team must verify what each component sends, accepts, rejects, and stops.

If the experiment ends in a plan, progress judgment, explanation, or declared tool call, test Gemini Robotics ER 2 now through the Gemini API or Google AI Studio. If it must produce robot actions, seek early access to Gemini Robotics 2. If those actions must be produced locally, seek trusted-tester access to Gemini Robotics On-Device 2. If the work is safety-critical, follow the model-card restriction and do not use this family.

Before a physical pilot, document where reasoning ends, where numerical action begins, and which controller, hardware safeguard, or person can block an unsafe request. That local work determines whether a capability can become a controlled experiment. Access to a model does not remove the integration and safety work around it.

For teams with a proposed robot workflow, BaristaLabs AI consulting can review the reasoning-to-action path, the required integrations, and the evidence needed for a go-or-hold decision. Discuss one robotics AI experiment.

Robotics AI follow-up

Apply this decision to one experiment

BaristaLabs can discuss one proposed robotics workflow, the required action and integration layers, and the safeguards that must remain outside model judgment.

Best fit for teams deciding whether a physical-AI idea needs a reasoning API, an action model, an on-device runtime, or a different safety boundary.

Turn this idea into a pilot

Which workflow should go first?

Use the readiness check to compare impact, effort, risk, owner, and next step before booking a call.

  • 3-5 minutes
  • Deterministic score
  • No sensitive data
Check workflow readiness

Practical AI Workflow Notes

Want more practical AI operations ideas?

Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.

A useful next step if you’re still exploring and not ready to book a 20-minute AI assessment.

Occasional emails. Practical workflow guidance only. Unsubscribe anytime.