TECHNICAL DEEP DIVE

Gemini Robotics for Product Managers: Google DeepMind's VLA Platform Explained

By Institute of AI PM·16 min read·Oct 3, 2026

TL;DR

Gemini Robotics is Google DeepMind's Vision Language Action (VLA) foundation model platform for robotics, announced at CES 2026 and now powering the next generation of Boston Dynamics' Atlas humanoid. Unlike earlier robotics AI that required separate perception, planning, and control models, Gemini Robotics handles all three from a single multimodal foundation. For AI PMs evaluating physical AI strategies, this represents a genuine platform shift: the stack looks more like foundation model APIs than traditional robotics software, which changes how products are specified, how teams hire, and how evaluation is done.

The AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.

What Gemini Robotics Actually Is

Gemini Robotics is not a robotics operating system. It is a family of foundation models built on the Gemini multimodal architecture that have been extended to understand physical space and output robot actions. The key difference from earlier approaches is the integration layer: traditional robotics AI stacks separate perception (what does the camera see?), reasoning (what should the robot do?), and control (how should the joints move?). Gemini Robotics collapses these into a single model that takes visual input and natural language instructions and outputs actions directly.

The platform was announced at CES 2026 in partnership with Boston Dynamics and represents DeepMind's application of the same scaling and transfer learning principles that made Gemini effective for text and image tasks. The core insight is that a model trained on internet-scale text and images already understands objects, spatial relationships, and human intent. Fine-tuning that foundation on robot trajectories produces a system that generalizes to new tasks with far less task-specific data than older approaches required.

1

Gemini Robotics Core

The base VLA model. Takes vision, language, and proprioceptive inputs (joint positions, force sensors), outputs action sequences. Handles manipulation tasks: pick and place, assembly, inspection. Runs on robot onboard compute or edge server depending on latency requirements.

2

Gemini Robotics-ER (Extended Reasoning)

The version used for complex multi-step tasks requiring planning. Before executing, the model generates a chain-of-thought reasoning trace about how to decompose the task. Slower to initiate but recovers from errors mid-task rather than requiring a human restart.

3

Gemini Robotics-Nav

Specialized for navigation and spatial reasoning. Used in Spot (Boston Dynamics' quadruped) for facility inspection and autonomous traversal. Understands floor plans, safety zones, and dynamic obstacle avoidance with natural language instruction following.

How VLA Models Differ From Earlier Robotics AI

To understand why Gemini Robotics is architecturally different, it helps to understand what the previous generation looked like. Before VLA models, production robotics AI typically used a pipeline of specialized models: a perception model (usually a computer vision model fine-tuned on the robot's camera view), a planning module (often rule-based or a smaller neural network), and a control policy trained via reinforcement learning on specific tasks.

Previous approach: task-specific policies

A model trained to pick a red block could not pick a blue cylinder without retraining. Each new task or environment required collecting new demonstration data (typically thousands of human-operated trajectories) and training a new control policy. Deployment was a one-task-per-model proposition.

VLA approach: general policies from foundation models

Gemini Robotics follows natural language instructions ('put the bolt in the leftmost bin') on tasks it was not explicitly trained for. The foundation model's language and spatial understanding generalizes. New task specification requires a prompt, not a dataset.

Previous approach: brittle to novel objects

If an object looked different from the training distribution (different lighting, orientation, packaging), perception failed and the manipulation attempt failed. Visual foundation models trained on internet-scale images are far more robust to this distribution shift.

VLA approach: reasoning about uncertainty

Gemini Robotics-ER can output 'I am not confident about the object pose' before attempting a grasp and request a better camera angle or human confirmation. Earlier policies had no uncertainty representation and would attempt the action regardless.

The Boston Dynamics Partnership: What It Means for Products

The CES 2026 announcement paired Gemini Robotics with Boston Dynamics' next generation Atlas humanoid robot, with Hyundai-owned Boston Dynamics planning to deploy Atlas units in Hyundai manufacturing facilities starting in 2028. For AI PMs, the commercial partnership is less important than what it signals about the product architecture:

Hardware is the new compute

The Atlas hardware serves roughly the same role in physical AI that GPUs serve in software AI: the constraint that determines what models can run and at what cost. Boston Dynamics controls the hardware interface, DeepMind controls the foundation model, and the enterprise customer builds the application on top. The same three-layer model as cloud AI, with hardware replacing cloud infrastructure.

Product specification shifts from behavior programming to instruction writing

Traditional industrial robotics required programming every motion path. Gemini Robotics-powered robots receive natural language task instructions from the application layer. For enterprise AI PMs, this means the product specification work shifts toward defining task instructions, success criteria, and safety boundaries in language rather than motion coordinates.

Evaluation is now similar to LLM evaluation

Earlier robotics products were evaluated on success rate for specific programmed tasks. VLA products require evals that span generalization to new objects, recovery from unexpected states, and appropriate uncertainty handling. This is much closer to how you evaluate an LLM agent than how you evaluate a motion-planning system.

Apply Physical AI Concepts to Real Products

The AI PM Masterclass covers emerging platforms like Gemini Robotics alongside software AI product strategy, taught live by a Salesforce Sr. Director PM.

Evaluating Gemini Robotics for Your Product: The PM Framework

If you are evaluating whether to build on Gemini Robotics or a competing VLA platform, these are the five dimensions that determine fit for a production deployment in 2026:

Task generalization vs. task-specific performance

Does your application require high success rates on a fixed set of tasks (use an older fine-tuned policy) or the ability to handle new task instructions without retraining (use a VLA platform like Gemini Robotics)? Do not pay for generalization you do not need.

Inference latency requirements

Gemini Robotics-Core runs inference in 100 to 400ms for manipulation decisions. Gemini Robotics-ER with reasoning can take 1 to 3 seconds. For high-speed assembly lines, this matters. For inspection or pick-and-place at human speed, it does not. Get this number before evaluating anything else.

Hardware lock-in tradeoffs

Gemini Robotics is tightly coupled to Boston Dynamics hardware and the DeepMind API stack. Competing VLA platforms (Pi Zero from Physical Intelligence, OpenVLA, Octo) offer more hardware portability. If you need to run across multiple robot form factors, the lock-in cost may outweigh Gemini Robotics' capability lead.

Safety certification complexity

VLA models produce less interpretable decision traces than programmed motion paths, which complicates safety certification in regulated environments (automotive, pharma, food processing). Ask vendors directly about their certification support before starting pilot programs in regulated facilities.

Demonstration data requirements

Even VLA models require some task-specific fine-tuning for reliable production performance. Estimate the number of demonstration trajectories you will need to collect for your task domain and include that cost in your build-vs-buy analysis.

The Competitive Landscape in 2026

Gemini Robotics is the most prominent VLA platform but not the only one. Understanding the competitive landscape matters for AI PMs making platform decisions and for those tracking the market for strategic reasons.

Physical Intelligence (Pi Zero and Pi Zero 2W)

Stanford spinout backed by OpenAI and Jeff Bezos. Pi Zero 2W achieves strong generalization across dexterous manipulation tasks with a flow matching architecture. Hardware-agnostic by design: the explicit strategic differentiator from Gemini Robotics.

OpenVLA

Open weight VLA from Stanford, Berkeley, and Toyota Research Institute. Released as an open weight model, enabling customization and on-premise deployment. Performance below frontier closed models but zero API cost and full fine-tuning control.

Figure AI with OpenAI partnership

Figure's humanoid robots run on an OpenAI multimodal backbone for task planning with Figure's proprietary motion control stack below it. Focuses on manufacturing and logistics deployments, competing directly with the Gemini Robotics plus Atlas offering.

1X Technologies (EVE and NEO)

Norwegian company with full vertical integration: builds hardware, collects proprietary training data, and trains their own foundation models. No third-party API dependency. Strong data flywheel thesis if deployments scale.

What AI PMs Need to Do Differently for Physical AI Products

Building on Gemini Robotics or any VLA platform requires adjusting several standard AI PM practices for the physical deployment context.

1

User research requires on-site fieldwork

Session recordings and user interviews do not capture how people interact with physical robots in real facilities. Plan for direct observation in the deployment environment. Operators form mental models of robot capability through physical proximity that are invisible in any remote research method.

2

Failure modes include physical harm

Your risk register needs a physical harm category. A hallucinating text LLM outputs a wrong answer. A hallucinating VLA moves a robot arm unexpectedly. Red-teaming a VLA product includes adversarial physical scenarios: unexpected objects in the workspace, people moving through the robot's path, sensor occlusion.

3

Rollback is not a config change

Updating a software model is a deployment. Updating a physical AI system may require recertification, new demonstration data, or updated operator training. Budget iteration cycles at 10x the length of software AI feature cycles. This changes sprint cadence expectations significantly.

4

Your acceptance criteria need physical tolerances

Software AI acceptance criteria are typically output quality thresholds. Physical AI acceptance criteria include grasp success rate on your specific object library, recovery time from error states, and cycle time consistency across an 8-hour shift. Define these with engineering before pilot planning, not after.

Build Products on the Next Wave of AI Platforms

The AI PM Masterclass covers how to evaluate and build on emerging AI platforms, from foundation models to physical AI. Taught live by a former Apple Group PM.

Before you go: get the AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.