Tech
Google DeepMind’s Gemini Robotics ER 2 Sits Above the Motor Control Systems, Brings Whole Body Intelligence
Google DeepMind released Gemini Robotics ER 2 as the high-level reasoning layer in a new trio of models built for robots that need to move, plan, and work alongside people or other machines. This version sits above the motor-control systems and acts as the coordinator.
Continuous video from the robot’s cameras continues to stream into the model, allowing it to monitor task progress in real time, decide where it is and what it needs to do next, all without freezing. In tests, it did a good job of classifying video frames into five phases of completion approximately 57% of the time, and pinpointing the exact moment something should stop, such as drinking a cup of coffee, with a staggering 91% accuracy and less than one second of average mistake.
Google Fitbit Air – Screenless Activity Tracker with Fitness, Heart Rate, and Sleep Tracking…
- Google Fitbit Air is the unbelievably comfortable, exceptionally smart way to transform your health[1]; and Google Health brings together effortless…
- Unlock more with Google Health Premium: With a premium membership, get personalized coaching that’s built with Gemini and adapts to your life…
- Comfortable fit – One Size Tracker (130-210 mm): The lightweight, micro-adjustable fit sits comfortably and quietly, so you can wear Google Fitbit Air…
That means the model can orchestrate these multi-step tasks, which take minutes to accomplish. Developers can offer lower-level controls, navigation routines, and even remote control tools, which ER 2 can use as needed while maintaining a watch on the scene. If something goes wrong, it can just retry the failed step rather than restarting the entire sequence. The same thing that allows for continual awareness also makes collaboration much easier. Robots can share a common understanding of what is going on in the workplace, allowing a small wheeled robot to hand off to a humanoid one that can handle uneven ground, or two arms to work next to each other without colliding.
The associated vision-language-action model accounts for the entire body motion. For the first time, this technology can drive a complete humanoid robot from its feet up. On Apptronik’s Apollo 2, a spoken order such as “put the watering can into the green bin on the bottom shelf” can be translated into action. It can move across the room, pick up the can, stroll across the floor, squat or stretch as needed, and place the item exactly where it should be. The model can manage these five-fingered hands with 22 degrees of freedom, which can tie a knot or seal a ziplock bag, as well as the simpler grippers seen on robots for packaging.
It adjusts to unfamiliar hardware in hours rather than weeks. With just 200 samples of motion data, the on-device model may learn skills and apply them to different hardware configurations, sensor layouts, and even joint counts. Early testing included robots from Apptronik, Franka, and a few research arms.
Of course, safety is just as important as any of this, and ER 2 has received great results on these benchmarks, which determine if the model sticks to physical limitations and detects when people are present. They showed a humanoid robot that stops dead in its tracks when a person enters the workstation and does not continue until the area is clear again. They’ve also created a new evaluation called ASIMOV-Agentic, which assesses the model’s willingness to say no to something just unsafe, to judge whether a work is even viable, or to seek human assistance if in doubt. The end result is a system that combines hardware safety with active risk management.
Developers can already access ER 2 via Google AI Studio and a private preview of the Gemini Enterprise Agent Platform. Lower level action models are still only available to a select group of partners. They’ve shared some sample code on GitHub that connects the reasoning layer to live robot feeds, as well as a video of Boston Dynamics’ Spot fetching a bag of popcorn using natural language control, which demonstrates the entire system in action.
You must be logged in to post a comment Login