A quantum leap in robotics with the Google Gemini Robotics ER 2 model
Google has officially unveiled its latest and most powerful robotics model, dubbed Gemini Robotics ER 2. The developers have designed it as a high-level brain that allows robots to talk to people, understand the physical environment, and plan multi-step tasks. The model then delegates the execution of movements to lower-level control systems. It can also independently use external tools, such as Google Search. The main advantage of this design is that the robot can plan its next steps while simultaneously performing current tasks, eliminating previously annoying “thinking” pauses.
The new version brings significant improvements over the previous Gemini Robotics ER 1.6 model. By analyzing continuous video, robots can now monitor their progress in real time, adjust for potential errors, and know exactly when it's time to move on to the next step. In tests, the system achieved 57.4 percent accuracy in determining the level of work completed and 91.3 percent accuracy in identifying key moments, such as the exact moment to stop pouring liquid. Another big innovation is support for multiple robots working together, which allows different devices to communicate with each other and perform more complex tasks together.
The developers have also improved spatial intelligence and a significantly higher level of safety. The model is better able to detect errors during execution, such as spills or slipping of material, and can read data from a wide variety of digital and analog measuring devices. In safety tests, it has proven to be excellent at recognizing the proximity of people. The robot automatically stops when a person approaches it and only resumes work when the area is safe again. The demonstration of its operation was performed on the well-known Spot robot from partner Boston Dynamics, which independently finds and brings a snack based on a voice command.
Robotic systems in the real world often encounter unpredictable circumstances that computer models in simulations cannot fully predict. How well the system will perform outside of the lab remains to be seen, as the model is currently available to developers via the Gemini API, Google AI Studio, and in a test version on the Gemini Enterprise Agent Platform.




















