All posts

Robotics

Gemini Robotics-ER 1.6 learns to read gauges for Boston Dynamics' Spot

Google DeepMind's new embodied reasoning model reads analog instruments with 93% accuracy on its own test when it can zoom and run code, up from 23% for version 1.5.

HackHoster Team · · 11 min read

Spot, Boston Dynamics' yellow four-legged robot with an arm mounted on its back, walking across a dark stage at Web Summit 2025
Photo: Web Summit / Wikimedia Commons, CC BY 4.0

At a glance

  • Google DeepMind released Gemini Robotics-ER 1.6 on April 14, 2026, as a preview in the Gemini API and Google AI Studio under the ID gemini-robotics-er-1.6-preview.
  • On DeepMind's instrument-reading test, ER 1.6 scores 86%, or 93% with agentic vision, against 23% for ER 1.5 and 67% for Gemini 3.0 Flash.
  • Agentic vision lets the model zoom into an image, point at needles and tick marks, and run code to work out where a reading falls on the scale.
  • Boston Dynamics made a Gemini-powered version of its AIVI-Learning inspection software live for all enrolled Spot customers on April 8.
  • DeepMind reports gains of 6% on text and 10% on video safety scenarios over Gemini 3.0 Flash, but no independent evaluation was available at launch.

Google DeepMind released Gemini Robotics-ER 1.6 on April 14. It is an update to the company's embodied reasoning model, the part of a robot's software stack that looks at camera images, works out what is where, plans the steps of a task and decides whether the task is finished. The new skill DeepMind is leading with is reading instruments: pressure gauges, thermometers, level indicators, chemical sight glasses and digital readouts.

DeepMind built that capability with Boston Dynamics, which uses it on Spot, its four-legged robot, for facility inspections. Developers can call the model today through the Gemini API and Google AI Studio under the ID gemini-robotics-er-1.6-preview, and DeepMind has published a starter Colab notebook and sample code in its robotics-samples repository on GitHub.

What DeepMind released

The announcement, written by DeepMind researchers Laura Graesser and Peng Xu, describes ER 1.6 as a reasoning-first model rather than a controller. It doesn't move motors. It takes text, images, video or audio, reasons about the scene, and returns text: coordinates of objects, step-by-step plans, answers about whether a task succeeded, or calls to tools. DeepMind says it can natively call Google Search, a vision-language-action model that turns instructions into movements, or functions that developers define for their own robots.

Four areas improved, according to DeepMind: pointing at objects, success detection from a single camera, success detection across several cameras, and instrument reading. The last one is new: DeepMind calls it a newly unlocked capability, and ER 1.5 managed only 23% on the company's test.

Boston Dynamics was quick to put it to work. Marco da Silva, who runs the Spot business as vice president and general manager, said in DeepMind's announcement that instrument reading and more reliable reasoning will let Spot see, understand and react to real-world problems completely on its own.

From Gemini Robotics to ER 1.6

DeepMind has been building this model family for about a year. The timeline explains why the company keeps the reasoning model separate from the model that drives a robot's joints.

DateReleaseWhat it added
March 12, 2025Gemini Robotics and Gemini Robotics-ERA vision-language-action model built on Gemini 2.0, plus a separate reasoning model for spatial understanding that plugs into existing robot controllers
September 25, 2025Gemini Robotics 1.5 and ER 1.5ER 1.5 became a planner that orchestrates tasks and calls tools; it opened to developers in the Gemini API, while the action model stayed with select partners
April 14, 2026Gemini Robotics-ER 1.6Instrument reading, agentic vision, better pointing and multi-camera success detection

In the March 2025 launch, DeepMind said the original ER model achieved two to three times the success rate of plain Gemini 2.0 in end-to-end robot control tests covering perception, planning and code generation. It listed Boston Dynamics among its trusted testers, alongside Agility Robotics, Agile Robots and Enchanted Tools, and released a dataset called ASIMOV for testing whether robot AI makes safe decisions. With version 1.5 in September 2025, the split became explicit: the ER model acts as the high-level brain, breaks a job into steps, looks things up when it needs to, and then hands each step to the action model as a plain-language instruction. DeepMind said then that ER 1.5 reached state-of-the-art results across 15 academic embodied reasoning benchmarks.

The Spot side of the story is older. Boston Dynamics revealed Spot in June 2016 and began selling it in June 2020 for $74,500, after a period of leasing it to early customers. The robot weighs about 25 kilograms, and its developer kit has been on GitHub since January 2020. Hyundai Motor Group agreed in December 2020 to buy an 80% stake in Boston Dynamics for about $880 million and completed the deal in June 2021. Industrial users, including an offshore oil vessel operated by Aker BP, put Spot to work on routine inspection rounds within months of its commercial launch.

A blue and white Spot robot in the colors of the Swedish mining company LKAB, standing on grass next to a person
A Spot robot in the colors of Swedish mining company LKAB, shown at Almedalen Week in Sweden in 2024. Industrial operators use Spot for inspection rounds. Photo: Bene Riobó / Wikimedia Commons, CC BY-SA 4.0

Why reading a gauge is hard for a robot

Many instruments on industrial sites were designed for human eyes, not for data networks. According to DeepMind, Boston Dynamics' customers send Spot around their facilities to photograph thermometers, pressure gauges and chemical sight glasses that people would otherwise read on foot. But a photo of a dial is not a number, and turning one into the other is surprisingly tricky.

DeepMind lists what the model has to perceive: the needle, any liquid level, the container's edges, the tick marks, and the printed labels and units. Each step has its own failure mode. A needle can sit between two marks. Some gauges carry two needles that together give a reading with decimal places. And a robot's camera is rarely square-on to the instrument, which skews what it sees.

Two round analog pressure gauges with needles and bar scales lying on a gray workbench
Analog pressure gauges like these, with needles, tick marks and units in bar, are the kind of instrument ER 1.6 is built to read. Photo: Bitjungle / Wikimedia Commons, CC BY-SA 4.0

Sight glasses are harder still. These are glass windows or tubes on tanks and pipes that show a liquid's level directly. To estimate how full one is, the model has to find the boundaries of the glass, find the liquid line, and allow for distortion caused by the camera angle. Boston Dynamics now asks its software to report sight-glass fullness as a percentage from 0 to 100.

A gray metal duct with a small vertical sight glass and a lever handle, in a Cold War era shelter in Rotterdam
A sight glass gauge on a manual ventilation system in a Cold War era shelter in Rotterdam. Reading the level means finding the glass edges and the liquid line, often from an angle. Photo: AgainErick / Wikimedia Commons, CC BY-SA 4.0

How agentic vision reads an instrument

The interesting part of ER 1.6 is how it handles this. DeepMind calls the method agentic vision, and describes it as visual reasoning combined with code execution. Instead of answering from a single look at a frame, the model works through the image the way a careful technician might:

  1. Look for the parts. It identifies the needle, liquid level, container boundaries, tick marks and labels.
  2. Zoom in. It crops into small regions to make out fine markings that are unreadable at the full frame's resolution.
  3. Point and measure. It places points on reference marks and on the needle, then writes and runs code to calculate proportions and the interval between marks.
  4. Interpret. It applies what it knows about units and instrument types, for example combining two needles into one value, to state what the reading means.

For sight glasses, DeepMind says the model also corrects for perspective when estimating the fill level. In effect, the model writes a small measurement routine for each image instead of eyeballing the dial.

Embodied reasoning model, defined. In DeepMind's design, an embodied reasoning (ER) model is a vision-language model that understands physical scenes and plans tasks but outputs text, not motor commands. A separate vision-language-action (VLA) model, or a conventional controller, carries out each step. ER 1.6 is the first kind.

An industrial dial thermometer with a needle pointing near 30 degrees Celsius on a scale from 0 to 150
A dial stem thermometer used for liquids and gases. Interpolating a needle between tick marks is the kind of measurement ER 1.6 does with pointing and code. Photo: Palagiri / Wikimedia Commons, CC BY-SA 3.0

Pointing, counting and knowing when a task is done

The other upgrades are less flashy and probably matter more for general robot work.

Pointing and counting. DeepMind treats pointing as a building block for spatial reasoning: detecting objects, counting them, comparing them, reasoning about motion and checking constraints. The model can also use points as intermediate steps on the way to a harder answer. In DeepMind's example, ER 1.6 looks at a pile of tools and correctly points to two hammers, one pair of scissors, one paintbrush and six pliers, and it declines to point at items that were requested but aren't in the picture, such as a wheelbarrow. On the same image, ER 1.5 miscounted the hammers and paintbrushes, missed the scissors entirely and pointed at a wheelbarrow that wasn't there.

Hand and power tools arranged in rows on a white surface, including a hammer, pliers, wrenches, screwdrivers and a cordless drill
A spread of tools similar to DeepMind's counting test, where ER 1.6 had to point at each requested item and skip ones that weren't there. Photo: Wilfredor / Wikimedia Commons, CC0

Success detection. DeepMind calls success detection the engine of autonomy, because a robot that can't tell whether it finished a step can't decide whether to move on or try again. ER 1.6 can combine several camera streams, such as an overhead camera and one mounted on the robot's wrist, and reason about how the views relate over time. That helps when one view is blocked or badly lit. In one example, the model decides whether a blue pen actually ended up in a black pen holder.

Safety. DeepMind says ER 1.6 follows physical safety constraints better than earlier models, including decisions about which objects can be picked up given gripper or material limits.

The numbers

DeepMind's instrument-reading chart compares four setups:

ModelInstrument-reading success
Gemini Robotics-ER 1.523%
Gemini 3.0 Flash67%
Gemini Robotics-ER 1.686%
Gemini Robotics-ER 1.6 with agentic vision93%

The fine print matters. DeepMind's notes say ER 1.5 doesn't support agentic vision, so its score was measured without it, and that the instrument-reading results in its main benchmark chart were run with agentic vision turned on for the other models. The jump from 23% to 93% therefore mixes a better model with a new tool. The 86% figure for ER 1.6 without agentic vision is the cleaner comparison with version 1.5. DeepMind also notes that its single-view and multi-view success-detection tests use different examples, so those two scores can't be compared with each other.

On safety, DeepMind used an upgraded version of its ASIMOV benchmark. It reports that ER 1.6 beats Gemini 3.0 Flash by 6% on text-based safety scenarios and by 10% on recognizing hazards in video, and that it is more accurate than ER 1.5 when following safety instructions that involve pointing. DeepMind publishes charts and examples for pointing and success detection, but the announcement doesn't spell out those scores in its text.

Boston Dynamics puts it into Orbit

Spot's inspection work runs through Orbit, Boston Dynamics' platform for automating facility monitoring with the robot. One of its tools, AIVI-Learning, analyzes the images Spot captures. Boston Dynamics says a version powered by Gemini and Gemini Robotics-ER 1.6 went live on April 8 for every customer enrolled in AIVI-Learning, a week before DeepMind's public launch.

According to Boston Dynamics and The Robot Report, the update adds or improves:

  • gauge readings and digital display readings,
  • sight-glass fullness, reported from 0 to 100%,
  • pallet counting,
  • detection of puddles and standing liquid,
  • lever position detection,
  • 5S audits, which check whether a workspace is sorted and tidy.

AIVI-Learning can run on Boston Dynamics' Site Hub hardware, on a customer's virtual machine or in the cloud, and The Robot Report notes that model upgrades arrive from the cloud without customers having to install anything. There is a condition attached: Boston Dynamics says customers must share data to use AIVI-Learning, and can't opt out if they want access to these AI models. The two companies' relationship also goes beyond Spot. The Robot Report recalls that in January 2026 DeepMind and Boston Dynamics announced a partnership to build the next generation of the Atlas humanoid on Gemini Robotics models.

Limitations and open questions

Treat the 93% figure with care. It comes from DeepMind's own evaluation, on a test set it hasn't described in detail, and no independent results were available at launch. On a safety-critical reading, a few percent error is a lot, and the errors that matter most are the rare ones: a reflection on the glass, a cracked dial, a gauge mounted upside down.

DeepMind is open about the gaps. It invites teams whose specialized instruments the model handles poorly to send 10 to 50 labeled images of the failures through a form, which tells you unusual equipment will still trip it up. The model is also a preview, so its behavior and ID can change.

Latency is the other practical limit. Agentic vision depends on code execution, and each zoom-and-measure loop adds time. For a robot that photographs a gauge and moves on, a few seconds may not matter. For a robot reacting in a closed loop, it does.

The data-sharing condition on AIVI-Learning will matter to some sites, especially in regulated industries, and the boundaries of what is shared aren't described in Boston Dynamics' post.

A red emergency stop button on a yellow background, mounted on a laboratory machine's control panel
An emergency stop on a lab machine. Google's own guidance says model-level safety is no substitute for safety engineering such as e-stops and collision avoidance. Photo: Cjp24 / Wikimedia Commons, CC BY-SA 3.0

Safety deserves the same caution. When it launched ER 1.5, Google said the model's safety features are not a substitute for rigorous safety engineering, and described a layered approach that includes emergency stops, collision avoidance and risk assessments. Better ASIMOV scores don't change that advice.

What you can build with it today

ER 1.6 is a planner, not a motor controller. The usual loop is to send camera frames and an instruction, get structured output back, and map it to calls on your robot's API or to a separate action model.

A few practical notes from the documentation:

  • Coordinates. Google's ER 1.5 guide shows points returned as [y, x] pairs normalized to 0 to 1000, so scale them by the frame's height and width before handing them to a controller.
  • Limits. The model page lists inputs of text, images, video and audio, an input limit of 131,072 tokens and up to 65,536 output tokens.
  • Features. Function calling, structured outputs, code execution, thinking and search grounding are all supported, which is what makes the orchestration and agentic vision patterns possible.
  • Latency. Google recommends trading thinking budget against response time, using little for simple tasks and more for complex reasoning.
A laboratory robot arm with a gripper lifting a white assay plate during a high-throughput screening run
A lab robot arm moving an assay plate at the US National Institute of Allergy and Infectious Diseases. Arms like this usually follow fixed scripts; ER models aim to let robots plan from images and instructions instead. Photo: NIAID, public domain

You don't need a robot to try the instrument skill. A webcam pointed at a utility meter, a lab thermometer or a printed gauge face is enough to prototype a reading logger, and a weekend hackathon is about the right size for one.

Practical tip. Build the checks around the model, not just the prompt. Ask for the reading, the units and the points used, reject values outside the instrument's range, take several frames and compare them, and route anything near an alarm threshold to a person.

Other projects that fit a short event: a counting assistant for a stockroom shelf, a "did the task succeed" checker for a desktop robot arm using two cameras, or a planner that turns a natural-language request into a list of function calls on an existing robot SDK. Keep a small labeled set of your own images so you can measure accuracy instead of trusting a headline number.

What to watch

As of April 15, the near-term questions are practical ones. Developers will be able to test ER 1.6 on their own instruments and see how the 93% figure holds outside DeepMind's test set. Boston Dynamics' AIVI-Learning customers have had the Gemini-powered version for only a week. DeepMind's feedback program for failure cases suggests further updates to the preview, and the January partnership on Atlas means the same reasoning models are headed for Boston Dynamics' humanoid as well.

Sources