Can We Trust AI to Rescue a Robot Alone at Sea?
When a robot sub's nose starts to drop under the ice, there's no human to call. Can an AI decide: keep going, or abort? Researchers at Johns Hopkins tested four AI models on exactly this, in 480 simulated trials.
They built SPAR, a simulation platform for testing how AI handles faults on autonomous underwater vehicles (AUVs). The test fault is based on real events on MBARI's Tethys-class AUVs, where a battery shifted forward and tipped the nose down. While the autopilot keeps flying, the AI reads a status report, works out likely causes and writes a new mission plan.
They tested two sizes of the shift. A small one, 5 millimeters: the fins trim it out, and the mission can safely continue. A big one, 5 centimeters: the right call is to abort. Same mission every time: dive to a 200-meter band, and cruise about a kilometer.
They tested one frontier model (GPT-5.5) and three small models that run on a single graphics card (Nemotron, GPT-OSS and Gemma 4). A few of the findings:
- For diagnosis, the choice of model mattered most. GPT-5.5 put the shifted weight in its top three suspects in 85–90% of trials; the best small model, Nemotron, in 60–78%.
- More detail in the prompt didn't clearly help.
- Finding the cause and making the right call were separate skills. When the big fault hit mid-dive, GPT-5.5 made the wrong call two times out of three, while Gemma 4 aborted every single time, even when it was safe to keep going.
On the big fault, GPT-5.5 aborted every time in cruise, but only a third of the time mid-dive.
The study's limits: simulation only, one fault type, one answer per trial.
The paper
"A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies"
Khalid Halba, Kylie Cooper, James G. Bellingham
Exploration Robotics Laboratory, Johns Hopkins Institute for Assured Autonomy
arXiv:2609.20620, September 2026
The research was supported by the Office of Naval Research and JHU's Bloomberg Distinguished Professorships Program.
This is an independent explainer. I'm not affiliated with or endorsed by the authors or Johns Hopkins, and the research is all theirs. Any mistakes in the explanation are mine.
I'm Bora Celik from Piccard. I explain new research in ocean science and robotics.