An Underwater Robot That Predicts Its Own Physics | AquaWAM Explained
Underwater, nothing stops when you tell it to. Cut the thrusters and a small robot keeps gliding for about two seconds; fire them gently and almost nothing happens. That makes a precise grab hard.
Researchers at Shandong University and the Beijing Institute of Technology built AquaWAM, a model that predicts the robot's own motion two seconds ahead, before every command, from its sensors rather than from camera images. In simulation, with the object's exact position handed to both, a hand-tuned controller made a tricky grab less than one time in five. AquaWAM's predicted pulses made it more than four times in five. In this independent explainer, I walk through how it works and what the paper shows.
In the video
- Why water makes control hard: the glide after the thrusters stop, the thrusters' dead band, and currents
- The AI models it's compared against: vision-language-action models, most of which commit to about a second and a half of commands without checking how it's going
- AquaWAM's idea: predict motion, not pictures, from the Doppler velocity log, the inertial sensor, depth and the arm's joints
- 128 candidate futures every half second, each imagined two seconds ahead and scored, with only the first half second of the winner run
- Where the physics comes from: "play", recordings of the robot moving around with no task, where the dead band and the glide show up
- Up close: short thrust pulses chosen by where the robot will come to rest, like a putt in golf
- The results on a 20-task simulator benchmark: success in nearly three trials out of four, against about half for the strongest baseline, U0. Without its play data, or without imagining futures, AquaWAM falls behind U0
- What happens when the velocity sensor loses the bottom
- What the paper doesn't show yet: it's all simulation, and carrying after the velocity sensor fails, a dead thruster and murky water still hurt
- What's next: the authors want to try it on a real remotely operated vehicle
The paper
"AquaWAM: A Dynamics-aware World Action Model for Underwater Embodied Agents"
Cunhao Zhu, Yifeng Wang, Dongliang Xu, Yunzhong Hou, Yue Yao, Chi Harold Liu
Shandong University and Beijing Institute of Technology
arXiv:2609.33299 [cs.RO], September 2026 (preprint)
The paper goes much deeper than this video: how the predictor and its disturbance estimate are built and trained, the full benchmark tables for all seven models, every ablation, the robustness tests, and the timing on an embedded computer. If you work on underwater manipulation or robot learning, it's worth reading in full.
This is an independent explainer. I'm not affiliated with or endorsed by the authors, Shandong University or the Beijing Institute of Technology, and the research is all theirs. The animations are my own illustrations of what the paper describes, not footage from the study: the robot's paths come from a simple physics sketch, and charts marked schematic are illustrations. Any mistakes in the explanation are mine.
I'm Bora Celik from Piccard. I explain new research in ocean science and robotics.