A World Model for Underwater Salvage: C3-JEPA Explained
A robot lies on the seabed, and another robot has been sent down to bring it back. In the last few meters, each camera sees only part of the target, and the ROV answers every push on the joystick late.
Researchers at Shanghai Jiao Tong University built C3-JEPA, a small world model that tracks just the target and the gripper across all of the ROV's cameras, and predicts how they'll move under the pilot's commands. In this independent explainer, I walk through how it works and what the paper shows.
In the video
- Tracking two things instead of mapping the whole scene
- Filling in a hidden camera's view from the other cameras
- Predicting the next few seconds and ranking candidate moves: in simulation, it picked the best of four approaches 53% of the time, vs. 25% by chance
- Real test-pool recordings: running on its own predictions, its error was about 30% lower than "assume nothing moves"
- The limits the authors spell out: the pool runs were replays, and it hasn't been tested at sea yet
- My take (not tested in the paper): nothing in the recipe needs a tether, so autonomous vehicles could use the same idea
The paper
"Underwater C³-JEPA: An Object-Centric Cross-View World Model for ROV Salvage"
Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
Underwater Engineering Institute, Shanghai Jiao Tong University
arXiv:2609.30214, September 2026
The research was supported by the National Key Research and Development Program of China (Grant 2023YFC2809701). The paper goes much deeper than this video: the architecture, the training setup, the full results and the ablations. If you work on underwater manipulation or world models, it's worth reading in full.
This is an independent explainer. I'm not affiliated with or endorsed by the authors or Shanghai Jiao Tong University, and the research is all theirs. The animations are my own illustrations of what the paper describes, not footage from the study. Any mistakes in the explanation are mine.
I'm Bora Celik from Piccard. I explain new research in ocean science and robotics.