TUHH mid-summer talk · delivered 2026-07-20

Teaching Robots to Swim

Replicating WHOI’s “Learning to Swim” reinforcement-learning pipeline (ICRA 2025) in simulation — scroll through the full deck.

01 Title dark

Reinforcement learning · underwater control

Teaching Robots
to Swim

From WHOI’s “Learning to Swim” method to TUHH’s underactuated HippoCampus.

Kyle Nelson  ·  TU Hamburg  ·  Summer 2026

02 Meet HippoCampus — TUHH’s micro-AUV dark · video slot

Meet HippoCampus

TUHH’s micro-AUV

Palm-sized & low-cost · built to swim in swarms

A 3D-printed, open-source micro submarine developed at TU Hamburg — agile enough for quadrotor-style maneuvers in a test tank.

Why it exists: one is cheap enough to build a fleet — swarms that monitor confined and shallow waters, map conditions, and track down sources of pollution.

Video: Duecker, Bauschmann, Hansen, Kreuzer & Seifried — “Towards Micro Robot Hydrobatics: An Integrated Framework for Guidance, Navigation, and Control for Agile Underwater Vehicles” · Hamburg University of Technology

03 How it swims today light

The Goal

Control System Overhaul

Current Control Laws

Attitude controllerhand-tuned
Depth controllerhand-tuned
Per-maneuver controllerone per task
One neural network
that learns to swim
trained in simulation · full 6-DOF pose

Current system: geometric drone-style attitude control on PX4 — Duecker et al., ICRA 2018

04 Meet CUREE — WHOI’s research AUV dark · the data vehicle

Meet CUREE

How the network is trained

1 · Randomize the robot, every episode Center of Buoyancy displaced randomly by 5 cm. Volume shifted by ±3 L (±13 %).
2 · 2 048 robots practise at once Each with its own random physics — only what works everywhere survives.
3 · One network, no hand-tuning Pose in → six thruster commands out; those outputs fly the real robot.
CUREE vector infographic

Fully actuated · holds any 6-DOF pose

F = k · |ω| · ω
motor physics equation · k = 0.001 N·s²/rad²

CUREE — Curious Underwater Robot for Ecosystem Exploration · WHOI WARPLab · arXiv:2410.00120

05 The campaigns — what & why dark · setup

Reproducing WHOI’s Results

The campaigns · what we’re doing and why

Neural Net Reward function

r  =  0.5 · e−|Δrot|  +  0.2 · e−‖Δpos‖²  +  0.2 · e−‖a‖²

termweight λimpact
Rotation0.5facing the goal
Position0.2sitting on the goal
Effort0.2not wasting thrust

CUREE code is open source. Can our own training on CUREE match the authors’ results?

The exam: hold 12 goal poses · 6 position · 6 rotation.

Four knobs: random start · training length · attempt length · randomization.

06 Learning 1 — the random-start lottery dashboard · dark

The seed lottery

Learning 1 of 4 · same setup, three random starts

Average miss over the exam’s two halves position m · rotation rad · mean of 6 goals each

Identical 2 500-round training, only the random start (“seed”) differs — a 1.6× rotation spread, repeatable, not noise.

07 Learning 2 — training length dashboard · dark

Longer training gained nothing

Learning 2 of 4 · Training Length

Scored at 2 500 vs 10 000 training rounds rotation error · rad²

The same three runs scored early vs late: 4× more training bought nothing.

08 Learning 3 — the trade-off dashboard · dark

Trading rotation for position

Learning 3 of 4 · more time & more randomization

What each knob does to the two scores mean of 3 paired runs · m² / rad²

Each knob against its own matched control — both help position and hurt rotation.

09 Learning 4 — reward ≠ skill dashboard · dark

Reward ≠ skill

Learning 4 of 4 · why we don’t trust the training score

The extended recipe: training score vs the exam reward · rotation error rad²

Longer attempts offer more reward points to collect — the score rose while skill fell.

10 Conclusion — the verdict light · conclusion

Results ≈2.5× WHOI’s

Conclusion · our best swimmers vs the authors’ released network

Best swimmer per campaign, both scores m² / rad²

Our best swimmer lands at 1.26× the pass line — the leftover rotation gap points at simulator physics, not the training.

11 Meet the two robots dark · generated robots in place

TUHH Implementation

Meet the vehicles

BlueROV2 Heavy vector infographic

BlueROV2 Heavy

8 thrusters · fully actuated

The validator — closely matches the paper.

HippoCampus vector infographic

HippoCampus

4 thrusters · underactuated

TUHH’s micro-AUV — the real challenge.

HippoCampus — TU Hamburg (hippocampusrobotics)

12 Where it’s going light · teal paths + image slots

Three ways to finish

The road ahead

waypoint following — HippoCampus and goal dot
Path 01

Follow waypoints

Reframe the task to route-following — something four thrusters can actually do.

position hold with free heading — goal at body center, rotation rings
Path 02

Hold position free rotation

Reproduce the WHOI result, minus the fixed-orientation requirement.

thruster-belt augmented HippoCampus
Path 03

Add a thruster belt

Bolt on four angled thrusters → fully actuated → run the paper’s exact task.

13 Thank you dark · contact

Thank you

Questions & contact

TUHH main building
Kyle Nelson
Kyle NelsonUC Berkeley
QR — kyle-nelson-berkeley.vercel.app
Nathalie Bauschmann
Nathalie BauschmannAdvisor · TU Hamburg
QR — tuhh.de Nathalie Bauschmann

Notes. Preview order (01 title · 02 meet HippoCampus · 03 the goal · 04 CUREE · 05–08 behavioral findings · 09 the two robots · 10 roadmap · 11 thank you). Mint on the white slide is nudged to #00A98F for legible text on white — the poster/graphic mint #5AFFC5 is fine as a fill but too light for type. Everything here is CSS/SVG stand-in; the real deck is built in Claude Design from the prompts in your plan file.