[Paper Review] Efficient collective swimming by harnessing vortices through deep reinforcement learning
This study uses deep reinforcement learning (DRL) to train autonomous swimmers that optimize collective propulsion by synchronizing with vortex wakes shed by a leader fish. The smart follower achieves up to 100% improvement in swimming efficiency by intercepting vortices at precise phase-locked positions, demonstrating that fish can harvest energy from hydrodynamic flows to reduce energy expenditure without compromising speed or stability.
Fish in schooling formations navigate complex flow-fields replete with mechanical energy in the vortex wakes of their companions. Their schooling behaviour has been associated with evolutionary advantages including collective energy savings. How fish harvest energy from their complex fluid environment and the underlying physical mechanisms governing energy-extraction during collective swimming, is still unknown. Here we show that fish can improve their sustained propulsive efficiency by actively following, and judiciously intercepting, vortices in the wake of other swimmers. This swimming strategy leads to collective energy-savings and is revealed through the first ever combination of deep reinforcement learning with high-fidelity flow simulations. We find that a `smart-swimmer' can adapt its position and body deformation to synchronise with the momentum of the oncoming vortices, improving its average swimming-efficiency at no cost to the leader. The results show that fish may harvest energy deposited in vortices produced by their peers, and support the conjecture that swimming in formation is energetically advantageous. Moreover, this study demonstrates that deep reinforcement learning can produce navigation algorithms for complex flow-fields, with promising implications for energy savings in autonomous robotic swarms.
Motivation & Objective
- To investigate whether fish can reduce energy expenditure by exploiting hydrodynamic vortices in the wake of conspecifics.
- To develop autonomous navigation strategies for swimmers that adapt to unsteady flow fields using reinforcement learning.
- To quantify the energetic benefits of coordinated swimming through high-fidelity fluid dynamics simulations.
- To reveal the physical mechanisms enabling energy extraction from vortex wakes in collective locomotion.
- To demonstrate the feasibility of DRL in discovering optimal, biologically plausible swimming strategies in complex fluid environments.
Proposed method
- Deep reinforcement learning (DRL) with Long Short-Term Memory (LSTM) networks is used to train self-propelled swimmers to learn optimal swimming policies from visual flow cues.
- High-fidelity direct numerical simulations (DNS) of the incompressible Navier-Stokes equations model the 2D flow field around two tandem swimmers (leader and follower) with realistic fish-like body deformations.
- Two distinct DRL agents are trained: IS η (efficiency-focused) and IS d (position-stabilization-focused), each with a custom reward function based on swimming efficiency or lateral deviation.
- The DRL agents learn policies through trial-and-error in a simulated environment, using state observations of local flow velocity and vorticity to make real-time decisions.
- Baseline control cases (Solitary Swimmers SS η and SS d) are used to isolate the energetic benefits arising from wake interactions.
- Energy metrics such as swimming efficiency (η), thrust-power (PThrust), deformation-power (PDef), and cost of transport (CoT) are computed and compared across configurations.
Experimental results
Research questions
- RQ1Can autonomous swimmers learn to improve swimming efficiency by actively interacting with vortex wakes generated by a leader fish?
- RQ2What physical mechanisms underlie the energy savings observed in collective swimming, particularly in relation to vortex synchronization?
- RQ3How does the choice of reward function in reinforcement learning influence the emergence of efficient swimming postures and trajectories?
- RQ4To what extent can a follower adapt its behavior to unsteady, complex flow fields without prior knowledge of the leader’s motion?
- RQ5What role does temporal memory (via LSTM) play in enabling stable, energy-efficient navigation in dynamic vortex environments?
Key findings
- The DRL-trained follower (IS η) achieves a swimming efficiency of η ≈ 1.0, representing a 100% improvement over the leader’s efficiency, by synchronizing its head motion with the lateral flow velocity in the vortex wake.
- IS η naturally settles at ∆x ≈ 2.2L behind the leader, a position that aligns with the periodic shedding of vortex rings, and also stabilizes at ∆x ≈ 1.5L, corresponding to the vortex spacing (0.7L apart).
- The follower’s body deformation is minimal during optimal vortex interception, indicating that energy savings arise from flow exploitation rather than increased muscular effort.
- Despite no direct reward for position, IS η maintains a stable lateral position (∆y ≈ 0) by leveraging temporal memory (LSTM), demonstrating robust adaptation to dynamic flow fields.
- The follower’s thrust-power is significantly enhanced in the mid-body region (0.2 < s/L < 0.4) due to favorable vortex interactions, while deformation-power remains low, confirming efficient energy harvesting.
- Even when the leader’s motion becomes erratic, the trained follower (IS η) autonomously adapts to remain in the wake and maximize long-term efficiency, proving generalization capability.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.