Open Access iconOpen Access

ARTICLE

Multi-View Latent Imitation Learning with Mamba-Based Action Encoding for Unmanned Surface Vehicle Navigation

Manh-Tuan Ha1, Nhu-Nghia Bui2, Dinh-Quy Vu1,*, Thai-Viet Dang2,*

1 Department of Vehicle and Energy Conversion Engineering, School of Mechanical Engineering, Hanoi University of Science and Technology, Hanoi, Vietnam
2 Department of Mechatronics, School of Mechanical Engineering, Hanoi University of Science and Technology, Hanoi, Vietnam

* Corresponding Authors: Dinh-Quy Vu. Email: email; Thai-Viet Dang. Email: email

(This article belongs to the Special Issue: Intelligent Perception, Decision-making and Security Control for Unmanned Systems in Complex Environments)

Computers, Materials & Continua 2026, 88(1), 47 https://doi.org/10.32604/cmc.2026.078280

Abstract

The development of Unmanned Surface Vehicles (USVs) has become a key focus in marine robotics, fueling the need for navigation systems capable of performing complex and delicate tasks with speed and precision. However, the end-to-end path tracking process often encounters challenges in learning efficiency, and generalization, and varying environmental conditions. To achieve sample-efficient and robust USV navigation in dynamic maritime environments, the paper proposes a novel hierarchical multi-view latent imitation learning (IL) architecture. By formulating a latent IL objective, the framework disentangles diverse navigation modalities through continuous variables, preventing mode collapse and enhancing behavioral adaptability to non-stationary conditions. High-dimensional multi-view observations are transformed via a ViT-based backbone into compressed latent features to minimize redundant environmental information. These representations are processed by a Mamba-based action encoder, which leverages selective state-space modeling to capture long-term temporal dependencies with high computational efficiency. A UNet-based decoder subsequently forecasts optimal action sequences by synthesizing spatial maps to infer critical environment-agent relationships. This preliminary multi-view latent IL-based trajectory ensures precise tracking and dynamic stability while adhering to physical vehicle constraints. Experimental results validate that this end-to-end approach achieves robust path planning effectiveness, obstacle avoidance capability, and model training efficiency in complex, multi-modal maritime scenarios.

Keywords

Latent imitation learning; unmanned surface vehicles; latent space model; state-space models; multi-modal feature fusion

Cite This Article

APA Style
Ha, M., Bui, N., Vu, D., Dang, T. (2026). Multi-View Latent Imitation Learning with Mamba-Based Action Encoding for Unmanned Surface Vehicle Navigation. Computers, Materials & Continua, 88(1), 47. https://doi.org/10.32604/cmc.2026.078280
Vancouver Style
Ha M, Bui N, Vu D, Dang T. Multi-View Latent Imitation Learning with Mamba-Based Action Encoding for Unmanned Surface Vehicle Navigation. Comput Mater Contin. 2026;88(1):47. https://doi.org/10.32604/cmc.2026.078280
IEEE Style
M. Ha, N. Bui, D. Vu, and T. Dang, “Multi-View Latent Imitation Learning with Mamba-Based Action Encoding for Unmanned Surface Vehicle Navigation,” Comput. Mater. Contin., vol. 88, no. 1, pp. 47, 2026. https://doi.org/10.32604/cmc.2026.078280



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 109

    View

  • 22

    Download

  • 0

    Like

Share Link