Home / Journals / CMES / Online First / doi:10.32604/cmes.2026.086003
Special Issues
Table of Content

Open Access

ARTICLE

MALT-Drive: Moment-Aligned Language-to-Trajectory Planning for Autonomous Driving

Ziheng Lu1, Yingfeng Cai1,*, Wei Dong1, Hai Wang2,*, Long Chen1
1 Automotive Engineering Research Institute, Jiangsu University, Zhenjiang, China
2 School of Automotive and Traffic Engineering, Jiangsu University, Zhenjiang, China
* Corresponding Author: Yingfeng Cai. Email: email; Hai Wang. Email: email
(This article belongs to the Special Issue: Multimodal Vision with Large Language Models)

Computer Modeling in Engineering & Sciences https://doi.org/10.32604/cmes.2026.086003

Received 22 May 2026; Accepted 21 July 2026; Published online 10 August 2026

Abstract

Vision-language-action models have recently gained attention in autonomous driving, as language can express high-level behavioral intent beyond geometric waypoints. However, existing planners often treat language as an external command or auxiliary explanatory signal, while the hidden states of generated control instructions are rarely aligned explicitly with the latent variables that produce executable trajectories. This paper presents MALT-Drive, a moment-aligned language-to-trajectory end-to-end planning framework for autonomous driving. Given front-camera visual input, ego-state history, route priors, and a driving-oriented VQA prompt, MALT-Drive first autoregressively generates a concise control instruction from a controlled instruction space and then decodes trajectory-level actions. To make the instruction actionable, planning is decomposed into a Maneuver Branch for path geometry and a Modulation Branch for speed regulation and executability estimation. For each branch, generated control-instruction tokens and planning tokens are represented by compact moment descriptors that combine first-order pooled statistics with grouped second-order relation sketches, thereby supporting distribution-aware alignment between language semantics and latent planning states. To reduce visual imitation shortcuts, MALT-Drive further constructs counterfactual control-instruction rollouts, in which the same scene is paired with multiple executable and non-executable canonical control instructions and corresponding motion targets. Experiments on Bench2Drive and NAVSIM show that MALT-Drive improves closed-loop driving performance and provides cross-benchmark evidence of instruction-action consistency.

Keywords

Autonomous driving; vision-language-action models; multimodal alignment; motion planning
  • 61

    View

  • 13

    Download

  • 0

    Like

Share Link