Junchen Huo1, Song Wang1,*, Enqing Chen1, Yingqiang Ding1, Shouyi Yang2
CMC-Computers, Materials & Continua, Vol.88, No.3, 2026, DOI:10.32604/cmc.2026.081658
- 23 July 2026
Abstract Multi-modal 3D object detection, which leverages the complementary strengths of LiDAR point clouds and camera RGB images, has emerged as a critical component of 3D perception in autonomous driving. As a critical challenge in multi-modal learning, modality alignment aims to establish accurate semantic correspondences across distinct modalities. However, existing methods encounter significant difficulties in achieving robust alignment when data from one modality is obscured, such as in the presence of object occlusion or adverse environmental conditions, including illumination variations and inclement weather. To alleviate this issue, we present CG-MAE, a dual-branch Bird’s-Eye-View (BEV) masked autoencoder… More >