Guest Editor(s)
Prof. Ki-Il Kim
Email: kikim@cnu.ac.kr
Affiliation: Department of Computer Science and Engineering, Chungnam National University, Daejeon, Republic of Korea
Homepage:
Research Interests: machine learning, internet of things

Dr. Babar Shah
Email: babar.shah@zu.ac.ae
Affiliation: College of Technological Innovation, Zayed University, Abu Dhabi, United Arab Emirates
Homepage:
Research Interests: AI-enabled security, network & IoT systems, cloud/edge computing

Summary
Human action recognition is a key technology for intelligent systems that perceive, understand, and interact with humans and dynamic environments. While conventional approaches primarily rely on a single modality such as RGB video, their performance can degrade under illumination changes, occlusion, viewpoint variations, sensor noise, and other real-world uncertainties. Multimodal action recognition addresses these limitations by integrating complementary information from RGB, depth, skeleton, audio, infrared, radar, IMU, wearable, and other sensing modalities.
This Special Issue aims to present recent advances in multimodal action recognition and understanding, with particular emphasis on intelligent and automated systems. It welcomes innovative research on multimodal representation and fusion, transformers and graph neural networks, self-supervised and contrastive learning, vision-language and multimodal foundation models, action anticipation, temporal action understanding, and robust recognition under missing or noisy modalities. Research on lightweight, explainable, privacy-preserving, and edge-based approaches is also encouraged.
Applications include intelligent surveillance, human-robot interaction, autonomous systems, smart manufacturing, sports analysis, and smart environments. The Special Issue seeks to advance multimodal action recognition from conventional classification toward robust, contextual, and semantic human activity understanding.
Keywords
multimodal action recognition, human activity understanding, multimodal learning, multimodal fusion, vision-language models