Home / Journals / CMC / Online First / doi:10.32604/cmc.2026.084514
Special Issues
Table of Content

Open Access

ARTICLE

Deep Hand Segmentation and Multi-Modal Gesture Recognition for Human-Robot Interaction via 3D Volumetric Encoding

Zarnab Kausar1, Shaheryar Najam2, Hadeel Alsolai3, Bayan Alabdullah3, Fatimah Alhayan3, Ahmad Jalal4,5,*, Hui Liu6,7,8,*
1 Department of Engineering Technology, Foundation University, Islamabad, Pakistan
2 Department of Electrical Engineering, Bahira University, Islamabad, Pakistan
3 Department of Information Systems, College of Computer and Information Sciences, Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia
4 Department of Computer Science, Air University, E-9, Islamabad, Pakistan
5 Department of Computer Science and Engineering, College of Informatics, Korea University, Seoul, Republic of Korea
6 Guodian Nanjing Automation Co., Ltd., Nanjing, China
7 Jiangsu Key Laboratory of Intelligent Medical Image Computing, School of Artificial Intelligence (School of Future Technology), Nanjing University of Information Science and Technology, Nanjing, China
8 Cognitive Systems Lab, University of Bremen, Bremen, Germany
* Corresponding Author: Ahmad Jalal. Email: email; Hui Liu. Email: email
(This article belongs to the Special Issue: Robotics Vision and Thinking)

Computers, Materials & Continua https://doi.org/10.32604/cmc.2026.084514

Received 23 April 2026; Accepted 17 June 2026; Published online 10 September 2026

Abstract

Hand gesture recognition (HGR) is essential for Human–Robot Interaction (HRI) but remains challenging due to variations in hand shape, motion, viewpoint, illumination, and background, while vision-based methods often suffer from sensitivity to skin tone, occlusions, deformations, and limited interpretability. To address these issues, we propose a unified framework integrating deep learning, geometry-driven analysis, and temporal motion modeling. We introduce Z-HandSegNet framework, involving a U-Net with a ResNet-34 encoder for robust hand segmentation, and the Ellipse-Guided Geometric Finger Segmentation and Keypoint Extraction (EG-FSKE) method, which decomposes hand silhouettes into palm and finger regions using distance transforms, Gaussian Mixture Models, ellipse fitting, and a fragment recovery strategy for occluded and deformed fingers. To capture the hierarchical structure and dense motion, adaptive octree-based volumetric representation and Ef-RAFT optical flow are used for dynamic modeling. These representations are combined with saliency and ridge-based descriptors into a multimodal feature and classified with a Transformer encoder for modelling temporal dependencies. The evaluations on three datasets (Jester, IPN Hand and EgoGesture) achieve state-of-the-art performance with accuracies of 93.22%, 92.34%, and 94.17%, respectively, and precision, recall and F1-scores above 0.87. Ablation studies confirm the contribution of each component, particularly Z-HandSegNet and EG-FSKE, highlighting the robustness, interpretability, and potential of the proposed approach for future real-time HRI applications.

Keywords

Human–robot interaction; hand gesture recognition; hand segmentation; geometric finger segmentation; optical flow; 3D volumetric representation; multi-modal feature fusion; transformer-based temporal modeling
  • 97

    View

  • 20

    Download

  • 0

    Like

Share Link