SAFE: A Semantic Audio-Visual Fusion Engine for Interpretable and Real-Time Crowd Anomaly Detection
Ravi Saharan, Akrisht Singh, Prakash Choudhary*
Computer Systems Science and Engineering, Vol.50, pp. 1-20, 2026, DOI:10.32604/csse.2026.081278
- 21 September 2026
Abstract The task of monitoring crowds for safety is a critical challenge, yet traditional surveillance systems are often visual only, error-prone, and lack interpretability. This paper presents SAFE (Semantic Audio-Visual Fusion Engine), a real-time, multi-modal framework that delivers interpretable, operator-facing alerts by fusing complementary audio-visual cues. The visual pipeline couples SSD–ResNet face detection and Hungarian tracking with facial emotion recognition and DBSCAN-based clustering of negative affect to compute a semantic visual anomaly score that includes an explicit overcrowding signal. In parallel, the audio pipeline extracts MFCCs (Mel-Frequency Cepstral Coefficients) and employs a lightweight one-dimensional convolutional neural… More >