TY - EJOU
AU - Makhmudov, Fazliddin
AU - Khamzaev, Jamshid
AU - Karimberdiyev, Jakhongir
AU - Rakhimov, Mekhriddin
AU - Djurayeva, Nigora
AU - Saymanov, Islambek
AU - Almuradov, Abdulla
AU - Kutlimuratov, Alpamis
TI - Video-Based Temporal Attention Network for Automated Analysis of Left Ventricular Ejection Fraction
T2 - Computer Modeling in Engineering \& Sciences
PY - 2026
VL - 148
IS - 3
SN - 1526-1506
AB - The left ventricular ejection fraction (LVEF) is one of the most important quantitative indicators in cardiovascular medicine. It serves as a key measure for diagnosing, assessing risk, and guiding treatment decisions in conditions such as heart failure, cardiomyopathies, and coronary artery disease. However, obtaining accurate LVEF values using echocardiography remains a major challenge. The process is often affected by differences between observers (8%–15%), time-consuming manual measurements that take several minutes per case, dependence on the operator’s skill, and sensitivity to image quality. To address these challenges, this study proposes a Temporal Attention Network (TAN)—a deep learning framework that combines three complementary components to improve the precision and efficiency of LVEF estimation. The model integrates a ResNet-50-based convolutional neural network to extract hierarchical spatial features, a bidirectional long short-term memory (BiLSTM) network with 512 hidden units per direction to learn temporal dependencies across complete cardiac cycles, and a Convolutional Block Attention Module (CBAM) that refines feature representations by highlighting diagnostically important cardiac structures. The proposed network was trained end-to-end using the EchoNet-Dynamic dataset, which includes 10,030 apical four-chamber echocardiography videos with expert-annotated LVEF values. Training utilized the Adam optimizer with a cosine annealing learning rate schedule and extensive data augmentation to enhance model robustness. When evaluated on 1277 unseen test videos representing diverse demographics and cardiac pathologies, the model achieved outstanding performance: a mean absolute error (MAE) of 4.12% (95% CI: 3.89%–4.35%), root mean squared error (RMSE) of 5.67%, and an R2 of 0.847. This marks an 11.3% improvement over previous state-of-the-art approaches (MAE = 4.65%) and is comparable to inter-observer variability among human experts (4%–5%). For classifying heart failure with reduced ejection fraction (LVEF < 40%), the model reached 94.7% accuracy, 92.8% sensitivity, 95.6% specificity, and an AUROC of 0.977. Ablation studies confirmed that temporal modeling reduced errors by 25% compared to frame-based methods, while the attention mechanisms provided an additional 16% improvement. Attention visualization maps consistently focused on the left ventricular region, aligning with expert-identified areas 82% of the time. Subgroup analyses demonstrated consistent performance across age groups (MAE 4.08%–4.34%), sexes (males: 4.18%, females: 4.05%, p = 0.51), and LVEF quartiles (MAE 4.02%–4.31%, p = 0.62), with expected degradation under poor image quality (MAE 6.73%). Overall, the Temporal Attention Network shows strong potential as a clinically applicable tool that can minimize diagnostic variability, streamline echocardiographic workflows, and open new possibilities for automated cardiac function assessment.
KW - EchoNet-Dynamic; ejection fraction; attention mechanism; LSTM; temporal analysis; cardiac function; computer-aided diagnosis; convolutional block attention module; cardiovascular assessment
DO - 10.32604/cmes.2026.085457