Open Access iconOpen Access

ARTICLE

A Novel Entropy-Based Framework for Hybrid Sampling in Imbalanced Learning

Ren-Jieh Kuo1,*, Muhammad Rizki1, Ferani Eva Zulvia2, Eddy Roflin3

1 Department of Industrial Management, National Taiwan University of Science and Technology, Taipei, Taiwan
2 Department of Industrial Engineering, Institut Teknologi Bandung, Bandung, Indonesia
3 Faculty of Medicine, Sriwijaya University, Palembang, Indonesia

* Corresponding Author: Ren-Jieh Kuo. Email: email

Computers, Materials & Continua 2026, 89(1), 90 https://doi.org/10.32604/cmc.2026.084436

Abstract

Imbalanced data remain a critical challenge in classification, as skewed distributions bias models toward majority classes and diminish sensitivity to minority classes, which are often the most critical. To address this issue, this paper proposes the Information Filtered Hybrid Algorithm (IF-HA), a novel entropy-based sampling method that integrates undersampling and oversampling guided by information theory. IF-HA quantifies instance importance through an instance-wise difference statistic. In the undersampling stage, majority of instances with low difference statistics in the border area are eliminated, while in the oversampling stage, synthetic samples are generated from two minority core points or two minority instances with high difference statistics located in the border area. This process removes noise, eliminates redundant majority border points, and generates synthetic minority samples in informative regions until an entropy-based imbalance threshold is reached. The proposed algorithm is evaluated on 20 benchmark datasets from the UCI and KEEL repositories. Results demonstrate that IF-HA consistently improves minority detection and achieves higher F1 Scores, recall, and AUC (Area Under the Curve) than other methods, including SMOTE, Borderline-SMOTE, ADASYN (Adaptive Synthetic Sampling), and SMOTE-TLNN-DEPSO. A real-world tuberculosis (TB) dataset from Indonesia was further used to validate the practical applicability of IF-HA using KNN, Random Forest, and XGBoost (eXtreme Gradient Boosting) classifiers. The results show consistent improvements after applying IF-HA. These findings indicate that entropy-based hybrid sampling is a promising approach for structured tabular imbalanced classification, while further validation on high-dimensional text and image datasets remains necessary to establish broader generalizability.

Keywords

Imbalanced data; hybrid sampling; information theory; entropy-based algorithm; minority class detection; tuberculosis classification

Cite This Article

APA Style
Kuo, R., Rizki, M., Zulvia, F.E., Roflin, E. (2026). A Novel Entropy-Based Framework for Hybrid Sampling in Imbalanced Learning. Computers, Materials & Continua, 89(1), 90. https://doi.org/10.32604/cmc.2026.084436
Vancouver Style
Kuo R, Rizki M, Zulvia FE, Roflin E. A Novel Entropy-Based Framework for Hybrid Sampling in Imbalanced Learning. Comput Mater Contin. 2026;89(1):90. https://doi.org/10.32604/cmc.2026.084436
IEEE Style
R. Kuo, M. Rizki, F. E. Zulvia, and E. Roflin, “A Novel Entropy-Based Framework for Hybrid Sampling in Imbalanced Learning,” Comput. Mater. Contin., vol. 89, no. 1, pp. 90, 2026. https://doi.org/10.32604/cmc.2026.084436



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 151

    View

  • 38

    Download

  • 0

    Like

Share Link