Open Access iconOpen Access

ARTICLE

Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation

Oyku Berfin Mercan1,2, Nezihe Turhan Turan3, Aytuğ Onan4,*

1 R&D Department, Türk Telekom, Ankara, Türkiye
2 Department of Computer Engineering, İzmir Katip Celebi University, İzmir, Türkiye
3 Department of Engineering Sciences, İzmir Katip Celebi University, İzmir, Türkiye
4 Department of Computer Engineering, İzmir Institute of Technology, İzmir, Türkiye

* Corresponding Author: Aytuğ Onan. Email: email

Computers, Materials & Continua 2026, 89(2), 77 https://doi.org/10.32604/cmc.2026.087209

Abstract

Biomedical texts naturally contain multiple biological and medical concepts within a document, resulting in a semantically rich and complex structure. Consequently, multi-label text classification (MLTC) has become a suitable framework for comprehensively modeling biomedical texts, including clinical reports, laboratory records, and scientific abstracts. However, relying solely on contextual language representations may be insufficient to explicitly reflect the broader scientific focus and conceptual orientation of a document. In this study, the MLTC problem in the biomedical domain is investigated using the Hallmarks of Cancer (HoC) dataset. Topic probability distributions obtained from CombinedTM are incorporated as an additional representational signal into pre-trained language models (PLMs), enabling the joint exploitation of contextual semantic information and document-level topical characteristics. Both general-purpose and biomedical language models were fine-tuned and evaluated within this unified framework. The experimental results demonstrate that incorporating topic information substantially improves the domain robustness of the models. In particular, the Macro-F1 score of the general-purpose BERT model increases from 0.7604 to 0.8458, indicating that it becomes competitive with biomedical language models on tasks that require biomedical domain knowledge. Similarly, biomedical models such as BioMed-RoBERTa also benefit from topic modeling, with Macro-F1 improving from 0.8665 to 0.8819; meanwhile, the highest overall performance is achieved by the PubMedBERT + CombinedTM configuration, with a Macro-F1 score of 0.8900. These findings indicate that integrating topic modeling–based thematic information with PLMs provides an effective, generalizable, and practical solution for biomedical MLTC tasks. Moreover, enriching general-purpose language models with topic information offers a promising alternative that reduces reliance on costly, data-intensive domain adaptation.

Keywords

Biomedical multi-label text classification; pre-trained language model; topic modeling; contextualized topic models; CombinedTM; hallmarks of cancer

Cite This Article

APA Style
Mercan, O.B., Turhan Turan, N., Onan, A. (2026). Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation. Computers, Materials & Continua, 89(2), 77. https://doi.org/10.32604/cmc.2026.087209
Vancouver Style
Mercan OB, Turhan Turan N, Onan A. Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation. Comput Mater Contin. 2026;89(2):77. https://doi.org/10.32604/cmc.2026.087209
IEEE Style
O. B. Mercan, N. Turhan Turan, and A. Onan, “Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation,” Comput. Mater. Contin., vol. 89, no. 2, pp. 77, 2026. https://doi.org/10.32604/cmc.2026.087209



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 163

    View

  • 50

    Download

  • 0

    Like

Share Link