Open Access
ARTICLE
Enhancing Biomedical Multi-Label Text Classification via Topic-Based Text Representation
1 R&D Department, Türk Telekom, Ankara, Türkiye
2 Department of Computer Engineering, İzmir Katip Celebi University, İzmir, Türkiye
3 Department of Engineering Sciences, İzmir Katip Celebi University, İzmir, Türkiye
4 Department of Computer Engineering, İzmir Institute of Technology, İzmir, Türkiye
* Corresponding Author: Aytuğ Onan. Email:
Computers, Materials & Continua 2026, 89(2), 77 https://doi.org/10.32604/cmc.2026.087209
Received 12 June 2026; Accepted 07 August 2026; Issue published 15 September 2026
Abstract
Biomedical texts naturally contain multiple biological and medical concepts within a document, resulting in a semantically rich and complex structure. Consequently, multi-label text classification (MLTC) has become a suitable framework for comprehensively modeling biomedical texts, including clinical reports, laboratory records, and scientific abstracts. However, relying solely on contextual language representations may be insufficient to explicitly reflect the broader scientific focus and conceptual orientation of a document. In this study, the MLTC problem in the biomedical domain is investigated using the Hallmarks of Cancer (HoC) dataset. Topic probability distributions obtained from CombinedTM are incorporated as an additional representational signal into pre-trained language models (PLMs), enabling the joint exploitation of contextual semantic information and document-level topical characteristics. Both general-purpose and biomedical language models were fine-tuned and evaluated within this unified framework. The experimental results demonstrate that incorporating topic information substantially improves the domain robustness of the models. In particular, the Macro-F1 score of the general-purpose BERT model increases from 0.7604 to 0.8458, indicating that it becomes competitive with biomedical language models on tasks that require biomedical domain knowledge. Similarly, biomedical models such as BioMed-RoBERTa also benefit from topic modeling, with Macro-F1 improving from 0.8665 to 0.8819; meanwhile, the highest overall performance is achieved by the PubMedBERT + CombinedTM configuration, with a Macro-F1 score of 0.8900. These findings indicate that integrating topic modeling–based thematic information with PLMs provides an effective, generalizable, and practical solution for biomedical MLTC tasks. Moreover, enriching general-purpose language models with topic information offers a promising alternative that reduces reliance on costly, data-intensive domain adaptation.Keywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools