TY - EJOU
AU - Ospan, Assel
AU - Mansurova, Madina
AU - Sailau, Aisha
AU - Sarsembayeva, Talshyn
AU - Mosavi, Amir
TI - Parameter Efficient Large Language Models for Dual Regime Text to Table Generation
T2 - Computer Modeling in Engineering \& Sciences
PY - 2026
VL - 148
IS - 3
SN - 1526-1506
AB - This study addresses the automated conversion of unstructured Kazakh journalistic text into structured tabular representations. Existing text-to-table approaches are commonly developed for high-resource languages, rely on fixed schemas, or require computationally expensive full-model fine-tuning. These limitations reduce their applicability to morphologically rich low-resource languages and heterogeneous news collections. We propose a morphology-aware dual-regime framework for Kazakh text-to-table generation. The framework extends a Text-Tuple-Table pipeline with semantic chunking based on cosine-similarity thresholds, thematic grouping using -means clustering, anchor-based factual cue detection, and large language model-based table generation. It supports both a static regime with a predefined five-column schema and a dynamic regime in which table headers and structure are inferred from the source article. A training resource was constructed from 149,624 articles published by Egemen Qazaqstan between 2017 and 2025. It contains 35,000 labeled text–table pairs, including 5000 manually annotated pairs and 30,000 semi-automatically generated pairs. A Qwen3.5-4B model was adapted using Low-Rank Adaptation under single-GPU constraints. Evaluation on a 1,000-record human-validated benchmark shows an in-domain trade-off between factual richness and structural regularity. The dynamic regime achieved higher Coverage (0.718), Accuracy (0.761), and supplementary Journalistic Value (0.932), whereas the static regime achieved higher Compression (0.969) and Structure (0.989). Ablation results further indicate that mixed supervision, thematic clustering, and Kazakh-specific preprocessing contribute to factual extraction quality. The findings demonstrate that combining morphology-aware processing, dual-regime schema generation, and parameter-efficient adaptation provides a practical approach to structured extraction from Kazakh journalistic text. Publicly available implementation materials and pretrained adapters support further reproducibility-oriented research in Kazakh and other low-resource language settings.
KW - Text-to-table generation; Kazakh language; large language models; low-resource NLP; low-rank adaptation; dual-regime schema generation
DO - 10.32604/cmes.2026.086051