TY - EJOU
AU - My, Linh Nguyen Thi
AU - Minh, Nhat Trinh Ngoc
AU - Vo, Tham
AU - Hoang, Vinh Truong
TI - MIRA-Pill: Multimodal Imprint Reading and Confidence-Gated Re-Ranking for Fine-Grained Pill Identification
T2 - Computers, Materials \& Continua
PY -
VL -
IS -
SN - 1546-2226
AB - Automated pill identification is a safety-critical task whose hardest cases involve visually near-identical medications that differ only in their printed or debossed imprint code. The widely used ePillID benchmark explicitly identified the reliable reading of such imprints as the most important open problem, but left it unsolved because the optical character recognition available at the time was not dependable on small, low-contrast pill surfaces. We revisit this problem with modern vision–language models (VLMs) and propose MIRA-Pill (Multimodal Imprint Reading and Confidence-Gated Re-Ranking), a framework that couples (i) a visual metric-learning branch (ResNet50 with Compact Bilinear Pooling) with (ii) a dedicated imprint-reading branch built on a compact VLM (Qwen2-VL-2B) fine-tuned with Low-Rank Adaptation, and fuses the two through a confidence-gated, log-linear product-of-experts re-ranking rule. The fusion is derived from a conditional-independence assumption and is accompanied by two provable properties: a no-harm guarantee when the imprint cannot be read, and a precise rescue condition that explains which errors the imprint corrects. Fine-tuning lifts imprint token recall on unseen pill types by 19.5 percentage points (pp) on hard consumer images. On ePillID (four cross-validation rounds over a fixed held-out test fold, both-sides protocol), MIRA-Pill raises Top-1 accuracy from 90.34% to 94.33% ( pp), a gain that is statistically significant at every test we apply (McNemar ; bootstrap 95% confidence interval (CI) of Top-1 pp), whereas a prior single-benchmark improvement of pp was not significant. The method generalises across datasets: pp on CURE with an OCR-consensus imprint database, and no aggregate degradation on the imprint-free OGYEIv2 dataset, where only 5 of 480 queries are hurt. The size of the gain correlates with imprint availability, confirming that the imprint signal, rather than incidental tuning, drives the improvement, while the confidence gate keeps the observed harm rate low when no imprint is present. For comparison, a zero-shot CLIP baseline reaches only 10.54% Top-1 under the same protocol, so the gain stems from the specialised reader and fusion rather than from generic vision–language capability. These results indicate that imprint reading, long considered unreliable, is now a dependable source of identification evidence on retrospective benchmarks; because the framework has not yet been evaluated in a real-world clinical or pharmacy setting, we position it as a promising candidate component, rather than a validated tool, for future automated pharmacy-verification and medication-safety systems.
KW - Pill identification; fine-grained recognition; vision–language model; imprint reading; multimodal fusion; metric learning; ePillID
DO - 10.32604/cmc.2026.087689