Open Access
ARTICLE
Mitigating Visual Noise in Multimodal AI: Selective Visual Grounding for Multimodal Machine Translation
1 Designovel Lab, Designovel, Pohang, Republic of Korea
2 Department of Convergence IT Engineering, POSTECH, Pohang, Republic of Korea
3 School of Information Convergence, Kwangwoon University, Seoul, Republic of Korea
* Corresponding Author: Kyudong Park. Email:
(This article belongs to the Special Issue: Explainable Multimodal AI: Interpretability, Generative Modeling, and Trustworthy Intelligent Systems)
Computer Modeling in Engineering & Sciences 2026, 148(1), 39 https://doi.org/10.32604/cmes.2026.083410
Received 03 April 2026; Accepted 11 June 2026; Issue published 27 July 2026
Abstract
Multimodal AI systems often suffer from “over-informing”, where excessive raw visual input introduces noise that distracts from task-relevant decisions. Motivated by selective human attention strategies, we propose ARS-MMT (Attention and Reasoning through Source Sentences for Multimodal Machine Translation), an architecture that operationalizes a “look-and-think” pipeline: a source-language encoder first builds contextualized linguistic representations, a relation reasoning network then produces a query-conditioned visual channel, and a multimodal decoder generates the translation conditioned in parallel on the encoded text and on this visual channel. We quantify the contribution of the visual modality through a controlled ablation: zeroing visual features reduces BLEU by 0.81 on test_2016_flickr En-De, while shuffling visual features across the batch changes BLEU by onlyKeywords
Cite This Article
Copyright © 2026 The Author(s). Published by Tech Science Press.This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.


Submit a Paper
Propose a Special lssue
View Full Text
Download PDF
Downloads
Citation Tools