Open Access iconOpen Access

ARTICLE

DDGEM: Diffusion Denoising and Generative Enhancement for Multimodal Recommendation

Weiwei Li*, Li Zhao, Chengshan Li, Wenjie Geng

School of Computer Science and Engineering, Chongqing University of Technology, Chongqing, China

* Corresponding Author: Weiwei Li. Email: email

Computers, Materials & Continua 2026, 89(2), 82 https://doi.org/10.32604/cmc.2026.085768

Abstract

In multimodal recommendation, sparse implicit feedback can lead to noisy collaborative graphs, while long-tail and cold-start items often lack reliable collaborative signals. Textual and visual features provide useful item-side information, but they may also be incomplete or inconsistent across modalities. These issues make robust user and item representation learning difficult. To address them, we propose Diffusion Denoising and Generative Enhancement for Multimodal Recommendation (DDGEM). DDGEM first applies node-wise diffusion denoising in the latent collaborative space to reduce unreliable user–item signals. It then uses relational diffusion to reconstruct adaptive item–item relations instead of relying on fixed semantic similarity. For sparse and cold-start items, a conditional variational autoencoder (CVAE)-based conditional generative module produces complementary item representations from textual and visual features, followed by popularity-aware generation balance. Finally, behavior-aware multimodal fusion dynamically combines collaborative, textual, and visual representations for each user–item pair. Experiments on three public Amazon datasets show that DDGEM outperforms representative collaborative filtering, generative recommendation, and multimodal recommendation baselines. Further significance tests, ablation studies, efficiency comparison, graph-noise robustness analysis, strict cold-start evaluation, modality-conflict analysis, and hyperparameter analysis verify the effectiveness and practical behavior of DDGEM.

Keywords

Multimodal recommendation; diffusion denoising; generative enhancement; behavior-aware fusion

Cite This Article

APA Style
Li, W., Zhao, L., Li, C., Geng, W. (2026). DDGEM: Diffusion Denoising and Generative Enhancement for Multimodal Recommendation. Computers, Materials & Continua, 89(2), 82. https://doi.org/10.32604/cmc.2026.085768
Vancouver Style
Li W, Zhao L, Li C, Geng W. DDGEM: Diffusion Denoising and Generative Enhancement for Multimodal Recommendation. Comput Mater Contin. 2026;89(2):82. https://doi.org/10.32604/cmc.2026.085768
IEEE Style
W. Li, L. Zhao, C. Li, and W. Geng, “DDGEM: Diffusion Denoising and Generative Enhancement for Multimodal Recommendation,” Comput. Mater. Contin., vol. 89, no. 2, pp. 82, 2026. https://doi.org/10.32604/cmc.2026.085768



cc Copyright © 2026 The Author(s). Published by Tech Science Press.
This work is licensed under a Creative Commons Attribution 4.0 International License , which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
  • 256

    View

  • 91

    Download

  • 0

    Like

Share Link