Special Issues
Table of Content

Foundation Models and Multimodal AI for Computer Vision: Algorithms and Real-World Applications

Submission Deadline: 31 March 2027 View: 33 Submit to Special Issue

Guest Editor(s)

Dr. Fabiana Di Ciaccio

Email: fabiana.diciaccio@unifi.it

Affiliation: Department of Civil and Environmental Engineering (DICEA), University of Florence, Florence, Italy

Homepage:

Research Interests: computer vision, artificial intelligence, foundation models, multimodal learning, vision transformers, remote sensing, digital twins, photogrammetry, environmental monitoring, cultural heritage


Dr. Paolo Russo

Email: paolo.russo@uniroma3.it

Affiliation: Department of Civil Engineering, Computer Science and Aeronautical Technologies, Roma Tre University, Rome, 00146, Italy

Homepage:

Research Interests: deep learning, computer vision, monocular depth estimation, signal processing, biomedical classification


Summary

Foundation Models and Multimodal Artificial Intelligence are rapidly transforming the field of Computer Vision, enabling intelligent systems to learn from massive amounts of heterogeneous data and generalize across a wide range of tasks. Recent advances in Vision Transformers, Vision-Language Models, self-supervised learning, and multimodal foundation models have significantly improved visual understanding, scene interpretation, object recognition, and decision-making, opening new opportunities across numerous scientific and industrial domains. These technologies are driving innovation in applications such as remote sensing, autonomous systems, robotics, medical imaging, digital twins, environmental monitoring, smart cities, manufacturing, and cultural heritage.


This Special Issue aims to collect original research articles addressing the latest advances in Foundation Models, Multimodal AI, and Computer Vision. We welcome contributions focusing on novel algorithms, efficient architectures, learning strategies, and real-world applications. The objective is to provide a comprehensive overview of recent developments while promoting innovative solutions capable of advancing the state of the art in intelligent visual systems.


Suggested Themes:
• Foundation Models for Computer Vision
• Vision-Language Models
• Vision Transformers
• Multimodal Learning
• Self-Supervised Learning
• Weakly Supervised Learning
• Large Vision Models
• Image Classification and Recognition
• Image Segmentation
• Object Detection and Tracking
• 3D Computer Vision
• Robotics and Autonomous Systems
• Digital Twins
• Smart Cities
• Industrial Inspection
• Explainable Artificial Intelligence
• Efficient Deep Learning


Keywords

foundation models, multimodal artificial intelligence, computer vision, vision transformers, vision-language models, self-supervised learning, deep learning, 3D vision, image analysis, real-world applications

Share Link