Submission Deadline: 31 March 2027 View: 33 Submit to Special Issue
Dr. Fabiana Di Ciaccio
Email: fabiana.diciaccio@unifi.it
Affiliation: Department of Civil and Environmental Engineering (DICEA), University of Florence, Florence, Italy
Research Interests: computer vision, artificial intelligence, foundation models, multimodal learning, vision transformers, remote sensing, digital twins, photogrammetry, environmental monitoring, cultural heritage
Dr. Paolo Russo
Email: paolo.russo@uniroma3.it
Affiliation: Department of Civil Engineering, Computer Science and Aeronautical Technologies, Roma Tre University, Rome, 00146, Italy
Research Interests: deep learning, computer vision, monocular depth estimation, signal processing, biomedical classification
Foundation Models and Multimodal Artificial Intelligence are rapidly transforming the field of Computer Vision, enabling intelligent systems to learn from massive amounts of heterogeneous data and generalize across a wide range of tasks. Recent advances in Vision Transformers, Vision-Language Models, self-supervised learning, and multimodal foundation models have significantly improved visual understanding, scene interpretation, object recognition, and decision-making, opening new opportunities across numerous scientific and industrial domains. These technologies are driving innovation in applications such as remote sensing, autonomous systems, robotics, medical imaging, digital twins, environmental monitoring, smart cities, manufacturing, and cultural heritage.
This Special Issue aims to collect original research articles addressing the latest advances in Foundation Models, Multimodal AI, and Computer Vision. We welcome contributions focusing on novel algorithms, efficient architectures, learning strategies, and real-world applications. The objective is to provide a comprehensive overview of recent developments while promoting innovative solutions capable of advancing the state of the art in intelligent visual systems.
Suggested Themes:
• Foundation Models for Computer Vision
• Vision-Language Models
• Vision Transformers
• Multimodal Learning
• Self-Supervised Learning
• Weakly Supervised Learning
• Large Vision Models
• Image Classification and Recognition
• Image Segmentation
• Object Detection and Tracking
• 3D Computer Vision
• Robotics and Autonomous Systems
• Digital Twins
• Smart Cities
• Industrial Inspection
• Explainable Artificial Intelligence
• Efficient Deep Learning


Submit a Paper
Propose a Special lssue