Multi-Modal and World Model Symposium 2026

August 15th - 16th, 2026
Location: CUHK

CUHK Campus (photo from CUHK website)

Multi-Modal and World Model Symposium 2026

The Multi-Modal and World Model Symposium 2026 brings together leading researchers to discuss the latest advances in Multimodal Large Models, Computer Vision, and World Models, which are increasingly recognized as a key foundation for Physical AI and embodied intelligence.

Hosted by LaVi Lab, Department of Computer Science and Engineering, The Chinese University of Hong Kong (CUHK), the two-day symposium will feature multiple keynote talks, invited presentations, and panel discussions from a distinguished lineup of internationally recognized scholars spanning computer vision, machine learning, multimodal AI, and embodied intelligence. By bringing together experts from diverse yet closely connected research areas, the symposium aims to provide a high-impact platform for exploring cutting-edge developments and emerging trends in next-generation AI.

Join us at CUHK to gain insights from world-class researchers and engage with the latest progress shaping the future of multimodal intelligence and World Models.

Abstract

In recent years, multimodal large models and World Models have witnessed rapid advancements. Among them, World Models—with their ability to infer physical laws and predict dynamic environments—are widely regarded as a critical technological pathway toward Physical AI and embodied intelligence.

While current research routes on World Models are diverse and flourishing, human cognition suggests that our deep understanding of and interaction with the environment largely stem from the continuous input and perceptual accumulation of multimodal information, particularly vision. Consequently, computer vision (CV) and multimodal technologies play a vital role in the construction and evolution of World Models.

For this symposium, we have invited academic experts from various research directions to focus on and discuss the interplay among computer vision, and World Models.

Program Schedule

August 15 (Morning Session)
9:00 - 9:15 Welcome and Opening Speech
9:15 - 10:00 David Forsyth (University of Illinois Urbana-Champaign), Keynote Talk
10:05 - 10:35 Jiajun Wu (Stanford University)
10:40 - 11:10 Xiaojuan Qi (The University of Hong Kong)
11:15 - 11:45 Manling Li (Northwestern University)
August 15 (Afternoon Session)
14:00 - 14:45 Xiaoming Liu (The University of North Carolina at Chapel Hill), Keynote Talk
14:50 - 15:20 Alex Schwing (University of Illinois Urbana-Champaign)
15:25 - 15:55 Jiajun Zhang (University of Chinese Academy of Sciences)
16:00 - 16:30 Xinggang Wang (Huazhong University of Science and Technology)
16:35 - 17:05 David Wipf (The University of Hong Kong)
17:10 - 18:00 Panel Discussion
August 16 (Morning Session)
9:00 - 9:45 Mohit Bansal (The University of North Carolina at Chapel Hill), Keynote Talk
9:50 - 10:20 Mengye Ren (New York University)
10:25 - 10:55 Bolei Zhou (University of California Los Angeles)
11:00 - 11:30 Qifeng Chen (The Hong Kong University of Science and Technology)
August 16 (Afternoon Session)
14:00 - 14:45 Zhuowen Tu (University of California San Diego), Keynote Talk
14:50 - 15:20 Qi Dou (The Chinese University of Hong Kong)
15:25 - 15:55 Fei Miao (Zhejiang University)
16:00 - 16:30 Ping Luo (The University of Hong Kong)
16:35 - 17:05 Long Chen (The Hong Kong University of Science and Technology)
17:10 - 18:00 Liwei Wang (The Chinese University of Hong Kong)
18:00 - 18:15 Closing Remarks

Location

Address and transportation guidance: Coming soon.