Yuncong Yang

I am currently a third-year PhD student at UMass Amherst, where I am supervised by Prof. Chuang Gan and collaborate with Prof. Yilun Du. Since June 2026, I have been a research intern at Meta MSL. Previously, I graduated from Columbia's Fu Foundation School of Engineering with both an M.S. and a B.S. in Computer Science (Summa Cum Laude). I was fortunate to work under the supervision of Prof. Shih-Fu Chang at Columbia University and with Dr. Jim Fan, Prof. Yuke Zhu, and Prof. Anima Anandkumar at NVIDIA Research.

My research interests lie in the area of Spatial Intelligence, Embodied AI, and Multi-modal Foundation Models.

CV  /  Google Scholar  /  Twitter  /  Github

Yuncong Yang
News
Older news
Research

(* indicates equal contribution)

SyncWorld: Visual Calibration Enables World Models as Zero-Shot Simulators
Yuncong Yang*, Zhengtao Han*, Furkan Özyurt*, Zeyuan Yang, Han Yang, Junyi Cao, Haoyu Zhen, Yilun Du, Chuang Gan
arXiv, 2026
project page / paper / code / model

We introduce SyncWorld, a visually calibrated world model that predicts the visual outcomes of numerical robot controls in unseen cameras, environments, and embodiments without downstream training.

MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
Yuncong Yang*, Jiageng Liu*, Zheyuan Zhang, Siyuan Zhou, Reuben Tan, Jianwei Yang, Yilun Du, Chuang Gan
NeurIPS 2025
project page / paper / code

We proposed MindJourney, a test-time scaling framework that improves spatial reasoning with the assistance of controllable world models.

Learning 3D Persistent Embodied World Models
Siyuan Zhou, Yilun Du, Yuncong Yang, Lei Han, Peihao Chen, Dit-Yan Yeung, Chuang Gan
NeurIPS 2025
paper / code

3D-Mem: 3D Scene Memory for Embodied Exploration and Reasoning
Yuncong Yang*, Han Yang*, Jiachen Zhou, Peihao Chen, Hongxin Zhang, Yilun Du, Chuang Gan
CVPR 2025
project page / paper / code / twitter

We proposed 3D-Mem, a framework that serves as 3D scene memory to empower embodied agents with lifelong exploration and reasoning abilities in 3D environments.

TempCLR: Temporal Alignment Representation with Contrastive Learning
Yuncong Yang*, Jiawei Ma*, Shiyuan Huang, Long Chen, Xudong Lin, Guangxing Han, Shih-Fu Chang
ICLR 2023
paper / code

We proposed TempCLR, a new contrastive learning framework that considers sequence-level temporal order consistency in Long-Video Understanding.

MineDojo: Building Open-Ended Embodied Agents with Internet-Scale Knowledge
Linxi Fan, Guanzhi Wang*, Yunfan Jiang*, Ajay Mandlekar, Yuncong Yang, Haoyi Zhu, Andrew Tang, De-An Huang, Yuke Zhu, Animashree Anandkumar
NeurIPS 2022 Datasets and Benchmarks Track   (Outstanding Paper Award, Featured Paper Presentation)
project page / paper / code

We introduce MineDojo, a new framework based on the popular Minecraft game for building generally capable, open-ended embodied agents.

Few-Shot End-to-End Object Detection via Constantly Concentrated Encoding across Heads
Jiawei Ma, Guangxing Han, Shiyuan Huang, Yuncong Yang, Shih-Fu Chang
ECCV 2022
paper

Teaching
Teaching Assistant: COMS 4732 Computer Vision II (Spring 2022)

Thanks for the template from Jon Barron!