Multimodal Learning
Learning unified representations across vision, language, and diverse real-world signals.
Hello, I’m
Ph.D. Student · Multimodal AI Researcher
I study how multimodal foundation models learn, reason, and align. I am a Ph.D. student at Shanghai AI Lab, jointly trained with the School of Artificial Intelligence, Shanghai Jiao Tong University, under the supervision of Yuhang Zang.

Research focus
My work sits at the intersection of representation learning, model efficiency, and alignment.
Learning unified representations across vision, language, and diverse real-world signals.
Improving visual understanding through scalable pre-training and effective data design.
Aligning foundation models with reliable rewards, preferences, and human intent.

Featured work · 2026
A unified framework for dense visual captioning and multimodal post-training.
Our latest work explores a scalable approach to stronger perception and alignment in multimodal large language models.
Read the paperUpdates
Received the Beijing Outstanding Graduate and Tsinghua Excellent Graduate awards.
Our work CapRL++ was submitted to TPAMI.
Started my internship at Shanghai AI Lab
Received the National Scholarship.
Won Second Prize in Tsinghua’s AGI Competition for a RAG project.
Background
Shanghai AI Lab × Shanghai Jiao Tong University
Shanghai Artificial Intelligence Laboratory
Building Environment & Energy Engineering, Tsinghua University
Recognition
Beyond the lab
Outside research, you’ll find me playing basketball or billiards, training at the gym, and exploring cafés around the city.