About Me
I build multimodal systems for visual alignment, model composition, and reliable reasoning in changing worlds.
I am currently a Postdoctoral Research Associate at Imperial College London, supervised by Prof. Jiankang Deng and Prof. Stefanos Zafeiriou. Previously, I completed my Ph.D. at Zhejiang University, advised by Prof. Chao Wu and co-supervised by Prof. Kun Kuang and Prof. Fei Wu. I was also fortunate to collaborate with Yibing Song from the Alibaba DAMO Academy.
I am always open to collaboration opportunities and conversations with fellow researchers. Recently in London, I have been collecting sunsets, parks, coffee conversations, and new research questions. 😊
Research Map
Align
Using structure and theory to make vision-language models generalize beyond familiar distributions.
Compose
Understanding how multimodal models forget, merge, and reuse specialized capabilities.
Reason
Eliciting reliable multimodal reasoning through lightweight reinforcement learning and curated supervision.
Recent Highlights
- 2026.04 🚀 Released LLaVA-OneVision-2, our next-generation 8B multimodal foundation model unifying image, long-form video, and spatial understanding, where I served as a core contributor.
- 2025.12 🚀 Released LLaVA-OneVision-1.5-RL, a fully open framework for democratized multimodal reinforcement learning, where I served as a core contributor.
- 2025.09 🎉 One first-author paper has been accepted to NeurIPS 2025 Multimodal Algorithmic Reasoning Workshop.
- 2025.09 🎉 One co-author paper has been accepted to NeurIPS 2025.
- 2025.07 🎉 One co-author paper has been accepted to ICCV 2025.
- 2025.07 🎉 One co-author paper has been accepted to KDD 2025.
Earlier updates
- 2025.05 🎉 Four co-author papers have been accepted to ICML 2025.
- 2025.04 🥳 I will attend the ICLR conference in Singapore, welcome to discuss in person.
- 2025.02 🎉 One co-author paper has been accepted to ICLR 2025 FL Workshop.
- 2025.01 🎉 Three papers (One first-author paper and two co-author papers) have been accepted to ICLR 2025.
- 2024.07 🥳 I went to Vienna, Austria to attend the ICML conference.
- 2024.05 🎉 One first-author paper has been accepted to KDD 2024.
- 2024.05 🎉 One first-author paper has been accepted to ICML 2024.
- 2023.10 🥳 I went to Paris, France to attend the ICCV conference.
- 2023.07 🎉 One first-author paper has been accepted to ICCV 2023.
- 2023.07 🎉 One first-author paper has been accepted to ACM Multimedia 2023.
- 2023.05 🎉 One paper has been accepted to KDD 2023.
- 2022.11 🎉 One paper has been accepted to IEEE Transactions on Big Data.
- 2021.05 🎉 One paper has been accepted to IJCAI 2021 FL Workshop.
📝 Publications

Reason
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence
Xiang An, Yin Xie, Feilong Tang, Yunyao Yan, Huajie Tan, Didi Zhu, Changrui Chen, Xiuwei Zhao, Bin Qin, Kaicheng Yang, et al.
Project Leaders: Bo Li, Ziyong Feng, Ziwei Liu, Zongyuan Ge, Jiankang Deng
- An 8B multimodal foundation model unifying image, long-form video, and 3D-aware spatial understanding under a single architecture via codec-aligned vision encoders.
- Fully open end-to-end release: data, encoders, training pipeline, checkpoints, and training logs.

Reason
LLaVA-OneVision-1.5-RL: Unlocking Multimodal Reasoning via Lightweight Reinforcement Learning
Didi Zhu (First Author, RL Section), Zhiyu Qu, Zerui Chen, Polydefkis Gkagkos, Xiang An, Bo Li
RL Section of Technical Report. Project Leaders: Changrui Chen, Jiankang Deng
- Leveraging lightweight RL framework (GRPO) on top of supervised instruct model to elicit latent reasoning capabilities.
- Using only 67K curated examples with discrepancy-based selection, significantly boosting performance on complex STEM, Coding, and Reasoning tasks.

Compose
REMEDY: Recipe Merging Dynamics in Large Vision-Language Models
.
Didi Zhu, Yibing Song, Tao Shen, Ziyu Zhao, Jinluan Yang, Min Zhang, Chao Wu
- First exploration of the LoRA fusion problem in Multimodal Large Language Models
- Proposing a dynamic fusion scheme enhances zero-shot generalization capability of MLLMs.

Compose
Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models.
Didi Zhu, Zhongyisun Sun, Zexi Li, Tao Shen, Ke Yan, Shouhong Ding, Chao Wu, Kun Kuang
- Pioneered the first comprehensive exploration and revelation of catastrophic forgetting in MLLMs such as InstructBLIP and LLaVa.
- Addressed the issue through an innovative training-free model grafting technique.

Align
Neural Collapse Anchored Prompt Tuning for Generalizable Vision-Language Models.
Didi Zhu, Zexi Li, Min Zhang, Junkun Yuan, Jiashuo Liu, Kun Kuang, Chao Wu
- The first exploration of large vision-language models through the lens of neural collapse in deep learning theory.
- Tackle class imbalance in generalization tasks for large vision-language models by leveraging neural collapse theory.

Align
Universal domain adaptation via compressive attention matching.
Didi Zhu, Yinchuan Li, Junkun Yuan, Zexi Li, Kun Kuang, Chao Wu
- Addressed the issue of inconsistent source-target label spaces in Universal Domain Adaptation directly using self-attention in ViT.

Align
Generalized Universal Domain Adaptation with Generative Flow Networks.
Didi Zhu, Yinchuan Li, Yunfeng Shao, Jianye Hao, Fei Wu, Kun Kuang, Jun Xiao, Chao Wu
- Introduced a comprehensive problem called Generalized Universal Domain Adaptation, achieving a unification of all Domain Adaptation sub-problems involving label heterogeneity.
- Implemented an exploration-aware active learning strategy based on Generative Flow Networks to effectively address GUDA.
Publications by Research Theme
AlignRobust VLMs, domain shift, and generalization
- KDD 2024 Neural Collapse Anchored Prompt Tuning for Generalizable Vision-Language Models, Didi Zhu, Zexi Li, Min Zhang, Junkun Yuan, Jiashuo Liu, Kun Kuang, Chao Wu.
- ICCV 2023 Universal domain adaptation via compressive attention matching, Didi Zhu, Yinchuan Li, Junkun Yuan, Zexi Li, Kun Kuang, Chao Wu.
- ACM MM 2023 Generalized Universal Domain Adaptation with Generative Flow Networks, Didi Zhu, Yinchuan Li, Yunfeng Shao, Jianye Hao, Fei Wu, Kun Kuang, Jun Xiao, Chao Wu.
- ICML 2025 ERICT: Enhancing Robustness by Identifying Concept Tokens in Zero-Shot Vision Language Models, Xinpeng Dong, Min Zhang, Didi Zhu, Ye Jun Jian, Keli Zhang, Aimin Zhou, Fei Wu, Kun Kuang.
- KDD 2025 FedGuCci: Making Local Models More Connected in Landscape for Federated Learning, Zexi Li, Jie Lin, Zhiqi Li, Didi Zhu, Tao Shen, Tao Lin, Chao Wu, Nicholas D. Lane.
- KDD 2023 Quantitatively Measuring and Contrastively Exploring Heterogeneity for Domain Generalization, Yunze Tong, Junkun Yuan, Min Zhang, Didi Zhu, Keli Zhang, Fei Wu, Kun Kuang.
- TBD Towards Effective Clustered Federated Learning: A Peer-to-peer Framework with Adaptive Neighbor Matching, Zexi Li, Jiaxun Lu, Shuang Luo, Didi Zhu, Yunfeng Shao, Yinchuan Li, Zhimeng Zhang, Yongheng Wang, Chao Wu.
- IJCAI WS Ensemble federated adversarial training with non-iid data, Shuang Luo, Didi Zhu, Zexi Li, Chao Wu.
ComposeModel merging, forgetting, and modular reuse
- ICLR 2025 REMEDY: Recipe Merging Dynamics in Large Vision-Language Models, Didi Zhu, Yibing Song, Tao Shen, Ziyu Zhao, Jinluan Yang, Min Zhang, Chao Wu.
- ICML 2024 Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models, Didi Zhu, Zhongyisun Sun, Zexi Li, Tao Shen, Ke Yan, Shouhong Ding, Chao Wu, Kun Kuang.
- NeurIPS 2025 Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model Merging, Jinluan Yang, Dingnan Jin, Anke Tang, Li Shen, Didi Zhu, Zhengyu Chen, Ziyu Zhao, Daixin Wang, Qing Cui, Zhiqiang Zhang, Jun Zhou, Fei Wu, Kun Kuang.
- ICML 2025 ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think, Tao Feng, Wei Li, Didi Zhu, Hangjie Yuan, Wendi Zheng, Dan Zhang, Jie Tang.
- ICML 2025 Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning, Wenke Huang, Jian Liang, Zekun Shi, Didi Zhu, Guancheng Wan, He Li, Bo Du, Dacheng Tao, Mang Ye.
- ICML 2025 Be Confident: Uncovering Overfitting in MLLM Multi-Task Tuning, Wenke Huang, Jian Liang, Guancheng Wan, Didi Zhu, He Li, Jiawei Shao, Mang Ye, Bo Du, Dacheng Tao.
- ICLR 2025 Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace, Jinluan Yang, Anke Tang, Didi Zhu, Zhengyu Chen, Li Shen, Fei Wu.
- ICLR 2025 Merging LoRAs like Playing LEGO: Pushing the Modularity of LoRA to Extremes Through Rank-Wise Clustering, Ziyu Zhao, Tao Shen, Didi Zhu, Zexi Li, Jing Su, Xuwu Wang, Kun Kuang, Fei Wu.
ReasonMultimodal reasoning and perceptual intelligence
- Tech Report 2026 LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence, Xiang An, Yin Xie, Feilong Tang, Yunyao Yan, Huajie Tan, Didi Zhu, Changrui Chen, Xiuwei Zhao, Bin Qin, Kaicheng Yang, et al.
- Tech Report 2025 LLaVA-OneVision-1.5-RL: Unlocking Multimodal Reasoning via Lightweight Reinforcement Learning, Didi Zhu, Zhiyu Qu, Zerui Chen, Polydefkis Gkagkos, Xiang An, Bo Li.
Educations
Ph.D. Student
Computer Science and Technology, Zhejiang University Hangzhou
Undergraduate
Computer Science and Technology, Beijing University of Chemical Technology Beijing
Research Experiences
Postdoctoral Research Associate
Imperial College London London, United Kingdom
Internships
Alibaba DAMO Academy
Hangzhou, China
Tencent Youtu Lab
Shanghai, China
Services
Reviewing
Honors and Awards
National Scholarship
Top 5%
Beijing Outstanding Graduates
Top 1%
National Scholarship
Top 1%
National Scholarship
Top 1%
National Scholarship
Top 1%
Miscellaneous
I recently arrived in London and have fallen in love with the city's sunsets, parks, and ever-changing light. I believe life is a grand experience to be savored, and I am always happy to share a coffee, a walk, or a conversation about what this next chapter might hold.