
I am a research scientist specializing in generative AI at the Institute of Artificial Intelligence (TeleAI), China Telecom. I received my Ph.D. from the University of Technology Sydney in 2024, advised by Dr. Linchao Zhu and Prof. Yi Yang.
I received a Master's degree from Xi'an Jiaotong University in 2020 and was a member of the SMILES LAB, advised by Prof. Xueming Qian and Prof. Li Zhu.
My research focuses on visual learning systems that generalize beyond fixed training distributions and continue to improve after pre-training. The work studies both the content of learned representations—discriminative visual features, semantic relationships, temporal dependencies, and scene state—and the evidence used to update them, including labels, preference data, reward estimates, and observed state transitions. This agenda spans visual recognition, video generation, world models, and visual post-training.
My long-term goal is to build agents that can predict the consequences of candidate actions, execute an action, compare predicted and observed state transitions, and use the discrepancy to improve their models or policies. I study this closed loop in digital environments, where interaction can be repeated and scaled, and in physical environments, where actions must satisfy geometric, dynamical, and safety constraints.
I am always looking for highly motivated research interns and long-term collaborators. We currently have multiple positions available, focusing on, but not limited to, multimodal large models, video generation/editing, and 3D generation. If you are interested in exploring these areas or discussing potential research collaborations, please feel free to contact me via email. (Applicants for internships are encouraged to include your CV.)
Note *: interns that I mentored. Note †: equal contribution.