About
I am a Ph.D. student in Computer Science at the School of Data Science, The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), advised by Prof. Kui Jia. Previously, I worked with Prof. Guyue Zhou at the Institute for AI Industry Research (AIR), Tsinghua University. I also conducted research at King Abdullah University of Science and Technology (KAUST) under the supervision of Prof. Mohamed Elhoseiny and worked as a Research Intern at DexForce Technology. Before that, I received my B.Eng. in Automation with honors from Harbin Institute of Technology, Weihai.
My research focuses on building embodied intelligence for perceiving and interacting with the physical world. In particular, I am interested in developing generalizable robotic manipulation policies and learning world models that capture how the physical world appears, behaves, and evolves.
News
- [2026.06] Our team won first place in the WorldArena Challenge at CVPR 2026. Champion
- [2026.06] One paper was accepted to IROS 2026.
- [2026.04] YOTO++ was accepted to IEEE TPAMI.
- [2026.03] EVA was released and open-sourced.
- [2025.10] One paper was accepted to IEEE Robotics and Automation Letters (RA-L).
- [2025.09] I started my Ph.D. at The Chinese University of Hong Kong, Shenzhen.
- [2025.06] GAT-Grasp was accepted to IROS 2025.
- [2025.06] One paper was accepted to ICCV 2025.
- [2025.04] YOTO was accepted to RSS 2025.
Education
Experience
Publications
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
arXiv preprint, 2026
EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards
arXiv preprint, 2026
YOTO++: Learning Long-Horizon Closed-Loop Bimanual Manipulation from One-Shot Human Video Demonstrations
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
Robotics: Science and Systems (RSS), 2025
Vid2WAM: Distilling Video Diffusion Priors into World Action Models
arXiv preprint, 2026
EVA: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards
arXiv preprint, 2026
YOTO++: Learning Long-Horizon Closed-Loop Bimanual Manipulation from One-Shot Human Video Demonstrations
IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI), 2026
You Only Teach Once: Learn One-Shot Bimanual Robotic Manipulation from Video Demonstrations
Robotics: Science and Systems (RSS), 2025
Research Interests
Robot Manipulation
Generalizable policies for reliable, long-horizon manipulation.
World Models
Predictive models that remain physically grounded and executable.
Embodied Intelligence
Connecting perception, imagination, and action in the physical world.