Boyu Cai

I am a joint M.S. student in Computer Science at ShanghaiTech University and the Institute of Automation, Chinese Academy of Sciences, advised by Weiming Hu, Yan Xu, and Li Yang.

My research focuses on learning generalizable visual representations for understanding open and dynamic environments. Across my work, I have found that language consistently provides semantic structure that strengthens grounding and generalization.

I am currently open to Ph.D. opportunities and industry positions. Please feel free to contact me!

Research

Across these projects, two lessons have shaped how I approach research: (1) end-to-end training often coordinates the full system more effectively than fine-tuning modules in isolation; and (2) well-designed model priors provide crucial structure for learning and generalization.