ECCV, 2026
project page / arXiv / code
SPAR jointly reconstructs photometric and open-vocabulary semantic scene content from sparse, unposed, and dynamic observations without explicit 3D supervision or ground-truth motion masks.
I am a joint M.S. student in Computer Science at ShanghaiTech University and the Institute of Automation, Chinese Academy of Sciences, advised by Weiming Hu, Yan Xu, and Li Yang.
My research focuses on learning generalizable visual representations for understanding open and dynamic environments. Across my work, I have found that language consistently provides semantic structure that strengthens grounding and generalization.
I am currently open to Ph.D. opportunities and industry positions. Please feel free to contact me!
Across these projects, two lessons have shaped how I approach research: (1) end-to-end training often coordinates the full system more effectively than fine-tuning modules in isolation; and (2) well-designed model priors provide crucial structure for learning and generalization.
ECCV, 2026
project page / arXiv / code
SPAR jointly reconstructs photometric and open-vocabulary semantic scene content from sparse, unposed, and dynamic observations without explicit 3D supervision or ground-truth motion masks.
CVPR, 2026
SRA-Det retrieves multiple semantic facets from token-level text features and uses attribute-augmented data to improve fine-grained open-vocabulary detection.
CVPR, 2026
TRE selectively suppresses dominant temporal patterns so spiking neural networks learn more diverse and complementary features across timesteps.
IEEE Transactions on Circuits and Systems for Video Technology, 2025
A guiding multi-task framework addresses imbalanced and conflicting optimization between detection and re-identification in end-to-end person search.
AAAI, 2025
Temporal-Self-Erasing supervision dynamically suppresses previously activated regions, encouraging complementary and more discriminative features over time.