My research aims to develop AI systems that understand complex visual environments and bridge vision and language. My interests include Video Understanding, Visual Question Answering, and Visual Grounding. Recently, I focus on Video Understanding and Multimodal LLM for real-world and healthcare applications, with an emphasis on efficiency, interpretability, and trustworthy AI. My goal is to build intelligent systems that understand the real world, support human decision-making, and enable effective human–AI interaction. I am actively seeking Research Interns/Visiting Students with backgrounds in CV/NLP/Multimodal.
* Equal contribution; # Core contributor; † Corresponding author
ECCV
CVPR
CVPR
TMM
TOMM
TIP
Last updated: 2026.10
Powered by Jekyll and Minimal Light theme.