My research focuses on developing AI systems that can understand complex visual environments and effectively bridge vision and language. My research interests include Video Understanding, Visual Question Answering, and Visual Grounding.
I currently focus on Video Understanding and Multimodal Large Language Models (MLLMs) for real-world and healthcare applications, with an emphasis on efficiency, interpretability, and trustworthy AI. My goal is to build intelligent systems that can perceive and reason about the real world, support human decision-making, and enable effective human–AI interaction.
I am actively seeking Research Interns and Visiting Students with backgrounds in Computer Vision, NLP, or Multimodal Learning.
* Equal contribution; # Core contributor; † Corresponding author
ECCV
CVPR
CVPR
TMM
TOMM
TIP
Last updated: 2026.10
Powered by Jekyll and Minimal Light theme.