Sheng Zhou (周晟)

Postdoctoral Researcher

King Abdullah University of Science and Technology

[Email] [Github] [Google Scholar]

About Me
Research

My research aims to develop AI systems that understand complex visual environments and bridge vision and language. My interests include video understanding, question answering, and visual grounding. Recently, I focus on egocentric video understanding and multimodal large language models for real-world and healthcare applications, with an emphasis on efficiency, interpretability, and trustworthy AI. My goal is to build intelligent systems that understand the real world, support human decision-making, and enable effective human–AI interaction. I am actively seeking Research Interns/Research Assistants/Visiting Students with backgrounds in CV/NLP/Multimodal.

My group at KAUST is actively looking for fully funded visiting students/interns (free housing, flight ticket, medical insurance, plus 1000 USD/month for stipend). If you are interested, please contact me.

🔥News
Publications

* Equal contribution; # Core Contributor; † Corresponding author

safeguard.png SafeGuard: A Multi-Agent Perception-Reasoning Framework for Social-Risk AI-Generated Video Detection
Wenlin Wu*, Sheng Zhou*, Peipei Song, Wenhao Wang, Junbin Xiao†, Xun Yang†.
ECCV'26 [Paper] [Code] [Dataset]
realunify.png RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
Yang Shi#, Yuhao Dong#, Yue Ding#, Yuran Wang#, Xuanyu Zhu#, Sheng Zhou#, Wenting Liu#, Haochen Tian#, Rundong Wang#, Huanqian Wang, Zuyan Liu, Bohan Zeng, Ruizhe Chen, Qixun Wang, Zhuoran Zhang, Xinlong Chen, Chengzhuo Tong, Bozhou Li, Qiang Liu, Haotian Wang†, Wenjing Yang, Yuanxing Zhang†, Pengfei Wan, YiFan Zhang†, Ziwei Liu†.
CVPR'26 [Paper] [Code] [Dataset]
egotextvqa.png EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
Sheng Zhou, Junbin Xiao†, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, Angela Yao.
CVPR'25 [Paper] [Project Page] [Code] [Dataset]
vitxtgqa Scene-Text Grounding for Text-Based Video Question Answering
Sheng Zhou, Junbin Xiao†, Xun Yang†, Peipei Song, Dan Guo†, Angela Yao, Meng Wang, Tat-Seng Chua.
IEEE TMM'25 [Paper] [Code] [Dataset]
gpin Graph Pooling Inference Network for Text-based VQA
Sheng Zhou, Dan Guo†, Xun Yang†, Jianfeng Dong, Meng Wang†.
ACM TOMM'24 [Paper] [Code]
ssgn Exploring Sparse Spatial Relation in Graph Inference for Text-Based VQA
Sheng Zhou, Dan Guo†, Jia Li, Xun Yang†, Meng Wang†.
IEEE TIP'23 [Paper] [Code]
Selected Honors and Awards
  • [2025.02] Tat-Seng Chua Scholarship
  • [2022 - 2025] First Class Academic Scholarship (three times)
  • [2020 - 2022] Second Class Academic Scholarship (two times)
  • [2020] Outstanding Graduate of Innovation and Entrepreneurship in Hunan Province
  • [2016 - 2019] National Encouragement Scholarship (three times)
Services
  • Reviewer for Conference: NeurIPS (26), CVPR (26), ECCV (26), ACM MM (25/26), EMNLP (26), IJCNN (25/26/27)
  • Reviewer for Journal: IEEE TIP, IEEE TCSVT, ACM TOMM, Information Fusion, Pattern Recognition, Neurocomputing, Journal of King Saud University Computer and Information Sciences, ...