About Me

Hi there, I am Shiyu Hu (胡世宇)!

I am a Research Fellow in the School of Physical and Mathematical Sciences (SPMS), Nanyang Technological University (NTU), working with Assoc. Prof. Kang Hao Cheong.

My research centers on reliable multimodal intelligence in open and interactive environments. I study how AI systems maintain object identity, world state, task-relevant evidence, and interaction context when observations unfold over time or humans participate. My current work spans open-world vision, multimodal reasoning, and human-centered agents.

I received my Ph.D. from the Institute of Automation, Chinese Academy of Sciences (CASIA) in January 2024 under the supervision of Prof. Kaiqi Huang and Prof. Xin Zhao. Before that, I completed my M.Sc. in Computer Science at the University of Hong Kong (HKU) with Prof. Choli Wang.

📣 I welcome research discussions and collaborations. Please contact me at shiyu.hu@ntu.edu.sg.

🔥 News

2026.06: 📝One paper (DASTrack) has been accepted by the 2026 European Conference on Computer Vision (ECCV, CCF-B Conference).

2026.05: 📣I am glad to have the opportunity to give a talk at the Chinese Congress on Image and Graphics (CCIG 2026) in Guangzhou, China. Many thanks to the special session Intelligent Evolution of Video and Image Security: Perception, Reasoning, and Adversarial Challenges. I sincerely look forward to exchanging ideas with everyone and hearing your valuable suggestions😄

2026.05: 🏆I received a Reviewer Award from the 43rd International Conference on Machine Learning (ICML 2026).

2026.04: 📝One paper (RGRL) has been accepted by the main conference of the 64th Annual Meeting of the Association for Computational Linguistics (ACL, CCF-A Conference).

2026.04: 📝One research paper has been accepted by the 48th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC, CAAI-B Conference).

2026.04: 📝One paper (COAL) has been accepted by the 35th International Joint Conference on Artificial Intelligence (IJCAI, CCF-B Conference).

2026.03: 📝One paper (TBDQ) has been accepted by the Pattern Recognition (PR, CCF-B Journal).

2026.02: 📝One paper (EARL) has been accepted by the 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR, CCF-A Conference).

2026.02: 📝One paper (MATrack) has been accepted by the 2026 IEEE International Conference on Robotics & Automation (ICRA, CCF-B Conference).

2026.02: 📝One paper (DAAWBench) has been accepted by the IEEE Transactions on Circuits and Systems for Video Technology (TCSVT, CCF-B Journal).

2026.02: 📝One review paper has been accepted by the IEEE Transactions on Network Science and Engineering (TNSE).

2026.01: 📣We presented two main-conference Oral papers (📹 CausalStep Slides and 📹 VerifyBench Slides) and three workshop papers (📹 SOEI Slides, 📹 EduVerse Slides, and 📹 EduPersona Slides) at AAAI 2026. Thanks to everyone who visited us at the Singapore EXPO for the discussions.

2026.01: 📝One paper (NarrLV) has been accepted by the 14th International Conference on Learning Representations (ICLR, CCF-A Conference).

2025.12: 📣We will conduct a Mini-Symposium (topic: Complex Network Systems and Large Language Models) on NODYCON 2026 (The Fifth International Nonlinear Dynamics Conference), more information will be released soon.

2025.11: 📝Three papers have been accepted by the AI for Education Workshop in the 40th Annual AAAI Conference on Artificial Intelligence (AAAIW).

2025.11: 📝Two papers (CausalStep and VerifyBench) have been accepted by the 40th Annual AAAI Conference on Artificial Intelligence (AAAI, CCF-A Conference, Oral).

2025.10: 📣We have conducted a tutorial at 28th European Conference on Artificial Intelligence (ECAI) (26th October, 2025, Bologna, Italy).

2025.10: 📣We have conducted a tutorial at 2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC) (5th October, 2025, Vienna, Austria).

2025.08: 📣We have conducted a tutorial at 34th International Joint Conference on Artificial Intelligence (IJCAI) (18th August, 2025, Montreal, Canada).

2025.07: 🏆Obtain IEEE SMCS TEAM Program Award.

2025.06: 📝One paper (ATCTrack) has been accepted by International Conference on Computer Vision (ICCV, CCF-A conference, Highlight).

▶️For More News

👩‍💻 Experiences

2024.08 - Present : Research Fellow, School of Physical and Mathematical Sciences (SPMS), Nanyang Technological University (NTU)

2018.03 - 2018.11 : Research Assistant at University of Hong Kong (HKU)

  • Direction: High Performance Computing, Heterogeneous Computing
  • PI: Prof. Choli Wang

2016.08 - 2016.09 : Research Intern at Institute of Electronics, Chinese Academy of Sciences (CASIE)

📖 Educations

2019.09 - 2024.01 : Ph.D. at the Institute of Automation, Chinese Academy of Sciences (CASIA)

  • Field: Computer Application Technology
  • Supervisor: Prof. Kaiqi Huang (IAPR Fellow, IEEE Senior Member, 10,000 Talents Program - Leading Talents)
  • Co-supervisor: Prof. Xin Zhao (IEEE Senior Member, Beijing Science Fund for Distinguished Young Scholars)
  • Thesis Title: Research of Intelligence Evaluation Techniques for Single Object Tracking
  • Thesis Committee: Prof. Jianbin Jiao, Prof. Yuxin Peng (The National Science Fund for Distinguished Young Scholars), Prof. Yao Zhao (IEEE Fellow, IET Fellow, The National Science Fund for Distinguished Young Scholars), Prof. Yunhong Wang (IEEE Fellow, IAPR Fellow, CCF Fellow), Prof. Ming Tang
  • Thesis Defense Grade: Excellent

2017.09 - 2019.07 : M.Sc. in the Department of Computer Science, University of Hong Kong (HKU)

  • Field: Computer Science
  • Supervisor: Prof. Choli Wang
  • Thesis Title: NightRunner: Deep Learning for Autonomous Driving Cars after Dark [🌐Project]
  • Thesis Defense Grade: A+

2013.09 - 2017.07 : B.E. in the Elite Class, School of Information and Electronics, Beijing Institute of Technology (BIT)

  • Field: Information Engineering
  • Undergraduate Thesis Supervisor: Prof. Senlin Luo
  • Thesis Title: Text Sentiment Analysis Based on Deep Neural Network
  • Thesis Defense Grade: Excellent

2015.07 - 2015.08 : Summer Session at the University of California, Berkeley (UC Berkeley)

  • Course: New Media
  • Course Grade: A

🔍️ Research Interests

Research Foundation

Research trajectory from visual perception tasks to human-grounded intelligence evaluation

My research began with visual object tracking and machine vision evaluation. I have studied task modeling, evaluation environments, measurement techniques, and human-machine comparison. Inspired by the Turing Test, I proposed the Visual Turing Test to evaluate dynamic visual intelligence against human abilities.

The central idea has remained consistent: AI evaluation should reveal what a system can perceive, understand, and reliably maintain in real environments, rather than report a benchmark score alone.


1️⃣ What abilities define human perception?

I used Visual Object Tracking as a representative task for studying dynamic visual ability. Traditional tracking assumes continuous motion and short-term observation. Global Instance Tracking (GIT) extends the task to long-term target retrieval, while Multi-modal GIT (MGIT) introduces hierarchical semantics and spatiotemporal reasoning. This work moved my research from perceptual localization toward cognitive visual tasks.


2️⃣ What environments do humans perceive?

Human visual environments are continuous, open, and semantically rich. I developed VideoCube to organize long videos through narrative structure, and SOTVerse as an open task space for testing visual generalization. BioDrone further examines reliable perception under motion disturbance and real-world physical constraints.


3️⃣ How large is the human-machine gap?

I constructed unified evaluation settings in which people and models perform comparable visual tasks. The results show different strengths: people use semantics and context more effectively, while machines often sustain precision and persistence. These comparisons shaped my interest in human-grounded evaluation and reliable human-AI collaboration.


Current Research Directions

Open-World Vision

I study dynamic visual perception and world-state modeling under occlusion, interference, and physical disturbance. Current topics include object tracking, vision-language grounding, long-horizon visual memory, and reliable perception for UAVs and embodied systems.

Multimodal Reasoning

I investigate whether multimodal foundation models select and use the right evidence in images and long videos. My work covers spatiotemporal and causal reasoning, streaming memory, adaptive visual computation, and process-level verification.

Human-Centered Agents

I develop agents that model cognitive and learning states, capability boundaries, and social interaction. Education provides a practical setting for personalized agents, virtual students, multi-agent simulation, and reliable human-AI collaboration.

The 3E framework connecting environment, evaluation, and executors

The 3E framework connects Environment, Evaluation, and Executors in one evaluation loop. Across the three directions above, I build open task spaces, human-grounded protocols, and process-level diagnostics to identify capability boundaries and failure mechanisms, then use those findings to improve model and system design.

📝 Publications

Book

Springer 2025
sym

Visual Object Tracking: An Evaluation Perspective
X. Zhao, Shiyu Hu, X. Yin
Springer, Part of the book series: Advances in Computer Vision and Pattern Recognition (ACVPR)
📌 Visual Object Tracking 📌 Intelligent Evaluation Technology
📃 Book

Accept

First Author / Corresponding Author

TPAMI 2023
sym

Global Instance Tracking: Locating Target More Like Humans
Shiyu Hu, X. Zhao, L. Huang, K. Huang
IEEE Transactions on Pattern Analysis and Machine Intelligence (CCF-A Journal)
📌 Visual Object Tracking 📌 Large-scale Benchmark Construction 📌 Intelligent Evaluation Technology
📃 Paper 📑 PDF 🪧 Poster 🌐 Platform 🔧 Toolkit 💾 Dataset

IJCV 2024
sym

SOTVerse: A User-defined Task Space of Single Object Tracking
Shiyu Hu, X. Zhao, K. Huang
International Journal of Computer Vision (CCF-A Journal)
📌 Visual Object Tracking 📌 Dynamic Open Environment Construction 📌 3E Paradigm
📃 Paper 📑 PDF 🪧 Poster 🌐 Platform

IJCV 2024
sym

BioDrone: A Bionic Drone-based Single Object Tracking Benchmark for Robust Vision
X. Zhao, Shiyu Hu✉️, Y. Wang, J. Zhang, Y. Hu, R. Liu, H. Lin, Y. Li, R. Li, K. Liu, J. Li
International Journal of Computer Vision (CCF-A Journal)
📌 Visual Object Tracking 📌 Drone-based Tracking 📌 Visual Robustness
📃 Paper 🌐 Platform 📑 PDF 🔧 Toolkit 💾 Dataset

NeurIPS 2023
sym

A Multi-modal Global Instance Tracking Benchmark (MGIT): Better Locating Target in Complex Spatio-temporal and causal Relationship
Shiyu Hu, D. Zhang, M. Wu, X. Feng, X. Li, X. Zhao, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
📌 Visual Language Tracking 📌 Long Video Understanding and Reasoning 📌 Hierarchical Semantic Information Annotation
📃 Paper 📃 PDF 🪧 Poster 📹 Slides 🌐 Platform 🔧 Toolkit 💾 Dataset

ICCV 2025
sym

ATCTrack: Aligning Target-Context Cues with Dynamic Target States for Robust Vision-Language Tracking
X. Feng*, Shiyu Hu*, X. Li, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang (*Equal Contributions)
International Conference on Computer Vision (CCF-A Conference, Highlight)
📌 Visual Language Tracking 📌 Multimodal Learning 📌 Adaptive Prompts
📃 Paper 📑 PDF

ICRA 2026
sym

MATrack: Efficient Multiscale Adaptive Tracker for Real-Time Nighttime UAV Operations
X. Li*, X. Li*, Shiyu Hu✉️
International Conference on Robotics and Automation (CCF-B Conference)
📌 Nighttime UAVs Tracking 📌 Multiscale Adaptive Tracker 📌 Visual Object Tracking
📃 Paper 📑 PDF

ICMR 2025
sym

DARTer: Dynamic Adaptive Representation Tracker for Nighttime UAV Tracking
X. Li*, X. Li*, Shiyu Hu✉️
International Conference on Multimedia Retrieval (CCF-B Conference)
📌 Nighttime UAVs Tracking 📌 Dark Feature Blending 📌 Dynamic Feature Activation
📃 Paper 📑 PDF

中国图象图形学报 2024
sym

Visual Intelligence Evaluation Techniques for Single Object Tracking: A Survey (单目标跟踪中的视觉智能评估技术综述)
Shiyu Hu, X. Zhao, K. Huang
Journal of Images and Graphics (《中国图象图形学报》, CCF-B Chinese Journal)
📌 Visual Object Tracking 📌 Intelligent Evaluation Technique 📌 AI4Science
📃 Paper 📑 PDF

IET-CVI 2025
sym

Improved SAR Aircraft Detection Algorithm Based on Visual State Space Models
Y. Wang, J. Zhang, Y. Wang, Shiyu Hu✉️, B. Shen, Z. Hou, W. Zhou
IET Computer Vision (CCF-C Journal)
📌 Synthetic Aperture Radar 📌 State Space Models 📌 Aircraft Object Detection

Collaborator (Arranged in Chronological Order)

CVPR 2026
sym

Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
X. Li*, X. Li*, Shiyu Hu, K. Huang
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CCF-A Conference)
📌 Video Large Language Models 📌 Video Reasoning 📌 Video Understanding
📃 Paper 📑 PDF

AAAI 2026
sym

CausalStep: A Benchmark for Explicit Stepwise Causal Reasoning in Videos
X. Li*, X. Li*, Shiyu Hu, K. Huang, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
📌 Video-based QA 📌 Video Reasoning 📌 Video Understanding
📃 Paper 📑 PDF 📹 Slides

AAAI 2026
sym

VerifyBench: A Systematic Benchmark for Evaluating Reasoning Verifiers Across Domains
X. Li*, X. Li*, Shiyu Hu, Y. Guo, W. Zhang
Proceedings of the AAAI Conference on Artificial Intelligence (CCF-A Conference, Oral)
📌 Verifiable Reward 📌 Reinforcement Learning
📃 Paper 📑 PDF 📹 Slides

ICLR 2026
sym

NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation Models
X. Feng, H. Yu, M. Wu, Shiyu Hu, J. Chen, C. Zhu, J. Wu, X. Chu, K. Huang
International Conference on Learning Representations (CCF-A Conference)
📌 Visual Understanding 📌 Video Generation 📌 Evaluation Technique
📃 Paper 📑 PDF

ACL 2026
sym

Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
X. Li*, X. Li*, J. Gao, R. Pi, Shiyu Hu, W. Zhang
Annual Meeting of the Association for Computational Linguistics (CCF-A Conference)
📌 Thinking-with-Image 📌 Vision-Language Models 📌 Pixel Reasoning
📃 Paper 📑 PDF

PR 2026
sym

Tracking by Detection and Query: An Efficient End-to-End Framework for Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Cao, F. Yang, X. Lu, X. Lu
Pattern Recognition (CCF-B Journal)
📌 Multi-object Tracking 📌 Tracking by Detection 📌 Tracking by Query
📃 Paper 📑 PDF

TCSVT 2026
sym

Talk with Your Fingers: A Depth-Aware Benchmark for Air-Writing Recognition
M. Wu, Y. Zhao, X. Li, Shiyu Hu, Y. Cai, J. Wu, W. Wang, K. Huang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
📌 Depth-aware Air-writing 📌 Benchmark Construction 📌 Human-machine Interaction
📃 Paper

IJCAI 2026
sym

COAL: Counterfactual and Observation-Enhanced Alignment Learning for Discriminative Referring Multi-Object Tracking
S. Jia, Shiyu Hu, Y. Wang, X. Cheng, Y. Cao, X. Lu
International Joint Conference on Artificial Intelligence (CCF-B Conference)
📌 Multi-object Tracking 📌 Referring Multi-object Tracking

ECCV 2026
Complete DASTrack figure

DASTrack: Rethinking Temporal Modeling in Visual Object Tracking via Decoupled Auxiliary Supervision
D. Zhang, Shiyu Hu, H. Fu, X. Feng, Y. Wang, KH Cheong, K. Huang
European Conference on Computer Vision (CCF-B Conference)
📌 Visual Object Tracking 📌 Temporal Modeling 📌 Auxiliary Supervision

EMBC 2026
sym

Global-Local Semi-Supervised Modeling for Retinal Layer Boundary Estimation in OCT
K. Li, B. Parikh, H. Yue, Shiyu Hu, S. W. Tan, W. Y. Low, X. Su, KH Cheong
Annual International Conference of the IEEE Engineering in Medicine and Biology Society (CAAI-B Conference)
📌 Optical Coherence Tomography 📌 Semi-supervised Learning

TNSE 2026
sym

Constraint-Driven Evolution of Multimodal Video Intelligence: A Network and System Perspective
X. Li*, X. Li*, Shiyu Hu, Z. Zhang, KH Cheong
IEEE Transactions on Network Science and Engineering
📌 Constraint-driven Video Intelligence 📌Multimodal Understanding and Reasoning
📃 Paper

Mathematics 2026
sym

CalcTutor: Multi-Agent LLM Grading of Handwritten Mathematics with RAG-Grounded Feedback for Adaptive Learning Support
L. Tan, B. Zhu, Shiyu Hu, A. Mishra, Darren J. Yeo, KH Cheong
Mathematics
📌 Adaptive Learning 📌 Multi-agent LLM 📌 Retrieval Augmented Generation
📃 Paper

ICML 2025
sym

CSTrack: Enhancing RGB-X Tracking via Compact Spatiotemporal Features
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
International Conference on Machine Learning (CCF-A Conference, Poster)
📌 Visual Object Tracking 📌 Multi-modal Learning
📃 Paper 📑 PDF

ICASSP 2025
sym

Enhancing Vision-Language Tracking by Effectively Converting Textual Cues into Visual Cues
X. Feng, D. Zhang, Shiyu Hu, X. Li, M. Wu, J. Zhang, X. Chen, K. Huang
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
📌 Visual Language Tracking 📌 Multi-modal Learning 📌 Grounding Model
📃 Paper 📃 PDF

C&E:AI 2025
sym

Artificial Intelligence-Enabled Adaptive Learning Platforms: A Review
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Computers & Education: Artificial Intelligence
📌 Adaptive Learning Platforms 📌 AI for Education 📌 Educational Technology
📃 Paper 📑 PDF

Mathematics 2025
sym

A Comprehensive Review on Automated Grading Systems in STEM Using AI Techniques
L. Tan, Shiyu Hu, Darren J. Yeo, KH Cheong
Mathematics
📌 Automated Grading Systems 📌 AI for Education 📌 Educational Technology
📃 Paper

Innovation and Emerging Technologies 2025
sym

Trustworthy AI in education: Framework, cases, and governance strategies
Y. Ma, X. Li, Shiyu Hu, S. Liu, KH Cheong
Innovation and Emerging Technologies
📌 Trustworthy Artificial Intelligence 📌 Educational Governance 📌 Algorithmic Fairness;
📃 Paper

中国心理卫生杂志 2025
sym

A Review of Intelligent Psychological Assessment Based on Interactive Environment (基于交互环境的智能化心理测评)
K. Huang, Y. Kang, C. Yan, Shiyu Hu, L. Wang, T. Tao, W. Gao
Chinese Mental Health Journal (《中国心理卫生杂志》, CSSCI Journal, Top Psychological Journal in China)
📌 Psychological Assessment System 📌 Gamified Assessment 📌 AI4Science

NeurIPS 2024
sym

Beyond Accuracy: Tracking more like Human via Visual Search
D. Zhang, Shiyu Hu, X. Feng, X. Li, M. Wu, J. Zhang, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
📌 Visual Object Tracking 📌 Visual Search Mechanism 📌 Visual Turing Test
📃 Paper 📑 PDF

NeurIPS 2024
sym

MemVLT: Vision-Language Tracking with Adaptive Memory-based Prompts
X. Feng, X. Li, Shiyu Hu, D. Zhang, M. Wu, J. Zhang, X. Chen, K. Huang
Conference on Neural Information Processing Systems (CCF-A Conference, Poster)
📌 Visual Language Tracking 📌 Human-like Memory Modeling 📌 Adaptive Prompts
📃 Paper 📑 PDF

ICASSP 2024
sym

Robust Single-particle Cryo-EM Image Denoising and Restoration
J. Zhang, T. Zhao, Shiyu Hu, X. Zhao
IEEE International Conference on Acoustics, Speech, and Signal Processing (CCF-B Conference, Poster)
📌 Medical Image Processing 📌 AI4Science 📌 Diffusion Model
📃 Paper 📑 PDF

TCSVT 2024
sym

Finger in Camera Speaks Everything: Unconstrained Air-Writing for Real-World
M. Wu, K. Huang, Y. Cai, Shiyu Hu, Y. Zhao, W. Wang
IEEE Transactions on Circuits and Systems for Video Technology (CCF-B Journal)
📌 Air-writing Technique 📌 Benchmark Construction 📌 Human-machine Interaction
📃 Paper 📃 PDF 🔧 Toolkit

PRCV 2024
sym

VS-LLM: Visual-Semantic Depression Assessment based on LLM for Drawing Projection Test
M. Wu, Y. Kang, X. Li, Shiyu Hu, X. Chen, Y. kang, W. Wang, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference)
📌 Psychological Assessment System 📌 Gamified Assessment 📌 AI4Science
📃 Paper 📃 PDF

PRCV 2023
sym

A Hierarchical Theme Recognition Model for Sandplay Therapy
X. Feng, Shiyu Hu, X. Chen, K. Huang
Chinese Conference on Pattern Recognition and Computer Vision (CCF-C Conference, Poster)
📌 Psychological Assessment System 📌 Gamified Assessment 📌 AI4Science
📃 Paper 📑 PDF 🔖 Supplementary 🪧 Poster

CSAI 2023
sym

Rethinking Similar Object Interference in Single Object Tracking
Y. Wang, Shiyu Hu, X. Zhao
International Conference on Computer Science and Artificial Intelligence (EI Conference, Oral)
📌 Visual Object Tracking 📌 Similar Object Interference 📌 Data Mining
📃 Paper 🗒 bibTex 📑 PDF

Neurocomputing 2022
sym

Revisiting Instance Search: A New Benchmark Using Cycle Self-training
Y. Zhang, C. Liu, W. Chen, X. Xu, F. Wang, H. Li, Shiyu Hu, X. Zhao
Neurocomputing (CCF-C Journal)
📌 Video Instance Search 📌 Benchmark Construction 📌 Data Mining
📃 Paper 📑 PDF 🌐 Project

图学学报 2021
sym

Visual Turing: The Next Development of Computer Vision in The View of Human-computer Gaming (视觉图灵:从人机对抗看计算机视觉下一步发展)
K. Huang, X. Zhao, Q. Li, Shiyu Hu
Journal of Graphics (《图学学报》, CCF-C Chinese Journal)
📌 Visual Object Tracking 📌 Intelligent Evaluation Technique 📌 AI4Science
📃 Paper 📑 PDF

Workshop

Diverse Text Generation for Visual Language Tracking Based on LLM, X. Li, X. Feng, Shiyu Hu, M. Wu, D. Zhang, J. Zhang, K. Huang, the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (Workshop in CCF-A Conference, Oral, Best Paper Honorable Mention), 📃 Paper 📃 PDF 🪧 Poster 📹 Slides 🌐 Platform 🔧 Toolkit 💾 Dataset 🏆 Award

Preprint

Preprint
Complete figure for the survey of streaming video understanding

How to Respond, How to Memorize, How to Be Fast: A Survey of Streaming Video Understanding
X. Li*, Shiyu Hu*, X. Feng, J. Zhao, K. Huang (*Equal Contributions)
📌 Streaming Video Understanding 📌 Memory Modeling 📌 Efficient Inference
📃 Paper

Preprint
sym

FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
Shiyu Hu*, X. Li*, X. Li, J. Zhang, Y. Wang, X. Zhao, KH Cheong (*Equal Contributions)
📌 Large Vision-Language Models 📌 Video Caption 📌 Video Understanding
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

When LLMs Learn to be Students: The SOEI Framework for Modeling and Evaluating Virtual Student Agents in Educational Interaction
Y. Ma*, Shiyu Hu*, X. Li, Y. Wang, Y. Chen, S. Liu, KH Cheong (*Equal Contributions)
📌 AI4Education 📌 LLMs 📌 LLM-based Agent
📃 Paper 📑 PDF

Preprint
sym

EduVerse: A User-Defined Multi-Agent Simulation Space for Education Scenario
Y. Ma*, Shiyu Hu*, B. Zhu, Y. Wang, Y. Kang, S. Liu, KH Cheong (*Equal Contributions)
📌 AI4Education 📌 LLMs 📌 LLM-based Agent
📃 Paper 📑 PDF

Preprint
sym

EduPersona: Benchmarking Subjective Ability Boundaries of Virtual Student Agents
B. Zhu*, Shiyu Hu*, Y. Ma, Y. Zhang, KH Cheong (*Equal Contributions)
📌 AI4Education 📌 LLMs 📌 LLM-based Agent
📃 Paper 📑 PDF

Preprint
sym

SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
Y. Wang*, Shiyu Hu*, S. Jia, P. Xu, H. Ma, Y. Ma, J. Zhang, X. Lu, X. Zhao (*Equal Contributions)
📌 Visual Object Tracking 📌 Similar Object Interference 📌 Multimodal Learning
📃 Paper 📑 PDF

Preprint
sym

How Texts Help? A Fine-grained Evaluation to Reveal the Role of Language in Vision-Language Tracking
X. Li*, Shiyu Hu*, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang (*Equal Contributions)
📌 Visual Language Tracking 📌 Multimodal Learning 📌 Evaluation Technique
📃 Paper 📑 PDF

Preprint
Complete STEMVerse diagnostic framework for STEM reasoning

STEMVerse: A Dual-Axis Diagnostic Framework for STEM Reasoning in Large Language Models
X. Li, X. Li, J. Zhao, Shiyu Hu✉️
📌 STEM Reasoning 📌 Large Language Models 📌 Diagnostic Evaluation
📃 Paper 📑 PDF

Preprint
sym

DTVLT: A Multi-modal Diverse Text Benchmark for Visual Language Tracking Based on LLM
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
📌 Visual Language Tracking 📌 Large Language Model 📌 Evaluation Technique
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

Visual Language Tracking with Multi-modal Interaction: A Robust Benchmark
X. Li, Shiyu Hu, X. Feng, D. Zhang, M. Wu, J. Zhang, K. Huang
📌 Visual Language Tracking 📌 Multi-modal Interaction 📌 Evaluation Technology
📃 Paper 📑 PDF 🌐 Project

Preprint
sym

Beyond Accuracy: Evaluating Grounded Visual Evidence in Thinking with Images
X. Li*, X. Li*, R. Pi, Shiyu Hu, J. Zhao,J. Gao
📌 Thinking-with-Image 📌 Vision-Language Models 📌 Agentic Models
📃 Paper 📑 PDF

Preprint
sym

Nearing or Surpassing: Overall Evaluation of Human-Machine Dynamic Vision Ability
Shiyu Hu, X. Zhao, Y. Wang, Y. Shan, K. Huang
📌 Visual Object Tracking 📌 Intelligent Evaluation Technique 📌 AI4Science
📑 PDF

⚙️ Projects

This section documents research software, evaluation platforms, academic challenges, and funded projects. Associated research outputs are listed under Publications.

Research Software

Darknet-Cross: Lightweight Deep Learning Framework for Heterogeneous Computing

2018.03 - 2018.11 · GitHub

A cross-platform acceleration framework developed for Android and Ubuntu across mobile and desktop GPUs. This work formed the engineering component of my master’s thesis at HKU.

Research Platforms

VideoCube / MGIT Platform

2019.11 - Present · Platform

Evaluation infrastructure developed and maintained for the Global Instance Tracking and Multi-modal Global Instance Tracking studies published at TPAMI 2023 and NeurIPS 2023. The platform supports dataset access, tracker registration, standardized result submission, and reproducible evaluation. As of July 2026, it had recorded more than 1.66 million visits, 1,900 data users, 720 registered trackers, and 960 result submissions.

SOTVerse / VLTVerse Platform

2021.07 - Present · Platform

A task-space evaluation platform developed and maintained for the SOTVerse and VLTVerse research program, including the study published in IJCV 2024. The platform provides structured evaluation across tracking scenarios, target categories, and linguistic specifications, and had recorded more than 267,000 visits as of July 2026.

BioDrone Platform

2022.05 - Present · Platform

Benchmark and evaluation infrastructure developed and maintained for real-world drone tracking research, including the BioDrone study published in IJCV 2024. The platform supports dataset access, standardized benchmarking, and comparative evaluation, and had recorded more than 541,000 visits as of July 2026.

GOT-10k Platform

2020.07 - Present · Platform

Long-term maintenance of the large-scale evaluation platform supporting the GOT-10k research published in TPAMI 2021. The platform provides standardized benchmark access, tracker registration, result submission, and evaluation services. As of July 2026, it had recorded more than 7.56 million visits, 11,000 data users, 32,700 registered trackers, and 406,000 result submissions.

Challenges

Hislopvision Challenge

2023.05 - 2023.11 · Platform

Organization of the Hislopvision track for the 3rd High-speed and Low-power Visual Understanding Challenge at PRCV 2023, with participating teams from Tsinghua University, Beijing Institute of Technology, and Jilin University.

Cell Tracking Challenge

2021.01 - 2021.04 · Challenge

Our method ranked second on Fluo-C2FL-MSC+ and third on Fluo-C2FL-Huh7, based on the challenge rankings recorded in October 2023.

Funded Research Project

Human-Computer Interaction in Intelligent Education

2024.01 - 2025.01

Proposal development and project delivery for a study funded by the 2023 Intelligent Education PhD Research Fund at the Shanghai Institute of AI Education, East China Normal University.

🏆 Honors and Awards

  • 2026 Reviewer Award, the 43rd International Conference on Machine Learning (ICML 2026)
  • 2025 IEEE SMCS TEAM Program Award by the IEEE Systems, Man, and Cybernetics Society
  • 2024 Best Paper Honorable Mention in the 3rd Workshop on Vision Datasets Understanding and DataCV Challenge in CVPR 2024 (CVPRW最佳论文提名)
  • 2024 Beijing Outstanding Graduates (北京市优秀毕业生, top 5%)
  • 2023 China National Scholarship (国家奖学金, top 1%, only 8 Ph.D. students in main campus of University of Chinese Academy of Sciences win this scholarship)
  • 2023 First Prize of Climbing Scholarship in Institute of Automation, Chinese Academy of Sciences (攀登一等奖学金, only 6 students in Institute of Automation, Chinese Academy of Sciences win this scholarship)
  • 2022 Merit Student of University of Chinese Academy of Sciences (中国科学院大学三好学生)
  • 2017 Excellent Innovative Student of Beijing Institute of Technology (北京理工大学优秀创新学生)
  • 2016 College Scholarship of Chinese Academy of Sciences (中国科学院大学生奖学金)
  • 2016 Excellent League Member on Youth Day Competition of Beijing Institute of Technology (北京理工大学优秀团员)
  • 2015 National First Prize in Contemporary Undergraduate Mathematical Contest in Modeling (CUMCM) (全国大学生数学建模竞赛国家一等奖, top 1%, only 1 team in Beijing Institute of Technology win this prize) [📑PDF] [📖Selected and Reviewed Outstanding Papers in CUMCM (2011-2015) (Chapter 9)]
  • 2015 First Prize of Mathematics Modeling Competition within Beijing Institute of Technology (北京理工大学数学建模校内选拔赛第一名)
  • 2015 Outstanding Individual on Summer Social Practice of Beijing Institute of Technology (北京理工大学暑期社会实践优秀个人)
  • 2015 Second Prize on Summer Social Practice of Beijing Institute of Technology (北京理工大学暑期社会实践二等奖, team leader)
  • 2015 Outstanding Student Cadre of Beijing Institute of Technology (北京理工大学优秀学生干部)
  • 2015 Outstanding League Cadre on Youth Day Competition of Beijing Institute of Technology (北京理工大学优秀团干部)
  • 2015 Outstanding Youth League Branch of Beijing Institute of Technology (北京理工大学优秀团支部, team leader)
  • 2015 Top-10 Activities on Youth Day Competition of Beijing Institute of Technology (北京理工大学十佳团日活动, team leader)
  • 2014 Outstanding Student of Beijing Institute of Technology (北京理工大学优秀学生)
  • 2014, 2015, 2016, 2017 Academic Scholarship of Beijing Institute of Technology (北京理工大学学业奖学金)

📣 Activities and Services

Tutorial

34th International Joint Conference on Artificial Intelligence (IJCAI)

  • Title: Human-Centric and Multimodal Evaluation for Explainable AI: Moving Beyond Benchmarks
  • Date & Location: 14:00-15:30, 18th August, 2025, Montreal, Canada

28th European Conference on Artificial Intelligence (ECAI)

  • Title: From Benchmarking to Trustworthy AI: Rethinking Evaluation Methods Across Vision and Complex Systems
  • Date & Location: 26th October, 2025, Bologna, Italy
    🌐 Webpage

2025 IEEE International Conference on Systems, Man, and Cybernetics (SMC)

  • Title: The Synergy of Large Language Models and Evolutionary Optimization on Complex Networks
  • Date & Location: 5th October, 2025, Vienna, Austria

31st IEEE International Conference on Image Processing (ICIP)

  • Title: An Evaluation Perspective in Visual Object Tracking: from Task Design to Benchmark Construction and Algorithm Analysis
  • Date & Location: 9:00-12:30, 27th October, 2024, Abu Dhabi, United Arab Emirates
  • Duration: Half-day
    📹 Slides 🌐 Webpage

27th International Conference on Pattern Recognition (ICPR)

  • Title: Visual Turing Test in Visual Object Tracking: A New Vision Intelligence Evaluation Technique based on Human-Machine Comparison
  • Date & Location: 14:30-18:00, 1st December, 2024, Kolkata, India
  • Duration: Half-day
    📹 Slides

17th Asian Conference on Computer Vision (ACCV)

  • Title: From Machine-Machine Comparison to Human-Machine Comparison: Adapting Visual Turing Test in Visual Object Tracking
  • Date & Location: 9:00-12:00, 9th December, 2024, Hanoi, Vietnam
  • Duration: Half-day
    📹 Slides 🌐 Webpage

Mini-Symposium

The Fifth International Nonlinear Dynamics Conference (NODYCON 2026)

  • Title: Complex Network Systems and Large Language Models
  • Date & Location: 20th-23rd September, 2026, Sapienza University of Rome, Italy
    🌐 Webpage

Talk

Chinese Congress on Image and Graphics (CCIG 2026)

  • Title: Visual Understanding Reliability in Open Environments: from Robust Perception to Semantic Consistency
  • Date & Location: 29th May, 2026, Guangzhou, China

TPC Member

Guest Editor

Associate Editor

Reviewer

  • Conferences: NeurIPS, ICML, ICLR, CVPR, ECCV, ICCV, ACL, AAAI, IJCAI, ACM MM, ICRA, and AISTATS.
  • Journals: ACM Computing Surveys, IEEE Transactions on Image Processing, SCIENCE CHINA Information Sciences, Pattern Recognition, Transactions on Machine Learning Research, IEEE Transactions on Network Science and Engineering, IEEE Transactions on Vehicular Technology, Information Fusion, Visual Intelligence, Engineering Applications of Artificial Intelligence, Expert Systems with Applications, Neurocomputing, and Knowledge-Based Systems.

Member

  • Societies: Institute of Electrical and Electronics Engineers (IEEE), China Society of Image and Graphics (CSIG), Chinese Association for Artificial Intelligence (CAAI), and China Computer Federation (CCF).

📄 CV

✉️ Contact

  • shiyu.hu@ntu.edu.sg (Main)
  • hushiyu199510@gmail.com (Personal)
  • hushiyu2019@ia.ac.cn (Valid from 2019.06 - 2024.07)
Visitor insights homepage visits Collecting geographic distribution
Top countries and regions
Country-level statistics will appear as visits accumulate.
Anonymous country-level aggregates only; IP addresses are not stored. Recorded from July 2026.


© Shiyu Hu | Last updated: 2026-07