Teaching foundation models to see, remember, and reason over long videos
PhD Candidate, NTU Singapore
A*STAR CIS Scholar, I²R
Expected Dec 2027
I'm a PhD candidate at Nanyang Technological University (School of Electrical & Electronic Engineering) and an A*STAR CIS Scholar at the Institute for Infocomm Research (I²R), advised by Prof. Xudong Jiang. I study multimodal video understanding, curious what a model can really afford to remember.
That curiosity shapes how I build: state-space models and spatiotemporal compression that let vision-language foundation models perceive and reason over long videos at tractable cost. This work has produced first-author papers at CVPR 2025 (TV3S), ECCV 2026 (STAC), and in Machine Intelligence Research. In late 2025 I spent a research visit with Prof. Yun Liu's group at Nankai University, Tianjin.
Before the PhD, I earned First-Class Honours at NTU while working full-time as a research assistant, leading R&D that secured over S$200K in government funding and mentoring 40+ students to national competition wins.
Multimodal Video Understanding
Video Reasoning Segmentation
Multimodal LLMs & VLMs
State Space Models (Mamba)
Spatiotemporal Compression
Model Efficiency