Teaching foundation models to see, remember, and reason over long videos

PhD Candidate, NTU Singapore A*STAR CIS Scholar, I²R Expected Dec 2027

I'm a PhD candidate at Nanyang Technological University (School of Electrical & Electronic Engineering) and an A*STAR CIS Scholar at the Institute for Infocomm Research (I²R), advised by Prof. Xudong Jiang. I study multimodal video understanding, curious what a model can really afford to remember.

That curiosity shapes how I build: state-space models and spatiotemporal compression that let vision-language foundation models perceive and reason over long videos at tractable cost. This work has produced first-author papers at CVPR 2025 (TV3S), ECCV 2026 (STAC), and in Machine Intelligence Research. In late 2025 I spent a research visit with Prof. Yun Liu's group at Nankai University, Tianjin.

Before the PhD, I earned First-Class Honours at NTU while working full-time as a research assistant, leading R&D that secured over S$200K in government funding and mentoring 40+ students to national competition wins.

Multimodal Video Understanding Video Reasoning Segmentation Multimodal LLMs & VLMs State Space Models (Mamba) Spatiotemporal Compression Model Efficiency

News

2026
🎉 ECCV 2026: STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation accepted to ECCV 2026, my second first-author paper at a top-tier vision conference.
2026
📗 Journal: Evaluating SAM2 for Video Semantic Segmentation accepted at Machine Intelligence Research (MIR).
Dec 2025
🔬 Research visit: Began a three-month visit (Dec 2025 – Feb 2026) to Prof. Yun Liu's group at Nankai University, Tianjin, working on multimodal understanding in video models.
Jun 2025
🎤 CVPR 2025: Presented TV3S at CVPR in Nashville.
2025
💰 Grant: Awarded the Lambda Research Grant for GPU computing resources.
2025
🤝 Leadership: Appointed Assistant President of the NTU Graduate Students' Association, after serving as Assistant Vice-President (Academic) in 2024.
Feb 2025
🎉 CVPR 2025: TV3S: Exploiting Temporal State Space Sharing for Video Semantic Segmentation accepted (22.1% acceptance rate), my first first-author CVPR paper.
Jan 2024
🎓 PhD started: Began my PhD at NTU with the A*STAR Computing and Information Science (ACIS) Scholarship.

Publications

Featured

STAC concept: dense video tokens aggregated and compressed into a compact set

STAC: Selective Spatiotemporal Aggregation and Compression for Video Reasoning Segmentation

Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Jing Yang, Henghui Ding, Xue Geng, Xudong Jiang

European Conference on Computer Vision, 2026 ECCV 2026

STAC compresses the spatiotemporal token stream feeding a multimodal LLM: roughly 85% fewer tokens and a 1.8× speedup in training and inference, while keeping reasoning-segmentation accuracy competitive.

TV3S qualitative results: video frames with segmentation maps compared against CFFM and ground truth

Exploiting Temporal State Space Sharing for Video Semantic Segmentation

Syed Ariff Syed Hesham, Yun Liu, Guolei Sun, Henghui Ding, Jing Yang, Ender Konukoglu, Xue Geng, Xudong Jiang

IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025 CVPR 2025 22.1% Acc. Rate

TV3S shares Mamba state spaces across time so temporal context propagates through a selective gate instead of a memory-heavy feature pool, enabling parallel training and inference over long videos.

Pipeline aggregating SAM object masks with semantic segmentation predictions

Evaluating SAM2 for Video Semantic Segmentation

Syed Hesham Syed Ariff, Yun Liu, Guolei Sun, Jing Yang, Henghui Ding, Xue Geng, Xudong Jiang

Machine Intelligence Research, accepted 2026 Journal

A systematic study of what SAM2's video foundation model does, and does not, bring to dense, class-aware video segmentation.

Taxonomy diagram organizing video scene parsing methods, datasets, and benchmarks

A Comprehensive Survey on Video Scene Parsing: Advances, Challenges, and Prospects

Guohuan Xie, Syed Ariff Syed Hesham, Bing Li, Wenya Guo, Guolei Sun, Ming-Ming Cheng, Yun Liu

arXiv:2506.13552, 2025 Under Review

A unified review of video scene parsing across semantic, instance, panoptic, tracking, and open-vocabulary segmentation, tracing the field from convolutional pipelines to foundation-model approaches.

Earlier Publications (2021 – 2024)

2024

Tracking and Monitoring of Underwater Object with SLAM

Lijun Jiang, Syed Ariff Syed Hesham*, Lan JunHang, Seah Kai Wen Kelvin, Yuhan Jiang, Yao Mengdi, Wei Dongliang, Bo Jiang

IEEE 19th Conference on Industrial Electronics and Applications (ICIEA)

2023

Adapted Lightweight MobileNet for Tire Pattern Classification

Syed Ariff Syed Hesham, Lijun Jiang, Keng Pang Lim, Sui JinZhou, Phua Zen Long, Lan JunHang, Zhao Hu, Yang Shanglin, Zhao Xu, Zhao Chuanfeng

IEEE 18th Conference on Industrial Electronics and Applications (ICIEA)

2023

Leveraging on Few-Shot Learning for Tire Pattern Classification in Forensics

Lijun Jiang, Syed Ariff Syed Hesham*, Keng Pang Lim, Changyun Wen

Journal of Automation and Intelligence

2022

A Versatile Application for Visual SLAM with Object Detection

Lijun Jiang, Syed Ariff Syed Hesham*, Keng Pang Lim, Yusong Wang, Hongkai Lin, Yuhang Zhao

IEEE 17th Conference on Industrial Electronics and Applications (ICIEA)

2022

YOLO Based Thermal Screening Using AI for Instinctive Human Facial Detection

Lijun Jiang, Syed Ariff Syed Hesham*, Keng Pang Lim, Krishnadas Manoj, Mohammed Razi, Zhang Zishuo, Ma Chenxin, Bo Jiang, Hamid Saeedipour

IEEE 17th Conference on Industrial Electronics and Applications (ICIEA)

2022

Borescope Tracking and Visualization of Internal Aero-Structure with VSLAM

Lijun Jiang, Syed Ariff Syed Hesham*, Keng Pang Lim, Sui JinZhou, Bo Jiang

IEEE 17th Conference on Industrial Electronics and Applications (ICIEA)

2021

Improving Recognition Performance for Low-Resolution Images Using DBPN

Lijun Jiang, Keng Pang Lim, Syed Ariff Syed Hesham*

IEEE 16th Conference on Industrial Electronics and Applications (ICIEA)

2021

Crack Detection on Aircraft Composite Structures Using Faster R-CNN

Lijun Jiang, Syed Ariff Syed Hesham*, Haoming Shi, Hamid Saeedipour

IEEE 16th Conference on Industrial Electronics and Applications (ICIEA)

2021

PIX2PT Map for Transfer-Based Few-Shot Learning

Syed Ariff Syed Hesham, Sui JinZhou*, Gabrielle Ee Song Xin, Lek Chen Ping, Lijun Jiang

IEEE International Conference on Multimedia & Expo Workshops (ICMEW)

* denotes equal contribution

View all publications on Google Scholar

Timeline

Dec 2025 – Feb 2026

Visiting Researcher

Nankai University, Tianjin, China

Prof. Yun Liu's group, enhancing multimodal understanding in video models.

Jan 2024 – Present

PhD, Electrical & Electronic Engineering

Nanyang Technological University & A*STAR I²R, Singapore

Topic: Effective Visual Perception with Large Foundation Models
Advisor: Prof. Xudong Jiang · Expected: Dec 2027

  • Efficient video semantic segmentation, video reasoning segmentation, and multimodal video understanding with large foundation models (CVPR 2025, ECCV 2026, MIR)
  • Designed TV3S (temporal state space sharing) and STAC (selective spatiotemporal aggregation and compression) to make long-video segmentation and reasoning tractable
  • Benchmarked SAM2 for video semantic segmentation and co-authored a survey on video scene parsing
GPA: 5.0/5.0 A*STAR CIS Scholar
Nov 2020 – Jan 2024

Research Assistant (Computer Vision)

Republic Polytechnic, School of Engineering, Singapore

Full-time research role held concurrently with my part-time bachelor's degree.

  • Developed a cost-effective near-field pose-estimation system that secured over S$200K in government funding
  • Built a real-time SLAM visualization pipeline in Python, C++, C#, and ROS
  • Mentored 40+ students across project teams, leading them to national competition wins and peer-reviewed publications
Aug 2020 – Dec 2023

BEng (Honours), Electrical & Electronic Engineering

Nanyang Technological University, Singapore

FYP: Real-time SLAM using deep features

GPA: >4.9/5.0 First-Class Honours Dean's List ×4
Mar 2018 – Aug 2018

Research Intern

Continental Automotive R&D, Singapore

  • Delivered practical solutions across three internal R&D projects, including bus passenger monitoring, an autonomous pick-and-place robot, and indoor localization with Bluetooth beacons
  • Presented results to cross-functional teams and at company open-house showcases
2016 – 2019

Diploma, Electrical & Electronic Engineering

Republic Polytechnic, Singapore

FYP: Improving Recognition Performance for Low-Resolution Images Using DBPN

Director's Roll of Honour GPA: >3.9/4.0

Skills

Research

Multimodal Video Understanding Video Reasoning & Segmentation Multimodal LLMs & VLMs Vision-Language Alignment State Space Models (Mamba) Spatiotemporal Compression Model Efficiency SLAM & 3D Perception

Technical

Python PyTorch C/C++ TensorFlow Distributed Training (Multi-node, Multi-GPU) Git Docker Linux ROS SLURM

Transferable

Technical Writing & Publication Mentorship & Team Leadership Cross-functional Collaboration

Languages

English (Fluent) Tamil (Fluent) Hindi (Fluent)

Get in Touch

Labs

A*STAR I²R, Fusionopolis

Rapid-Rich Object Search (ROSE) Lab, NTU Singapore

Profiles