Yao Xiao

CS Ph.D. candidate @ UIUC

Google Scholar
GitHub
X
Linkedin

About Me

I am a final-year Ph.D. candidate in Computer Science at University of Illinois Urbana-Champaign, advised by Prof. Derek Hoiem.

I aim to build intelligent agents that plan under partial information, remember what matters, and draw on past experience to avoid repeating failures. More specifically:

  • Closed-loop Planning: How can models make decisions under partial information? Each decision must build on past observations while shaping what comes next.

  • Memory Representation: How can an agent compress its past into a compact, updatable state that preserves what matters and turns experience into reusable lessons?

  • Experience Retrieval: How do we recall useful past experiences while filtering out distractions?

I'm seeking full-time opportunities starting Summer 2027. If you're interested in my research or would simply like to connect, feel free to email me!

Selected Work

For the full publication list, please refer to my Google Scholar.

NeurIPS 2026

ReToken: One Token to Improve Vision–Language Models for Visual Retrieval

Yao Xiao, Reuben Tan, Zhen Zhu, Yuqun Wu, Jianfeng Gao, Derek Hoiem

NeurIPS, 2026.

[TL;DR]
Attention tells you where a VLM looks, so shouldn't it tell you which frames matter? It doesn't: the retrieval signal lives in the value space, and a single learned token is enough to read it out.
[Paper] [Code]
ECCV 2026🏅Spotlight

How to Teach Large Multimodal Models New Skills?

Zhen Zhu, Yiming Gong, Yao Xiao, Yaoyao Liu, Derek Hoiem

ECCV, 2026. Spotlight.

[Paper] [Code]
TMLR 2025🏅J2C

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu, Zhen Zhu, Derek Hoiem

TMLR, 2025. J2C certification (top 10%).

[TL;DR]
A simple, training-free approach to get region tokens directly comparable to text embeddings, achieving state-of-the-art zero-shot region-level understanding.
[Paper] [Code]

Blog

A casual blog for sharing thoughts and tracking my journey. No tech talk, just real life.

→ Full list

Teaching

Services

Co-organizer: UIUC Computer Vision External Speaker Series, 2025 - 2026.

Reviewer: CVPR, ICCV, ECCV, ICLR, NeurIPS, WACV, IEEE Access.

Last Updated: 9/24/2026, 6:33:35 PM