Experience
Amazon — AGI Foundations (Audio / Multimodal LLM)
Applied Scientist Intern
- Full-stack distributed RL post-training: sole owner of the end-to-end pipeline for multimodal reasoning LLMs — training infrastructure, algorithm and reward design, evaluation and model delivery — and sole first author of the resulting publication (CESAR, ICLR 2026).
- Built and scaled a multi-node, multi-GPU GRPO training stack for omni-modal LLMs, and designed a process-reward framework that moves RLVR beyond outcome verification and removes any dependence on human-annotated reasoning traces.
Education
University of Illinois Urbana-Champaign
Ph.D. in Computer Science GPA 4.0/4.0
- Research: Autonomous RL post-training for flow/diffusion models & multimodal reasoning LLMs
- Service: Reviewer @ ICML 2025–26, ICLR 2025–26, NeurIPS 2024–26, COLM 2026, CVPR 2026, WACV 2027, AAAI 2027, AISTATS 2025
Tsinghua University
M.Eng. in Computer Technology GPA 3.97/4.0 · Top 1.3%
- Service: Reviewer @ NeurIPS 2022–23, ICML 2023–24, ICLR 2024
Nankai University
B.Eng. in Intelligent Science and Technology Rank 1/83
- GPA 93.28/100 (3.9/4.0) · National Scholarship ×2 (Top 1%)
Selected Publications * = first / co-first author
ICML 2026
Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation
+10.3% over strong RAG and reranking baselines
ICML 2026
15.8–62.6% fewer tokens at equal or better accuracy
COLM 2026
Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models
3B model competitive with GPT-5-mini and Claude-Haiku-4.5 on BFCL-V3
TMLR 2026
Online RL for self-improving protein design, without human-curated data
ICLR 2026
SOTA MMAU · Outperforms Gemini 2.5 Pro & GPT-4o Audio
ICLR 2026
1.5× lossless speedup (LIBERO) · 2.4× (SimplerEnv)
NeurIPS 2025
2B SD3 surpasses 4.8B & 12B models
NeurIPS 2025
SOTA 79.36% Top-1 on ImageNet-1K (ResNet-50)
ICLR 2025
First online RLHF for flow matching models
2026
2.55× / 3.77× speedup · 13.8 → 26.3 Hz on real-world GR00T tasks
TPAMI 2026
ICLR 2023 Oral
Ranked 5/4176 · 24 Atari world records · 78× data efficiency
ICML 2022
Beats Agent57 with 500× less data · 22 world records
Awards & Honors
- 2024GPA 4.0/4.0, UIUC Ph.D. Program
- 2024GPA 3.97/4.0, Top 1.3% — Tsinghua University M.Eng.
- 2021Outstanding Graduates (Top 1%) — Nankai UniversityTop 1%
- 2021Excellent Graduation Thesis — Nankai University
- 2021Tang Lixin ScholarshipTop 1%
- 2020National Scholarship (Rank 1/83)Top 1%
- 2020Nomination for Zhou Enlai Scholarship
- 2019National Scholarship (Rank 1/83)Top 1%
- 20193rd Prize, RoboCup@HOME Education World Final — Sydney, Australia
- 2019Bronze Medal, ACM/ICPC Asia Regional — Xuzhou
- 2018National 2nd Prize, Mathematical Contest in ModelingTop 5%