Jiajun Fan
Publications

Peer-reviewed work

Reinforcement learning for generative models and reasoning LLMs, agentic RL, and efficient inference. Listed newest first; my name is bold in each author list.

Google Scholar 482 citations 12 h-index 12 i10-index 23 papers

Research overview

I study how to train multimodal large generative/agentic models with reinforcement learning in a way that is stable (e.g., lifelong continual learning) — avoiding the diversity and performance collapse that RL fine-tuning tends to cause — and progressively autonomous (e.g., self-evolving), steadily removing the human data collection and labelling the pipeline depends on. That work spans these domains:

Data-distribution optimizationgame agents
Text-to-image generationflow matching & diffusion
Audio reasoning LLMsmultimodal reasoning
Tool-use agentsagentic RL
Procedure-Aware RL COLM 2026 PyRAG 2026
AI for Scienceprotein & materials design
Math reasoningtoken efficiency
VLA & roboticsefficient inference
Robot designmorphology & control

2026
COLM 2026

Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models

Qinglong Zheng, Jiajun Fan, Chaoran Cheng, Ge Liu. "Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models." COLM 2026.
@inproceedings{zheng2026procedureaware,
  title={Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models},
  author={Zheng, Qinglong and Fan, Jiajun and Cheng, Chaoran and Liu, Ge},
  booktitle={Conference on Language Modeling},
  year={2026}
}
TMLR 2026

ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning

Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu. "ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning." Transactions on Machine Learning Research (TMLR), 2026.
Paper
@article{wang2025proteinzero,
  title={ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning},
  author={Wang, Ziwen and Fan, Jiajun and Guo, Ruihan and Nguyen, Thao and Ji, Heng and Liu, Ge},
  journal={Transactions on Machine Learning Research},
  issn={2835-8856},
  year={2026},
  url={https://arxiv.org/abs/2506.07459}
}
arXiv

ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models

Ye Li, Huanan Liu, Kangye Ji, Yuan Meng, Jiajun Fan, Yuansong Wang, Shiyu Qin, Chenglei Wu, Shu-Tao Xia, Zhi Wang. "ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models." arXiv preprint arXiv:2605.29438, 2026.
Paper
@article{li2026elegantvla,
  title={ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models},
  author={Li, Ye and Liu, Huanan and Ji, Kangye and Meng, Yuan and Fan, Jiajun and Wang, Yuansong and Qin, Shiyu and Wu, Chenglei and Xia, Shu-Tao and Wang, Zhi},
  journal={arXiv preprint arXiv:2605.29438},
  year={2026},
  url={https://arxiv.org/abs/2605.29438}
}
arXiv

Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation

Jiashuo Sun, Jimeng Shi, Yixuan Xie, Saizhuo Wang, Jash Rajesh Parekh, Pengcheng Jiang, Zhiyi Shi, Jiajun Fan, Qinglong Zheng, Peiran Li, Shaowen Wang, Ge Liu, Jiawei Han. "Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation." arXiv preprint arXiv:2605.12975, 2026.
Paper
@article{sun2026pyrag,
  title={Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation},
  author={Sun, Jiashuo and Shi, Jimeng and Xie, Yixuan and Wang, Saizhuo and Parekh, Jash Rajesh and Jiang, Pengcheng and Shi, Zhiyi and Fan, Jiajun and Zheng, Qinglong and Li, Peiran and Wang, Shaowen and Liu, Ge and Han, Jiawei},
  journal={arXiv preprint arXiv:2605.12975},
  year={2026},
  url={https://arxiv.org/abs/2605.12975}
}
ICML 2026

Batched Contextual Reinforcement

Bangji Yang, Hongbo Ma, Jiajun Fan, Ge Liu. "Batched Contextual Reinforcement." ICML 2026.
Paper
@inproceedings{yang2026bcr,
  title={Batched Contextual Reinforcement},
  author={Yang, Bangji and Ma, Hongbo and Fan, Jiajun and Liu, Ge},
  booktitle={International Conference on Machine Learning},
  year={2026},
  url={https://openreview.net/forum?id=8Oc3Mx754M}
}
ICML 2026

Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation

Jiashuo Sun, Pengcheng Jiang, Saizhuo Wang, Jiajun Fan, Heng Wang, Siru Ouyang, Ming Zhong, Yizhu Jiao, Chengsong Huang, Xueqiang Xu, Pengrui Han, Peiran Li, Jiaxin Huang, Ge Liu, Heng Ji, Jiawei Han. "Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation." ICML 2026.
Paper
@inproceedings{sun2026barrag,
  title={Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation},
  author={Sun, Jiashuo and Jiang, Pengcheng and Wang, Saizhuo and Fan, Jiajun and Wang, Heng and Ouyang, Siru and Zhong, Ming and Jiao, Yizhu and Huang, Chengsong and Xu, Xueqiang and Han, Pengrui and Li, Peiran and Huang, Jiaxin and Liu, Ge and Ji, Heng and Han, Jiawei},
  booktitle={International Conference on Machine Learning},
  year={2026},
  url={https://openreview.net/forum?id=Tt8lCe1NrW}
}
ICLR 2026

SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration

Ye Li, Yuan Meng, Zewen Sun, Kangye Ji, Chen Tang, Jiajun Fan, Xinzhu Ma, Shu-Tao Xia, Zhi Wang, Wenwu Zhu. "SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration." ICLR 2026.
Paper
@inproceedings{li2026spvla,
  title={{SP-VLA}: A Joint Model Scheduling and Token Pruning Approach for {VLA} Model Acceleration},
  author={Li, Ye and Meng, Yuan and Sun, Zewen and Ji, Kangye and Tang, Chen and Fan, Jiajun and Ma, Xinzhu and Xia, Shu-Tao and Wang, Zhi and Zhu, Wenwu},
  booktitle={International Conference on Learning Representations},
  year={2026},
  url={https://openreview.net/forum?id=RwdGIIjPlC}
}
TPAMI 2026

PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference

Ye Li, Chen Tang, Yuan Meng, Jiajun Fan, Zenghao Chai, Xinzhu Ma, Zhi Wang, Wenwu Zhu. "PRANCE: Joint Token-Optimization and Structural Channel-Pruning for Adaptive ViT Inference." TPAMI 2026.
Paper
@article{li2026prance,
  title={{PRANCE}: Joint Token-Optimization and Structural Channel-Pruning for Adaptive {ViT} Inference},
  author={Li, Ye and Tang, Chen and Meng, Yuan and Fan, Jiajun and Chai, Zenghao and Ma, Xinzhu and Wang, Zhi and Zhu, Wenwu},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
  year={2026},
  url={https://arxiv.org/abs/2407.05010}
}
ICLR 2026

Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards

Jiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey, Prashanth G. Shivakumar, Ivan Bulyko, Ankur Gandhe, Ge Liu, Yile Gu. "Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards." ICLR 2026.
Project page Paper
@inproceedings{fan2026cesar,
  title={Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio {LLMs} via Reasoning Process Rewards},
  author={Fan, Jiajun and Ren, Roger and Li, Jingyuan and Pandey, Rahul and Shivakumar, Prashanth G. and Bulyko, Ivan and Gandhe, Ankur and Liu, Ge and Gu, Yile},
  booktitle={International Conference on Learning Representations},
  year={2026},
  url={https://openreview.net/forum?id=DUr48hxO2h}
}
2025
arXiv

Fine-tuning Flow Matching Generative Models with Intermediate Feedback

Jiajun Fan, Chaoran Cheng, Shuaike Shen, Xiangxin Zhou, Ge Liu. "Fine-tuning Flow Matching Generative Models with Intermediate Feedback." arXiv:2510.18072, 2025.
Project page Paper
@article{fan2025acflow,
  title={Fine-tuning Flow Matching Generative Models with Intermediate Feedback},
  author={Fan, Jiajun and Cheng, Chaoran and Shen, Shuaike and Zhou, Xiangxin and Liu, Ge},
  journal={arXiv preprint arXiv:2510.18072},
  year={2025},
  url={https://arxiv.org/abs/2510.18072}
}
NeurIPS 2025

Variational Supervised Contrastive Learning

Ziwen Wang, Jiajun Fan, Thao Nguyen, Heng Ji, Ge Liu. "Variational Supervised Contrastive Learning." NeurIPS 2025.
Paper
@inproceedings{wang2025varcon,
  title={Variational Supervised Contrastive Learning},
  author={Wang, Ziwen and Fan, Jiajun and Nguyen, Thao and Ji, Heng and Liu, Ge},
  booktitle={Advances in Neural Information Processing Systems},
  year={2025},
  url={https://openreview.net/forum?id=uOOlHOq500}
}
NeurIPS 2025

Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models

Jiajun Fan, Tong Wei, Chaoran Cheng, Yuxin Chen, Ge Liu. "Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models." NeurIPS 2025.
Project page Paper
@inproceedings{fan2025adrpo,
  title={Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models},
  author={Fan, Jiajun and Wei, Tong and Cheng, Chaoran and Chen, Yuxin and Liu, Ge},
  booktitle={Advances in Neural Information Processing Systems},
  year={2025},
  url={https://openreview.net/forum?id=aXO0xg0ttW}
}
ICLR 2025

Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization

Jiajun Fan, Shuaike Shen, Chaoran Cheng, Yuxin Chen, Chumeng Liang, Ge Liu. "Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization." ICLR 2025.
Project page Paper
@inproceedings{fan2025orwcfmw2,
  title={Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization},
  author={Fan, Jiajun and Shen, Shuaike and Cheng, Chaoran and Chen, Yuxin and Liang, Chumeng and Liu, Ge},
  booktitle={International Conference on Learning Representations},
  year={2025},
  url={https://openreview.net/forum?id=2IoFFexvuw}
}
2024
AI4Mat-NeurIPS-2024

Efficient Design-and-Control Automation with Reinforcement Learning and Adaptive Exploration

Jiajun Fan, Hongyao Tang, Michael Przystupa, Mariano Phielipp, Santiago Miret, Glen Berseth. "Efficient Design-and-Control Automation with Reinforcement Learning and Adaptive Exploration." AI4Mat-NeurIPS-2024.
Paper
@inproceedings{fan2024edison,
  title={Efficient Design-and-Control Automation with Reinforcement Learning and Adaptive Exploration},
  author={Fan, Jiajun and Tang, Hongyao and Przystupa, Michael and Phielipp, Mariano and Miret, Santiago and Berseth, Glen},
  booktitle={AI for Accelerated Materials Design (AI4Mat), NeurIPS 2024 Workshop},
  year={2024},
  url={https://openreview.net/forum?id=stiehhc5y6}
}
2023
NeurIPS 2023

Optimal Transport for Treatment Effect Estimation

Hao Wang, Jiajun Fan, Zhichao Chen, Haoxuan Li, Weiming Liu, Tianqiao Liu, Quanyu Dai, Yichao Wang, Zhenhua Dong, Ruiming Tang. "Optimal Transport for Treatment Effect Estimation." NeurIPS 2023.
Paper
@inproceedings{wang2023ot,
  title={Optimal Transport for Treatment Effect Estimation},
  author={Wang, Hao and Fan, Jiajun and Chen, Zhichao and Li, Haoxuan and Liu, Weiming and Liu, Tianqiao and Dai, Quanyu and Wang, Yichao and Dong, Zhenhua and Tang, Ruiming},
  booktitle={Advances in Neural Information Processing Systems},
  year={2023},
  url={https://arxiv.org/abs/2310.18286}
}
ICLR 2023 Oral · Top 5%

Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection

Jiajun Fan, Yuzheng Zhuang, Yuecheng Liu, Jianye Hao, Bin Wang, Jiangcheng Zhu, Hao Wang, Shu-Tao Xia. "Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection." ICLR 2023, oral (ranked 5/4176).
Project page Paper
@inproceedings{fan2023lbc,
  title={Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection},
  author={Fan, Jiajun and Zhuang, Yuzheng and Liu, Yuecheng and Hao, Jianye and Wang, Bin and Zhu, Jiangcheng and Wang, Hao and Xia, Shu-Tao},
  booktitle={International Conference on Learning Representations},
  year={2023},
  note={Oral, Top 5\%},
  url={https://openreview.net/forum?id=FeWvD0L_a4}
}
2022
ICML 2022

Generalized Data Distribution Iteration

Jiajun Fan, Changnan Xiao, "Generalized Data Distribution Iteration." In the proceedings of International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, 2022.
Project page Paper
@inproceedings{fan2022gdi,
  title={Generalized Data Distribution Iteration},
  author={Fan, Jiajun and Xiao, Changnan},
  booktitle={International Conference on Machine Learning},
  year={2022},
  url={https://proceedings.mlr.press/v162/fan22c.html}
}
arXiv

Entire Space Counterfactual Learning: Tuning, Analytical Properties and Industrial Applications

Hao Wang, Zhichao Chen, Jiajun Fan, Yuxin Huang, Weiming Liu, Xinggao Liu, "Entire Space Counterfactual Learning: Tuning, Analytical Properties and Industrial Applications." Arxiv, 2022.
Paper
@article{wang2022escl,
  title={Entire Space Counterfactual Learning: Tuning, Analytical Properties and Industrial Applications},
  author={Wang, Hao and Chen, Zhichao and Fan, Jiajun and Huang, Yuxin and Liu, Weiming and Liu, Xinggao},
  journal={arXiv preprint arXiv:2210.11039},
  year={2022}
}
NeurIPS 2022 Workshop

CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration

Changnan Xiao, Haosen Shi, Jiajun Fan, Shihong Deng, Haiyan Yin, "CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration." In the proceedings of Deep Reinforcement Learning Workshop NeurIPS 2022, 2022.
Paper
@inproceedings{xiao2022casaneurips,
  title={CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration},
  author={Xiao, Changnan and Shi, Haosen and Fan, Jiajun and Deng, Shihong and Yin, Haiyan},
  booktitle={Deep Reinforcement Learning Workshop, NeurIPS 2022},
  year={2022}
}
AAAI-22 Workshop

GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning

Jiajun Fan, Changnan Xiao, Yue Huang, "GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning." In the proceedings of AAAI-22 Workshop on Reinforcement Learning in Games, 2022.
Project page Paper
@inproceedings{fan2022gdiworkshop,
  title={GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning},
  author={Fan, Jiajun and Xiao, Changnan and Huang, Yue},
  booktitle={AAAI-22 Workshop on Reinforcement Learning in Games},
  year={2022}
}
AAAI-22 Workshop

A Review for Deep Reinforcement Learning in Atari: Benchmarks, Challenges, and Solutions

Jiajun Fan, "A Review for Deep Reinforcement Learning in Atari: Benchmarks, Challenges, and Solutions." In the proceedings of AAAI-22 Workshop on Reinforcement Learning in Games, 2022.
Paper
@inproceedings{fan2022atarireview,
  title={A Review for Deep Reinforcement Learning in Atari: Benchmarks, Challenges, and Solutions},
  author={Fan, Jiajun},
  booktitle={AAAI-22 Workshop on Reinforcement Learning in Games},
  year={2022}
}
2021
arXiv

An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning

Changnan Xiao, Haosen Shi, Jiajun Fan, Shihong Deng, "An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning." Arxiv, 2021.
Paper
@article{xiao2021entropy,
  title={An Entropy Regularization Free Mechanism for Policy-based Reinforcement Learning},
  author={Xiao, Changnan and Shi, Haosen and Fan, Jiajun and Deng, Shihong},
  journal={arXiv preprint arXiv:2106.00707},
  year={2021}
}
2020
arXiv

Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning

Jiajun Fan, He Ba, Xian Guo, Jianye Hao, "Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning." Arxiv, 2020.
Paper
@article{fan2020criticpi2,
  title={Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning},
  author={Fan, Jiajun and Ba, He and Guo, Xian and Hao, Jianye},
  journal={arXiv preprint arXiv:2011.06752},
  year={2020}
}

Citation counts update daily from Google Scholar. Preprints and workshop papers are included; see Google Scholar for the complete record.