Reinforcement learning for generative models and reasoning LLMs, agentic RL, and efficient inference. Listed newest first; my name is bold in each author list.
Google Scholar482 citations12 h-index12 i10-index23 papers
Research overview
I study how to train multimodal large generative/agentic models with reinforcement learning
in a way that is stable (e.g., lifelong continual learning) — avoiding the diversity and
performance collapse that RL fine-tuning tends to cause — and progressively autonomous
(e.g., self-evolving), steadily removing the human data collection and labelling the pipeline
depends on. That work spans these domains:
Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models
Qinglong Zheng, Jiajun Fan, Chaoran Cheng, Ge Liu. "Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models." COLM 2026.
@inproceedings{zheng2026procedureaware,
title={Procedure-Aware Reinforcement Learning for Tool-Augmented Large Language Models},
author={Zheng, Qinglong and Fan, Jiajun and Cheng, Chaoran and Liu, Ge},
booktitle={Conference on Language Modeling},
year={2026}
}
Ziwen Wang, Jiajun Fan, Ruihan Guo, Thao Nguyen, Heng Ji, Ge Liu. "ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning." Transactions on Machine Learning Research (TMLR), 2026.
@article{wang2025proteinzero,
title={ProteinZero: Self-Improving Protein Generation via Online Reinforcement Learning},
author={Wang, Ziwen and Fan, Jiajun and Guo, Ruihan and Nguyen, Thao and Ji, Heng and Liu, Ge},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2026},
url={https://arxiv.org/abs/2506.07459}
}
@article{li2026elegantvla,
title={ElegantVLA: Learning When to Think for Efficient Vision-Language-Action Models},
author={Li, Ye and Liu, Huanan and Ji, Kangye and Meng, Yuan and Fan, Jiajun and Wang, Yuansong and Qin, Shiyu and Wu, Chenglei and Xia, Shu-Tao and Wang, Zhi},
journal={arXiv preprint arXiv:2605.29438},
year={2026},
url={https://arxiv.org/abs/2605.29438}
}
@article{sun2026pyrag,
title={Retrieval is Cheap, Show Me the Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation},
author={Sun, Jiashuo and Shi, Jimeng and Xie, Yixuan and Wang, Saizhuo and Parekh, Jash Rajesh and Jiang, Pengcheng and Shi, Zhiyi and Fan, Jiajun and Zheng, Qinglong and Li, Peiran and Wang, Shaowen and Liu, Ge and Han, Jiawei},
journal={arXiv preprint arXiv:2605.12975},
year={2026},
url={https://arxiv.org/abs/2605.12975}
}
@inproceedings{sun2026barrag,
title={Rethinking the Reranker: Boundary-Aware Evidence Selection for Robust Retrieval-Augmented Generation},
author={Sun, Jiashuo and Jiang, Pengcheng and Wang, Saizhuo and Fan, Jiajun and Wang, Heng and Ouyang, Siru and Zhong, Ming and Jiao, Yizhu and Huang, Chengsong and Xu, Xueqiang and Han, Pengrui and Li, Peiran and Huang, Jiaxin and Liu, Ge and Ji, Heng and Han, Jiawei},
booktitle={International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=Tt8lCe1NrW}
}
Ye Li, Yuan Meng, Zewen Sun, Kangye Ji, Chen Tang, Jiajun Fan, Xinzhu Ma, Shu-Tao Xia, Zhi Wang, Wenwu Zhu. "SP-VLA: A Joint Model Scheduling and Token Pruning Approach for VLA Model Acceleration." ICLR 2026.
@inproceedings{li2026spvla,
title={{SP-VLA}: A Joint Model Scheduling and Token Pruning Approach for {VLA} Model Acceleration},
author={Li, Ye and Meng, Yuan and Sun, Zewen and Ji, Kangye and Tang, Chen and Fan, Jiajun and Ma, Xinzhu and Xia, Shu-Tao and Wang, Zhi and Zhu, Wenwu},
booktitle={International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=RwdGIIjPlC}
}
@article{li2026prance,
title={{PRANCE}: Joint Token-Optimization and Structural Channel-Pruning for Adaptive {ViT} Inference},
author={Li, Ye and Tang, Chen and Meng, Yuan and Fan, Jiajun and Chai, Zenghao and Ma, Xinzhu and Wang, Zhi and Zhu, Wenwu},
journal={IEEE Transactions on Pattern Analysis and Machine Intelligence},
year={2026},
url={https://arxiv.org/abs/2407.05010}
}
Jiajun Fan, Roger Ren, Jingyuan Li, Rahul Pandey, Prashanth G. Shivakumar, Ivan Bulyko, Ankur Gandhe, Ge Liu, Yile Gu. "Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio LLMs via Reasoning Process Rewards." ICLR 2026.
@inproceedings{fan2026cesar,
title={Incentivizing Consistent, Effective and Scalable Reasoning Capability in Audio {LLMs} via Reasoning Process Rewards},
author={Fan, Jiajun and Ren, Roger and Li, Jingyuan and Pandey, Rahul and Shivakumar, Prashanth G. and Bulyko, Ivan and Gandhe, Ankur and Liu, Ge and Gu, Yile},
booktitle={International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=DUr48hxO2h}
}
@inproceedings{wang2025varcon,
title={Variational Supervised Contrastive Learning},
author={Wang, Ziwen and Fan, Jiajun and Nguyen, Thao and Ji, Heng and Liu, Ge},
booktitle={Advances in Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=uOOlHOq500}
}
@inproceedings{fan2025adrpo,
title={Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative Models},
author={Fan, Jiajun and Wei, Tong and Cheng, Chaoran and Chen, Yuxin and Liu, Ge},
booktitle={Advances in Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=aXO0xg0ttW}
}
@inproceedings{fan2025orwcfmw2,
title={Online Reward-Weighted Fine-Tuning of Flow Matching with Wasserstein Regularization},
author={Fan, Jiajun and Shen, Shuaike and Cheng, Chaoran and Chen, Yuxin and Liang, Chumeng and Liu, Ge},
booktitle={International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=2IoFFexvuw}
}
@inproceedings{fan2024edison,
title={Efficient Design-and-Control Automation with Reinforcement Learning and Adaptive Exploration},
author={Fan, Jiajun and Tang, Hongyao and Przystupa, Michael and Phielipp, Mariano and Miret, Santiago and Berseth, Glen},
booktitle={AI for Accelerated Materials Design (AI4Mat), NeurIPS 2024 Workshop},
year={2024},
url={https://openreview.net/forum?id=stiehhc5y6}
}
@inproceedings{wang2023ot,
title={Optimal Transport for Treatment Effect Estimation},
author={Wang, Hao and Fan, Jiajun and Chen, Zhichao and Li, Haoxuan and Liu, Weiming and Liu, Tianqiao and Dai, Quanyu and Wang, Yichao and Dong, Zhenhua and Tang, Ruiming},
booktitle={Advances in Neural Information Processing Systems},
year={2023},
url={https://arxiv.org/abs/2310.18286}
}
@inproceedings{fan2023lbc,
title={Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection},
author={Fan, Jiajun and Zhuang, Yuzheng and Liu, Yuecheng and Hao, Jianye and Wang, Bin and Zhu, Jiangcheng and Wang, Hao and Xia, Shu-Tao},
booktitle={International Conference on Learning Representations},
year={2023},
note={Oral, Top 5\%},
url={https://openreview.net/forum?id=FeWvD0L_a4}
}
Jiajun Fan, Changnan Xiao, "Generalized Data Distribution Iteration." In the proceedings of International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA, 2022.
@inproceedings{fan2022gdi,
title={Generalized Data Distribution Iteration},
author={Fan, Jiajun and Xiao, Changnan},
booktitle={International Conference on Machine Learning},
year={2022},
url={https://proceedings.mlr.press/v162/fan22c.html}
}
@article{wang2022escl,
title={Entire Space Counterfactual Learning: Tuning, Analytical Properties and Industrial Applications},
author={Wang, Hao and Chen, Zhichao and Fan, Jiajun and Huang, Yuxin and Liu, Weiming and Liu, Xinggao},
journal={arXiv preprint arXiv:2210.11039},
year={2022}
}
Changnan Xiao, Haosen Shi, Jiajun Fan, Shihong Deng, Haiyan Yin, "CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration." In the proceedings of Deep Reinforcement Learning Workshop NeurIPS 2022, 2022.
@inproceedings{xiao2022casaneurips,
title={CASA: Bridging the Gap between Policy Improvement and Policy Evaluation with Conflict Averse Policy Iteration},
author={Xiao, Changnan and Shi, Haosen and Fan, Jiajun and Deng, Shihong and Yin, Haiyan},
booktitle={Deep Reinforcement Learning Workshop, NeurIPS 2022},
year={2022}
}
Jiajun Fan, Changnan Xiao, Yue Huang, "GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning." In the proceedings of AAAI-22 Workshop on Reinforcement Learning in Games, 2022.
@inproceedings{fan2022gdiworkshop,
title={GDI: Rethinking What Makes Reinforcement Learning Different From Supervised Learning},
author={Fan, Jiajun and Xiao, Changnan and Huang, Yue},
booktitle={AAAI-22 Workshop on Reinforcement Learning in Games},
year={2022}
}
Jiajun Fan, "A Review for Deep Reinforcement Learning in Atari: Benchmarks, Challenges, and Solutions." In the proceedings of AAAI-22 Workshop on Reinforcement Learning in Games, 2022.
@inproceedings{fan2022atarireview,
title={A Review for Deep Reinforcement Learning in Atari: Benchmarks, Challenges, and Solutions},
author={Fan, Jiajun},
booktitle={AAAI-22 Workshop on Reinforcement Learning in Games},
year={2022}
}
@article{fan2020criticpi2,
title={Critic PI2: Master Continuous Planning via Policy Improvement with Path Integrals and Deep Actor-Critic Reinforcement Learning},
author={Fan, Jiajun and Ba, He and Guo, Xian and Hao, Jianye},
journal={arXiv preprint arXiv:2011.06752},
year={2020}
}
Citation counts update daily from Google Scholar. Preprints and workshop papers are included; see Google Scholar for the complete record.