Batched Contextual Reinforcement
Published in International Conference on Machine Learning (ICML 2026), 2026
Recommended citation: Bangji Yang, Hongbo Ma, Jiajun Fan, Ge Liu. "Batched Contextual Reinforcement." ICML 2026. https://openreview.net/forum?id=8Oc3Mx754M
BCR is a single-stage training paradigm that trains a model to solve N problems at once inside a shared context window, rewarded purely by per-instance accuracy. The implicit token budget this creates cuts token usage by 15.8–62.6% while maintaining or improving accuracy across five mathematical benchmarks, and avoids the optimization collapse that explicit length penalties induce.
