TL;DR — RL with verifiable rewards teaches a tool-using agent what a correct call looks like, but never how to sequence calls. We add a procedural reward — invocation-order quality plus a final-system-state check — and an error-aware penalty to GRPO. A 3B model then beats correctness-only RL on all three benchmarks, and lands in the same range as much larger proprietary models.

The problem: correct calls, wrong order

In state-aware environments the order of tool calls decides whether the final system state is right. Wrapping a gift before looking up the address is fine; shipping it first is not. Correctness-only rewards score the tool name and its parameters, so a perfectly-formed call in the wrong position earns the same credit as the right one — and one misordered call cascades into every later step.

Correctness-only RLVR

Matches tool names and parameter values against the ground truth. Execution order, failed calls and the resulting system state are invisible to the reward, so the model fits what a call looks like rather than how to reach the goal.

Procedure-aware RL

Scores the invocation chain itself: graded credit for partial ordering, a check that the trajectory ends in the correct state, and an explicit penalty for structural call errors — tool calling becomes a trainable procedural skill.

Method

The policy is fine-tuned with GRPO (KL-regularized) on a weighted sum of four reward terms. Only the reward changes — the RL algorithm stays standard.

1

Format

Reasoning and tool-call tags are well-formed and parsable.

2

Correctness

Tool names, parameter keys and values match the ground truth.

3

Procedural

NDCG over the invocation order, plus a check that the trajectory ends in the correct state.

4

Error-aware

Penalty for structural invocation errors: wrong name, wrong parameter keys, wrong types.

2026-09-30T20:54:16.488050 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ procedure-aware reward, evaluated on every rollout u p d a t e   ,   r o l l   o u t   a g a i n θ query + tools multi-turn history p o l i c y     π θ think · tool_call · response GRPO update weighted total reward format reward tags well-formed correctness reward tool name · params procedural reward NDCG order + final state error-aware penalty name / key / type errors
Figure 1. Procedure-aware agentic RL. Every rollout is scored on four axes; the weighted total drives a standard GRPO update. The procedural reward (blue) and the error-aware penalty (red) are what correctness-only RL leaves out.

The procedural term uses normalized discounted cumulative gain so that partially-correct orderings are ranked above invalid ones, instead of collapsing to a binary match. The state check gives credit whenever the agent eventually reaches the correct final state, which keeps reasonable detours — listing a directory before creating a file, or recovering from a failed call — from being punished as mistakes.

Results

+2.6
BFCL-V3 avg. acc over ToolRL (3B)
+5.7
API-Bank avg. acc over ToolRL
+8.4
Bamboogle accuracy over ToolRL
3B
matches far larger proprietary models
2026-09-30T20:59:04.482014 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ BFCL-V3 (avg. acc) API-Bank (avg. acc) Bamboogle (accuracy) 0 20 40 60 80 accuracy (%) 57.3 27.3 48.8 63.1 57.9 48.4 65.7 63.7 56.8 Qwen2.5-3B (base) ToolRL (correctness-only RL) Procedure-Aware RL (ours)
Figure 2. Qwen2.5-3B-Instruct trained with procedure-aware rewards, against the same base model and ToolRL's released correctness-only checkpoint, on three tool-use benchmarks.
ModelBFCL-V3 avg. accMulti-turn acc
Qwen2.5-3B-Instruct (base)57.31%1.00%
ToolRL (correctness-only RL, 3B)63.07%3.88%
GPT-5-mini57.11%5.50%
GPT-4.1-mini64.65%2.50%
Gemma-3-12b-it64.69%5.75%
Procedure-Aware RL (ours, 3B)65.67% Best7.25%

Proprietary and larger open models are evaluated under BFCL's official per-model protocols, so they are a reference point rather than a matched comparison. The matched comparison is the base model and ToolRL, which share the Qwen2.5 backbone.

Both new terms are load-bearing

2026-09-30T20:59:04.500856 image/svg+xml Matplotlib v3.10.8, https://matplotlib.org/ ToolRL-3B (no procedural, no penalty) w/o procedural reward w/o error-aware penalty Full model (ours) 50 52 54 BFCL-V3 avg. acc (%) 50.73 51.65 51.58 53.92
Figure 3. Removing either term costs more than two points, and both together account for the whole gap to correctness-only RL. Scores here average the three BFCL sub-categories (non-live, live, multi-turn), so they are not directly comparable with the table above.

Dropping the error-aware penalty mostly hurts single-turn accuracy, where invocation errors go unpunished; dropping the procedural reward mostly hurts the multi-turn split, where ordering is what matters. All differences are significant under a paired bootstrap on per-sample outcomes.

What the model learns

Rewarding sequence quality buys behaviour nobody wrote into the reward. On a multi-turn file-system task the trained agent navigates before it creates — cd first, then the file — while the correctness-only baseline recreates a directory that already exists and then cascades through the resulting state errors. The same training also improves recovery: after a failed call the agent retries along a corrected path instead of repeating the failure.

BibTeX

@inproceedings{zheng2026procedure,
  title     = {Procedure-Aware Reinforcement Learning for Tool-Augmented
               Large Language Models},
  author    = {Qinglong Zheng and Jiajun Fan and Chaoran Cheng and Ge Liu},
  booktitle = {Conference on Language Modeling (COLM)},
  year      = {2026},
  url       = {https://openreview.net/forum?id=d4wrBuJ4xo}
}