TL;DR — RL with verifiable rewards teaches a tool-using agent what a correct call looks like, but never how to sequence calls. We add a procedural reward — invocation-order quality plus a final-system-state check — and an error-aware penalty to GRPO. A 3B model then beats correctness-only RL on all three benchmarks, and lands in the same range as much larger proprietary models.
The problem: correct calls, wrong order
In state-aware environments the order of tool calls decides whether the final system state is right. Wrapping a gift before looking up the address is fine; shipping it first is not. Correctness-only rewards score the tool name and its parameters, so a perfectly-formed call in the wrong position earns the same credit as the right one — and one misordered call cascades into every later step.
Correctness-only RLVR
Matches tool names and parameter values against the ground truth. Execution order, failed calls and the resulting system state are invisible to the reward, so the model fits what a call looks like rather than how to reach the goal.
Procedure-aware RL
Scores the invocation chain itself: graded credit for partial ordering, a check that the trajectory ends in the correct state, and an explicit penalty for structural call errors — tool calling becomes a trainable procedural skill.
Method
The policy is fine-tuned with GRPO (KL-regularized) on a weighted sum of four reward terms. Only the reward changes — the RL algorithm stays standard.
Format
Reasoning and tool-call tags are well-formed and parsable.
Correctness
Tool names, parameter keys and values match the ground truth.
Procedural
NDCG over the invocation order, plus a check that the trajectory ends in the correct state.
Error-aware
Penalty for structural invocation errors: wrong name, wrong parameter keys, wrong types.
The procedural term uses normalized discounted cumulative gain so that partially-correct orderings are ranked above invalid ones, instead of collapsing to a binary match. The state check gives credit whenever the agent eventually reaches the correct final state, which keeps reasonable detours — listing a directory before creating a file, or recovering from a failed call — from being punished as mistakes.
Results
| Model | BFCL-V3 avg. acc | Multi-turn acc |
|---|---|---|
| Qwen2.5-3B-Instruct (base) | 57.31% | 1.00% |
| ToolRL (correctness-only RL, 3B) | 63.07% | 3.88% |
| GPT-5-mini | 57.11% | 5.50% |
| GPT-4.1-mini | 64.65% | 2.50% |
| Gemma-3-12b-it | 64.69% | 5.75% |
| Procedure-Aware RL (ours, 3B) | 65.67% Best | 7.25% |
Proprietary and larger open models are evaluated under BFCL's official per-model protocols, so they are a reference point rather than a matched comparison. The matched comparison is the base model and ToolRL, which share the Qwen2.5 backbone.
Both new terms are load-bearing
Dropping the error-aware penalty mostly hurts single-turn accuracy, where invocation errors go unpunished; dropping the procedural reward mostly hurts the multi-turn split, where ordering is what matters. All differences are significant under a paired bootstrap on per-sample outcomes.
What the model learns
Rewarding sequence quality buys behaviour nobody wrote into the reward. On a multi-turn file-system task the trained agent navigates before it creates — cd first, then the file — while the correctness-only baseline recreates a directory that already exists and then cascades through the resulting state errors. The same training also improves recovery: after a failed call the agent retries along a corrected path instead of repeating the failure.
BibTeX
@inproceedings{zheng2026procedure,
title = {Procedure-Aware Reinforcement Learning for Tool-Augmented
Large Language Models},
author = {Qinglong Zheng and Jiajun Fan and Chaoran Cheng and Ge Liu},
booktitle = {Conference on Language Modeling (COLM)},
year = {2026},
url = {https://openreview.net/forum?id=d4wrBuJ4xo}
}