Reasoning Made Audio LLMs Worse. We Flipped the Sign.
Chain-of-thought is the reflex answer for making a model smarter. In Audio LLMs it backfires: tell the model to think first and it gets worse, and the longer it thinks the worse it gets. This note is about why that happens, why it is not an argument against reasoning, and what changes when you supervise the reasoning process instead of only the final answer.
Flip the switch. Every number is MMAU Test-mini total accuracy from Table 1 of the paper — the left state is the model answering directly, the right state is the same model told to reason first. The untrained base model falls 3.40 points for thinking. After training on the reasoning process, CESAR gains 3.40 for it: the same margin, sign reversed. That reversal is what “solving test-time inverse scaling” means.
The short version
Six things this paper pinned down
Each is stated with the number that carries it, and shown in full further down.
Ask Qwen2.5-Omni-7B to think before answering and MMAU accuracy falls from 68.60 to 65.20. Longer chains kept making it worse. We name this test-time inverse scaling.
Models never taught how to reason produce hallucinatory, inconsistent chains whose errors compound. The capacity is there; the supervision was not.
Outcome-only RLVR scores the final token and ignores the road to it. CESAR’s reward has five terms, three of them process-level — reasoning–answer consistency; structure, logic and domain grounding in one term; and a penalty on overthinking.
Under the same test, CESAR gains +3.40 from reasoning — exactly the margin the base model lost. Reasoning stops being a liability and becomes the source of the gain.
77.10 on MMAU Test-mini against 71.60 and 62.50, and the top-scoring 7B model on the harder MMAU-Pro at 56.4.
On MMSU, reasoning reaches 81.07 against a human 86.77 — and perception rose too, though it stays far behind humans. That gap is the real bottleneck.
01 — The paradox
More thinking, less accuracy
Text LLMs made chain-of-thought look like a free lunch: o1 and DeepSeek-R1 turned longer deliberation into better answers. Carrying the same prompt into audio produces the opposite. On the open Audio LLM this work starts from, switching reasoning on costs 3.40 points, and an outcome-only RL baseline recovers only to break-even; sweeping the maximum thinking length makes the loss deepen rather than recover. We call it test-time inverse scaling, and it is the first thing a process-level view has to explain.
02 — The diagnosis
The chains were never trained, only permitted
It would be easy to read that as evidence that audio reasoning is a dead end. The chains themselves say otherwise. Supervised fine-tuning on CoT data teaches a model to imitate the shape of reasoning; outcome-only RL rewards it only for landing on the right option. Neither ever looks at the middle. Three failure modes follow directly, and all three are visible in the traces.
03 — The fix
Grade the road, not just the destination
CESAR keeps GRPO and keeps verifiable correctness, then stops treating the reasoning trace as a black box between prompt and answer. Five terms make up the total reward. The first two are the conventional verifiable pair; the last three are the ones that make reasoning a trainable skill rather than an emergent accident.
Training is GRPO on Qwen2.5-Omni-7B, 8 sampled responses per example, on AVQA augmented with answer-invariant rephrasings so the model has to learn the reasoning rather than the wording.
The whole of Table 1, not just the headline. Switch tracks and the pattern holds: on every one of the four, the base model loses ground by reasoning and CESAR gains it. Sound is where the gap is widest — 83.48 against the base model’s 69.07 — and Speech is where the overthinking penalty costs something, which is why the variant without it leads there. Proprietary systems are shown for scale and have no reasoning switch to flip.
04 — What changed
The sign flips, and a 7B model clears the proprietary field
The headline is not that accuracy went up. It is that the relationship between thinking and accuracy inverted — the mirrored bars at the top of this page. What follows from that inversion is the part that matters for anyone choosing a model.
05 — The sweet spot
Trained reasoning has an optimal depth, and finds it
Sweeping the maximum thinking length from 0 to 250 tokens separates the two regimes cleanly. Baselines either collapse or wander with no reliable gain. Our variant without the overthinking penalty climbs steadily to a 76.50% peak. The full method, penalised for rambling, peaks higher at 77.10% using a chain of only about 35–40 tokens. The penalty is not a tax on thinking; it is what teaches the model when to stop.
06 — The ceiling
Near-human reasoning, and a perceptual wall behind it
MMSU separates what a model hears from what it concludes, and the split is stark. CESAR’s reasoning lands within six points of expert humans — and beats them outright on semantic reasoning, 88.72 against 82.16. Perception is another story: 48.45 against a human 91.24. Training the process lifted perception too, which is the genuinely surprising part, but the gap that remains is the one that will decide how far audio reasoning can go.
Reasoning is close to human; perception is nowhere near. Each point is one system on MMSU, reasoning on the horizontal axis and perception on the vertical. Every model sits far below the dashed parity line — they all reason far better than they hear, and only the human reference clears it. CESAR posts the best reasoning score of any model here, 81.07, within six points of the human 86.77 — while its perception, at 48.45, is 42.79 points short of the human 91.24. That asymmetry is the finding: process rewards move reasoning close to human parity and lift perception with it, but the perceptual bottleneck is what remains.
07 — The verdict
Three thousand human judgements, blind
Accuracy says the answers improved. It cannot say the reasoning did. So the full 1,000-question MMAU Test-mini set was judged by three independent expert annotators — over 3,000 individual judgements, blind to which model wrote which trace and blind to the correct answer, scoring only which reasoning process was sounder.
08 — What it means
Reasoning is a skill you supervise, not a switch you flip
The tidy reading of test-time inverse scaling was that audio is simply not a reasoning-friendly modality. The result here says something narrower and more useful: what fails is reasoning nobody trained. Give the process its own reward — consistency with its own conclusion, analytical structure, domain grounding, and a real cost for rambling — and the same 7B model that lost 3.40 points to thinking gains 3.40, passes systems many times its size, and reasons within touching distance of expert humans.
The wall it runs into next is not cognitive. It is perceptual: 48.45 against 91.24. The next gain in audio reasoning will not come from thinking harder about what the model heard. It will come from hearing it better.