Jiajun Fan
Research Notes

Reasoning Made Audio LLMs Worse. We Flipped the Sign.


Chain-of-thought is the reflex answer for making a model smarter. In Audio LLMs it backfires: tell the model to think first and it gets worse, and the longer it thinks the worse it gets. This note is about why that happens, why it is not an argument against reasoning, and what changes when you supervise the reasoning process instead of only the final answer.

Same model, same benchmark — one instruction apart · bars span 62–79%
Qwen2.5-Omni-7Bbase model
65.20−3.40
Ke-Omni-Routcome-only RL
74.60+0.10
CESAR w/o OPours
76.50+1.50
CESARours
77.10+3.40

Flip the switch. Every number is MMAU Test-mini total accuracy from Table 1 of the paper — the left state is the model answering directly, the right state is the same model told to reason first. The untrained base model falls 3.40 points for thinking. After training on the reasoning process, CESAR gains 3.40 for it: the same margin, sign reversed. That reversal is what “solving test-time inverse scaling” means.

The short version

Six things this paper pinned down

Each is stated with the number that carries it, and shown in full further down.

1
Reasoning made audio models worse

Ask Qwen2.5-Omni-7B to think before answering and MMAU accuracy falls from 68.60 to 65.20. Longer chains kept making it worse. We name this test-time inverse scaling.

§01 · the paradox
2
The culprit is untrained reasoning, not reasoning

Models never taught how to reason produce hallucinatory, inconsistent chains whose errors compound. The capacity is there; the supervision was not.

§02 · the diagnosis
3
So reward the process, not just the answer

Outcome-only RLVR scores the final token and ignores the road to it. CESAR’s reward has five terms, three of them process-level — reasoning–answer consistency; structure, logic and domain grounding in one term; and a penalty on overthinking.

§03 · the fix
4
The sign flips

Under the same test, CESAR gains +3.40 from reasoning — exactly the margin the base model lost. Reasoning stops being a liability and becomes the source of the gain.

→ §04 · what changed
5
A 7B model passes Gemini 2.5 Pro and GPT-4o Audio

77.10 on MMAU Test-mini against 71.60 and 62.50, and the top-scoring 7B model on the harder MMAU-Pro at 56.4.

→ §04 · what changed
6
Better reasoning also sharpened perception

On MMSU, reasoning reaches 81.07 against a human 86.77 — and perception rose too, though it stays far behind humans. That gap is the real bottleneck.

§06 · the ceiling

01 — The paradox

More thinking, less accuracy

Text LLMs made chain-of-thought look like a free lunch: o1 and DeepSeek-R1 turned longer deliberation into better answers. Carrying the same prompt into audio produces the opposite. On the open Audio LLM this work starts from, switching reasoning on costs 3.40 points, and an outcome-only RL baseline recovers only to break-even; sweeping the maximum thinking length makes the loss deepen rather than recover. We call it test-time inverse scaling, and it is the first thing a process-level view has to explain.

MMAU Test-mini — total accuracy1k expertly annotated questions · 27 reasoning skills · bars span 58–80%
CESARours · with reasoning
77.10
CESAR w/o OPours · no overthinking penalty
76.50
Ke-Omni-Routcome-only RL
74.60
Gemini 2.5 Flashproprietary
71.80
Gemini 2.5 Proproprietary
71.60
Qwen2.5-Omni-7Bbase, answering directly
68.60
GPT-4o Audioproprietary
62.50
CESAR (7B, ours)baselines and proprietary systems

02 — The diagnosis

The chains were never trained, only permitted

It would be easy to read that as evidence that audio reasoning is a dead end. The chains themselves say otherwise. Supervised fine-tuning on CoT data teaches a model to imitate the shape of reasoning; outcome-only RL rewards it only for landing on the right option. Neither ever looks at the middle. Three failure modes follow directly, and all three are visible in the traces.

failure mode 01
Reasoning appears at random
Nothing in an outcome-only objective makes a reasoning pattern reliable. Useful analysis shows up when it happens to, and cannot be summoned on demand.
failure mode 02
The answer contradicts the reasoning
The most damaging one. A model correctly identifies “three rings” in its trace and then emits 2 as its answer. Outcome-only reward is blind to this: it grades only the 2.
failure mode 03
No analytical structure
Without pressure toward elimination, comparison or multi-step deduction, chains drift into free association — and every extra token is another chance to hallucinate.
consequence
Errors compound with length
Put the three together and length becomes a liability: a longer chain is simply more unsupervised steps, each able to derail the next. That is the inverse-scaling curve.
If the deficit were reasoning capacity, the fix would be a bigger model. It is reasoning supervision — and that is something a reward can carry.

03 — The fix

Grade the road, not just the destination

CESAR keeps GRPO and keeps verifiable correctness, then stops treating the reasoning trace as a black box between prompt and answer. Five terms make up the total reward. The first two are the conventional verifiable pair; the last three are the ones that make reasoning a trainable skill rather than an emergent accident.

verifiable · weight 5.0
Answer correctness
The anchor. A binary check that the chosen option is right, weighted 5× everything else so process credit can never buy a wrong answer.
verifiable · weight 1.0
Format compliance
Output must carry a real <think> block and a real <answer> block. Without it a model simply routes around the reasoning it is being trained on.
process · weight 1.0
Reasoning–answer consistency
Concept overlap measured twice: thought against answer, and thought against the question. This is the term that kills the classic failure where a model reasons its way to “three rings” and then outputs “2”.
process · weight 1.0
Structure, logic, domain
One reward with three parts: analytical patterns (sequential organisation, comparison, elimination), logical markers (deduction, hypothesis, evidence), and audio-domain vocabulary — acoustic, musical, phonetic terms.
process · weight 1.0
Overthinking penalty
A linear cost on chain length, 1 − |t| / 256. Rambling is where hallucinations accumulate; this is what buys the short, decisive chains.

Training is GRPO on Qwen2.5-Omni-7B, 8 sampled responses per example, on AVQA augmented with answer-invariant rephrasings so the model has to learn the reasoning rather than the wording.

MMAU Test-mini, every track — and what reasoning is worth on each · bars span 55–86%
The model is:

oursevery other systemreasoning costs this model points on the selected track

The whole of Table 1, not just the headline. Switch tracks and the pattern holds: on every one of the four, the base model loses ground by reasoning and CESAR gains it. Sound is where the gap is widest — 83.48 against the base model’s 69.07 — and Speech is where the overthinking penalty costs something, which is why the variant without it leads there. Proprietary systems are shown for scale and have no reasoning switch to flip.

04 — What changed

The sign flips, and a 7B model clears the proprietary field

The headline is not that accuracy went up. It is that the relationship between thinking and accuracy inverted — the mirrored bars at the top of this page. What follows from that inversion is the part that matters for anyone choosing a model.

77.10%
MMAU Test-mini — SOTA, above Gemini 2.5 Pro (71.60) and GPT-4o Audio (62.50)
56.4%
MMAU-Pro average — best of any 7B model on the in-the-wild benchmark
+3.40
points gained from reasoning, against −3.40 for the untrained base model

05 — The sweet spot

Trained reasoning has an optimal depth, and finds it

Sweeping the maximum thinking length from 0 to 250 tokens separates the two regimes cleanly. Baselines either collapse or wander with no reliable gain. Our variant without the overthinking penalty climbs steadily to a 76.50% peak. The full method, penalised for rambling, peaks higher at 77.10% using a chain of only about 35–40 tokens. The penalty is not a tax on thinking; it is what teaches the model when to stop.

Test-time scaling was never unavailable to Audio LLMs. It was unavailable to untrained reasoning — and it returns the moment the process is supervised.

06 — The ceiling

Near-human reasoning, and a perceptual wall behind it

MMSU separates what a model hears from what it concludes, and the split is stark. CESAR’s reasoning lands within six points of expert humans — and beats them outright on semantic reasoning, 88.72 against 82.16. Perception is another story: 48.45 against a human 91.24. Training the process lifted perception too, which is the genuinely surprising part, but the gap that remains is the one that will decide how far audio reasoning can go.

7075808590405060708090equal on both42.79 pointsof perception missing123456the six points1CESARours · 81.07 / 48.452Qwen2.5-Omni-7Bbase · 79.83 / 42.503Ke-Omni-Routcome-only RL · 78.06 / 47.094Gemini 1.5 Proproprietary · 76.16 / 46.105GPT-4o Audioproprietary · 71.96 / 39.676Humanexpert · 86.77 / 91.24reasoning →perception →

Reasoning is close to human; perception is nowhere near. Each point is one system on MMSU, reasoning on the horizontal axis and perception on the vertical. Every model sits far below the dashed parity line — they all reason far better than they hear, and only the human reference clears it. CESAR posts the best reasoning score of any model here, 81.07, within six points of the human 86.77 — while its perception, at 48.45, is 42.79 points short of the human 91.24. That asymmetry is the finding: process rewards move reasoning close to human parity and lift perception with it, but the perceptual bottleneck is what remains.

07 — The verdict

Three thousand human judgements, blind

Accuracy says the answers improved. It cannot say the reasoning did. So the full 1,000-question MMAU Test-mini set was judged by three independent expert annotators — over 3,000 individual judgements, blind to which model wrote which trace and blind to the correct answer, scoring only which reasoning process was sounder.

Human preference, majority vote1,000 questions · 3 annotators each · win / lose / tie
vs Qwen2.5-Omni-7Bbase model
88.60% win
6.60% lose
4.80% tie
vs Ke-Omni-Routcome-only RL
63.10% win
14.80% lose
22.10% tie
CESAR preferredbaseline preferredtiethe second row is the one that matters: it beats an outcome-only RL baseline trained on the same AVQA corpus — plus extra in-domain music data

08 — What it means

Reasoning is a skill you supervise, not a switch you flip

The tidy reading of test-time inverse scaling was that audio is simply not a reasoning-friendly modality. The result here says something narrower and more useful: what fails is reasoning nobody trained. Give the process its own reward — consistency with its own conclusion, analytical structure, domain grounding, and a real cost for rambling — and the same 7B model that lost 3.40 points to thinking gains 3.40, passes systems many times its size, and reasons within touching distance of expert humans.

The wall it runs into next is not cognitive. It is perceptual: 48.45 against 91.24. The next gain in audio reasoning will not come from thinking harder about what the model heard. It will come from hearing it better.

ICLR 2026 — read the paper →Qwen2.5-Omni-7BGRPOMMAU · MMAU-Pro · MMSUUIUC & Amazon AGI Foundations