Blog — Jiajun Fan
Writing
Blog
Notes on reinforcement learning, generative models, and research observations — written to think through ideas, not to summarize papers.
What's happening in my field →2026
Voice Agents Could Be Ranked, Never Improved. We Closed the Loop.
Voice agents could be measured, never improved. We closed the loop - two open models talking in native audio - and task success more than doubled.
Read →
For the First Time, We Can See an Audio LLM’s Brain Map and Consciousness Space
A speech model answers a spoken question in one word. We read its middle layers while it listened - the answer was already there, in words it never says.
Read →
2025
Reasoning Made Audio LLMs Worse. We Flipped the Sign.
Telling an Audio LLM to think cost it 3.40 points. Rewarding the reasoning process instead earns 3.40 - the same margin, mirrored.
Read →
One Subtraction Ends the Exploration–Exploitation Trade‑off
RL post-training runs on one fixed coefficient that must both protect the model and get out of its way. Subtract the advantage - one line.
Read →