Blog
Writing
Blog
Notes on reinforcement learning, generative models, and research observations — written to think through ideas, not to summarize papers.
What's happening in my field →2025
Test-Time Inverse Scaling in Audio LLMs
Chain-of-thought reasoning helps text LLMs but hurts Audio LLMs. This post explains why — and how process rewards fix it.
Read →
The Exploration-Exploitation Dilemma in RLHF for Generative Models
A deep dive into why fixed regularization in RLHF leads to diversity collapse, and how adaptive sample-level control resolves the exploration-exploitation dilemma.
Read →