For the First Time, We Can See an Audio LLM’s Brain Map and Consciousness Space
An audio language model takes a spoken question and answers it in one word. Between the sound going in and the word coming out sit forty-eight transformer layers we normally treat as a black box — and unlike a reasoning model, it writes nothing down on the way. This note is about what we found when we read those layers directly, while the model was still listening.
This is the readout itself. Every column is one audio position in time, every row one of the 49 readout depths — input at the bottom, output at the top. A cell is coloured where the model is poised to say the concept there, darker for a better rank, and hovering one shows the six words it was actually closest to producing. Two things read at a glance. The signal is sparse in time: most positions stay blank and a handful of moments carry the concept, which is why the readout takes the best rank over positions rather than averaging. And the bottom of the map is empty — on three of these seven clips not a single cell below 35% depth is lit. Nothing the sound adds is sayable until the workspace band.
The short version
Five things the readout shows
Each one is stated with the number that carries it, and shown further down. If you read nothing else, read these.
An archaeology clip reconstructs a whole unspoken scene of a dig — burial, coffin, ruins, excavation, treasures — none of them in the audio, the question or the options. On some clips the model’s own transcription and its own caption both miss the decisive concept and the readout assembles it regardless. It is not re-transcribing.
Clip and prompt are entirely in English, yet one concept surfaces in several scripts at once — music / música / 音乐 / Musik / musica. Across 863,770 workspace readouts, 52.6% are English and 38.5% are Chinese. It holds a concept, not a word.
Given the same clip as audio and as the model’s own emotion-free caption, only the audio mind forms the true sound source, the speaker’s role, or the affect in a voice. Correct speaker role on 88.9% of disagreement clips versus 70.4% for the caption mind.
The audio-driven gap over the text prior is null at the input, turns on at about 12% depth, and is largest in the workspace band. Transplant real-audio activations into a run whose audio was deleted and the answer comes back — 10/10 patching early, 9/10 mid-stack, 0/10 if you wait for the last fifth.
Deleting the entry layer costs 65 points of accuracy. Deleting any single interior layer costs at most 10, and the band means are within a point of zero — the network routes around whichever one you remove.
01 — The hidden chain
It works out an answer the clip never says
The method is deliberately boring. Take the model’s own output head — the matrix it uses to turn a final hidden state into a word — and apply it early, at every layer, at the positions where the audio lives. That is the logit lens: no probe is trained, nothing is fine-tuned, and it costs one forward pass. What comes back is which word the model is poised to produce at that depth, long before any answer token exists.
What comes back is not a transcript. The readout carries concepts that appear in neither the question, the options, nor the model’s own written description of the clip. On an archaeology recording the workspace band reconstructs an entire unspoken scene of a dig before resolving to the archaeologist the question asks about. On a clip about a telephone patent it assembles telecommunications and wired on the way to the year.
The strongest cases go further: on some clips the model’s own two routes for turning audio into text — its verbatim transcription and its free-form caption — both miss the decisive concept, and the middle layers assemble it anyway. Two independent audio-to-text routes missing what the readout finds makes plain re-transcription an unlikely explanation.
The pattern is not unique to one clip. Pick any of these seven and read the chain it assembles, step by step, in the order each word becomes readable.
Each step is a word the readout reaches at the audio positions, ordered by the depth at which it becomes readable. Concepts in the middle of a chain appear in neither the question nor any option — they are inferred from the sound and then used. Where the chain ends in a name, that name commits at the output slot rather than over the audio; we mark that rather than overclaim it.
Is it reading the answer, or just surfacing frequent words?
A lens can light up on generically common tokens, so the same readout was run for ten unrelated placebo words — banana, guitar, planet, umbrella, tuesday and others. Across ten clips the true answer concept fills 1–34% of workspace cells; the placebos fill 0.12%. Essentially never. The single visible exception is trumpet on the whip clip at 1.3%, where a whip-crack genuinely does resemble a brass transient.
The signal is also sparse in time. Unlike text, where readable content spreads over many tokens, only 4–30% of audio positions carry the concept — 5 of 34 positions on one clip, 10 of 37 on another. A randomly chosen audio position often reads as noise. That is why the readout aggregates by taking the best rank over positions rather than averaging, which would wash the signal out.
02 — A concept, not a word
The clip is in English. The thought is not.
Is the model quietly re-transcribing the sound, or holding an idea? On a music clip the notion of music appears in the workspace band as English music, Chinese yīnyuè, Spanish música, German Musik and Italian musica — all at once, at the same audio positions. On birdsong the leading readout is the Chinese niǎolèi. The model is not reaching for an English string it memorised; it is holding an idea and rendering it into whichever token happens to sit nearest.
Tallying the top-1 readout across all audio-region workspace cells — 863,770 cells over 948 clips — the inner vocabulary is 52.6% English, 38.5% Chinese and at most 4.3% anything else. More than a third of the model’s leading guesses are in a language that appears nowhere in its input.
If the model had memorised an English string, the readout would be English. Instead one concept surfaces in several scripts at the same audio positions, most of them at rank 1. It is holding an idea and rendering it into whichever token sits nearest.
The inputs are entirely in English. This is the language of the top-1 readout at the audio positions across the whole corpus — 948 clips, 863,770 workspace cells, lens-noise fragments dropped. More than a third of what the model is poised to say is Chinese. The paper reports the two leading figures, 52.6% English and 38.5% Chinese, and bounds everything else at 4.3%; the per-language split below that ceiling comes from the project’s own tally of the same cells.
That is not the lens’s known fondness for frequent CJK tokens. A frequency control shows the Chinese form is clip-specific: yīnyuè is rank 1 on 30% of music clips but only 6% of the 610 non-music clips — a 5× lift — and Chinese bird forms show a 9–20× lift on birdsong, where a globally frequent token would lift near 1.
03 — Spreading activation
It free-associates, and always in the same order
Beyond the task-relevant concept, semantic neighbours of what the model hears light up spontaneously — spoken nowhere in the clip, in no question and no option — and are then set aside. On a skateboard clip a single audio position reads down the layers as a schedule: the raw sound is verbalized (gǔndòng, “rolling”), the category forms (skate / skateboard), and only then a neighbour ignites — huáxuě, skiing, a sibling board-sport — held for several layers, never rank 1, then dropped. The model answers Skateboard.
Across every family we could verify, the associate always follows its evoker, and the farthest one (老虎, tiger) lands latest at about 92% depth. The richest cases are whole neighbourhoods: an archaeology clip reconstructs a dig scene — burial, coffin, ruins, excavation, treasures — none of them spoken in the clip. And this is largely not a geometry artifact of the output embedding: five of six associates fall outside their evoker’s top-100 cosine neighbours, and the one near pair (skiing, the 7th neighbour of skate) still ignites strictly after its evoker and is then dropped — a schedule that static similarity does not encode. It is the audio analogue of spreading activation, read layer by layer, with depth playing the role of time.
Every one of these is spoken nowhere in the clip — and in neither the question nor any option. Box size scales with how often the word fires in this clip; the second figure is how many of the 948 corpus grids contain it at all. All of them are rare, so this is genuine association rather than a globally common token. The archaeology clip reconstructs an entire unspoken scene of a dig.
Four phases: verbalize, categorize, ignite, dismiss. The rolling sound is put into words first. The category arrives at L32. Only then does 滑雪 — skiing, a sibling board-sport named nowhere in the clip — ignite beside it, ride along for several layers without ever reaching rank 1, and drop out of the readout entirely by L41. The model answers Skateboard.
The associate never comes first. Across every family we could verify, the neighbour becomes readable at or after its evoker — and the farthest, evidence-free one (老虎, tiger) lands latest at about 92% depth. Two clocks are visible: a perceptual competitor rides the evidence and arrives early but only on the frames that support it — howl is readable from ~65% and fires on the howl frames, not the growl frames — while a memory associate needs no acoustic support and arrives late. Both are set aside; the model answers Lion.
04 — What words throw away
Two minds, one clip
Sound carries how something was said — an emotion, a speaker’s role, an acoustic source — that no transcript preserves. So feed the model the same clip twice: once as the waveform, once as the model’s own emotion-free caption of it, with the question and options identical in both runs.
On the clips where the two minds disagree, only the audio mind forms the true, speech-specific concept. A roar reads lion and the model answers Lion; the caption run never forms the concept and answers Wolf. A clip reads priest where the caption hears only “father”. A sarcastic voice reads the inferred affect frustration, a word named nowhere in the prompt and nowhere in the caption.
Two minds, one clip, side by side. Every column is a token position across the whole prompt — the system preamble, then the audio, then the question and the printed options — and every row a sampled depth. The pale layer shows where the readout is a real word at all; the marked cells are where it is the concept lion. Given the waveform the concept lights up across the audio span and through the workspace band, and the model answers Lion. Given only its own written description of the same clip — which says “deep growls” and “howl” but never names the animal — the concept barely forms at all, and the model answers Wolf. What a sound is lives in the model’s mind only while it can hear it.
And the boundary, stated rather than hidden
Hearing does not always help. On 500 spoken TriviaQA questions the answer is a stored fact and the spoken clip carries words identical to the written question — the sound adds nothing the text does not already supply. There the text mind wins: reading both minds on the 87 items it gets right from text and wrong from speech, the answer concept appears in the text mind’s workspace on 83% of them versus 62% for the speech mind. We treat that task as the boundary of the effect, not as a clean mechanism.
05 — Where the thought lives
Null at the input, largest in the middle, already decided by the end
This is the control the whole paper rests on. The question and the four printed options are byte-identical in all three runs — only the sound is swapped. If the readout were just a prior over the printed options it would not move. It moves all the way to chance. Accuracies are the workspace band of Table 2; ranks are the threshold-free median over the 152,064-token vocabulary.
The whole map on one axis. Nothing the sound adds is answer-relevant through the first eighth of the network. The gap over the text prior then turns significant and never closes; the answer concept stabilizes about a third of the way up and holds; and by the motor layers the decision is already made — patching there recovers nothing. Depths are measured, not assumed.
The findings above read the workspace band by assumption — layers 17–38, taken a priori from the text-model literature, not tuned on this data. So the next question is whether the audio-driven signal actually lives there. Sweeping the readout band by band, the real−silence gap — the part of the answer the printed text alone cannot supply — is negligible in the sensory band and opens up sharply in the workspace.
Threshold-free, the same ordering shows up in the raw rank of the correct answer among all 152,064 vocabulary tokens: #2,297 with the real clip, #5,040 with a mismatched one, #26,241 with silence. And reading that gap layer by layer, the two conditions are statistically indistinguishable through the first several layers — the encoder has deposited acoustic features, but nothing the sound adds is answer-relevant yet — and then, a little past a tenth of the way up, the gap opens at about 12% depth and stays open. Thirty-seven of 49 readout depths reach significance, 35 surviving multiple-comparison control, concentrated in the workspace band (18 of 22 layers, versus 9 of 17 sensory).
Readable is not the same as used
Decodability need not mean the model relies on it. So: take clips the model gets right with audio and wrong once the audio’s features are zeroed, then copy the clean run’s activations into the corrupted run at one band and let it finish.
So the audio content is causally used, and committed before the last fifth of the network. And it is not a quirk of one checkpoint: repeating the band-localized audio-swap on an architecturally different model — Qwen2.5-Omni-7B, a dense 28-layer Thinker against our sparse mixture-of-experts — reproduces the same signature, with the gap small in the sensory band and large in the workspace (p < 10−3).
06 — The functional map
Delete one layer at a time and the pipeline separates
Reading is correlational; deleting is causal. Bypassing one layer at a time — replacing its output with an identity skip — and measuring accuracy over 40 clips gives a clean functional map.
Only the two ends are irreplaceable. Deleting the layer where the audio is read in costs 65 points of accuracy — so far off the scale that it is drawn separately above. Every one of the other 47 layers costs at most 10 points, several cost nothing, and 20 of them slightly help. The band means excluding that entry layer are all within a point of zero: sensory -0.2, workspace +0.7, motor +1.0. Retrieval is not carried by any one layer — the network routes around whichever one you remove.
Breaking accuracy is one thing; which ability breaks is another. Tracking three separable functions per deleted layer on one clip: perception fails only for layers 0–1, delivery fails only for layer 47, and no single layer stops the answer concept from surfacing at all. The pipeline is therefore read the sound in (L0–L1), hold and recall it across the interior, deliver it out (L47) — and only the two ends are irreplaceable.
Which tells you where a hallucination is born
The same readout localizes errors. Take the clips the model answers correctly from real audio but wrongly from its own caption — the caption dropped the decisive acoustic detail, so the text-only mind falls back on a plausible neighbour. Across 73 such clips, the depth at which the wrong concept first takes over is early: median onset 12% depth, 58.9% formed by 20% depth, 46 of 73 inside the sensory band, and 84.9% commit before the output layer rather than slipping at the end.
A text-mode hallucination is thus typically not a late slip. The model locks onto a plausible-but-wrong concept at roughly the same depth where correct answers first emerge, and then never revises it. Reading the two minds side by side both detects the hallucination — the audio mind holds the right concept where the text mind holds the wrong one — and shows where it entered.
07 — Scope
What we do not claim
The account is deliberately conservative, and the limits are worth stating plainly rather than burying. The logit lens is a cheap proxy, not a faithful one — against a faithful lens our readouts should, if anything, under-detect, and early-layer readouts are noisy by construction, which is why the claims concern the middle band. This is one model family, on a sound-dominated slice of clips, with a signal that is sparse across audio positions. The patching test localizes coarsely: it shows the content is committed before the motor band, not that the workspace band alone is responsible.
And “workspace” is used in the functional, information-posted-for-report sense from cognitive science. No claim is made about subjective experience. The quantities reported here are controls — they exist to rule out the model reading a prior off the printed options — not benchmark scores.
Why it matters
This is the regime a chain-of-thought monitor cannot see
Monitoring a model’s stated reasoning only works if the model states it — and a speech agent answering a spoken question directly writes nothing down. Worse, stated reasoning need not be faithful. What this readout shows is that in an audio LLM the decision-relevant concepts are already legible, in words, before a single token is emitted: multilingual, paralinguistic, free-associating, and causally used.
Two uses follow if that holds under a faithful lens. As a training signal, a workspace readout offers a process-level target where a scalar reward sees only the final answer. As a monitor, it is close to free — one matrix product per position read — against deployed multimodal systems that are already latency- and memory-bound. The natural next test is whether safety-relevant decisions — tool calls, refusals, fabrication — are readable this way before a speech agent acts.