Skip to content

The thought is made of parts

We took a language model's thought — "the capital of Italy" — apart into three averaged ingredients, put them back together, wrote the assembled state into a single spot inside the model… and it said "Rome". Swap the Italy part for France and it says "Paris". Push the same knob 4× too hard and it babbles — yet still knows the answer.

A thought assembled from three averaged parts flies across the model's map of thoughts, lands beside the real measured state — and the console types the model's actual output: "Rome, and the currency of the United"

The numbers

52% ≈ 53%a thought assembled from averaged parts makes the model say the answer at its own accuracy ceiling (8B: 62% vs 68%)
20/20 × 3every ordered relation swap flips the answer margin in Qwen3-1.7B, Qwen3-8B and Gemma-2-9B; permuted-label nulls ≈ 0
80%even at a 4× overdose — where fluent speech collapses into token loops — the right answer still wins a forced choice
3 × 2three models, two domains (geography, animal taxonomy) where it works — and two (arithmetic, logic) where it provably doesn't

The honest null that started this

This project began as a replication of a "readable subspace" claim — which did not replicate under matched controls. The structure reported here is causal organization, validated against nulls that preserve every statistic of the method while destroying its meaning. Every claim sits next to the control that could have killed it in Evidence & controls.


Paper: Operator–Operand Factorization in LLM Residual Streams: Causal Influence and Compositional SufficiencyPDF · abstract & BibTeX · code & reproducibility · what's next

Matias Podeley · independent researcher · mpodeley@gmail.com · MIT license · learn the field in Spanish: Interpretabilidad Mecanicista