Qwen3-1.7B, real persisted outputs — nothing simulated. Pick an entity and a question. We average the model's own states into three ingredients — a generic base μ, an entity part, and a question part — add them up, write the sum into a single position inside the model while it reads a different question… and it speaks.
Every output shown is a persisted k=8 greedy generation from the real intervention (op_audit.py / op_patch_decomp.py); coordinates are the measured workspace states projected to 2D (export_showcase_data.py). Paper: “Operator–Operand Factorization in LLM Residual Streams: Causal Influence and Compositional Sufficiency”.