15 min read
Attention Heads Learn Routing–Transport Orchestration
Abstract
An attention head binds a routing operator (QK) to a transport operator (OV), but these two sides are usually analyzed separately. We test whether trained models organize this binding at the population level, without task labels or named head types. Within each layer, we compare the head-by-head geometry of complete QK and OV operators and construct counterfactual layers that preserve every learned route and transport while permuting only their assignment. Pythia-70M, 160M, and 410M show positive QK–OV population relations, whereas GPT-2 shows a negative relation. An exact thin-SVD control shows that this model-dependent relation survives removing singular gains but disappears when left-to-right correspondence is scrambled while preserving supports, spectra, and ranks. Static QK and OV geometry separately predicts held-out attention and source-response geometry. Most importantly, the learned route–transport assignment makes normalized population outputs more reinforcing than matched permutations in Pythia, but more cancelling in GPT-2; random initialization is null. Reinforcement replicates in all six controlled Pythia-160M initialization and data-order variants. During Pythia-70M training, direct behavioral QK–OV similarity peaks early and disappears, while assignment-level reinforcement persists. These results identify head orchestration as a learned object distinct from isolated-head taxonomy: training organizes not only the available routes and transports, but which route carries which transport.
Keywords: mechanistic interpretability · attention · operator geometry · unsupervised discovery
1. Introduction
Multi-head attention is usually described as parallel computation followed by concatenation. In residual-stream coordinates, however, each head is more naturally split into two low-rank operators: QK decides where information is read, and OV decides what is moved and where it is written [1]. This decomposition motivates a question between single-head interpretation and whole-layer analysis: is the pairing of routes with transports itself learned population structure?
The question is not answered by showing that Q and K, or V and O, are coupled within a head. Nor is it answered by head diversity alone. A trained layer could contain the same multiset of routing maps and transport maps under many different pairings. Those pairings would preserve each side in isolation but could change whether the resulting head outputs reinforce, cancel, or span different directions.
We develop three complementary tests. First, we compare the relational geometries of complete QK and OV operators. Second, we separate support, singular spectrum, and left-to-right orientation using exact low-rank factorization and matched nulls. Third, we compose every observed attention pattern with every observed transport response, allowing the actual route–transport assignment to be compared with exact within-layer permutations. Discovery uses no prompts, task labels, linguistic categories, or known head classes; held-out activations are used only for validation and the counterfactual composition assay.
Across five pretrained decoder-only models, the result is not a universal geometry but a family of learned regimes. Pythia models show positive static QK–OV geometry and reinforcing assignments. GPT-2 shows negative static and behavioral relations and a cancelling assignment. DistilGPT2 is intermediate. This heterogeneity is scientifically useful: it rules out an architectural identity and suggests that head populations admit different orchestration strategies.
2. Operators and population geometry
2.1 Routing and transport
Use row-vector residual states . For head , define
Ignoring positional transformations for notation, the pre-softmax score from destination to source is , while the source-side transport is . The implemented experiments use the model's exact positional mechanism and softmax when collecting behavior.
Equation 1 removes non-identifiability of the skinny factorization: replacing factors of by leaves the complete operator unchanged. We fold the attention-input LayerNorm scale into readers and remove residual mean directions following standard interpretability preprocessing. Raw-factor audits in Pythia-70M, GPT-2, and DistilGPT2 recover the same signs and similar magnitudes.
2.2 Relational similarity
Normalize each complete operator by its Frobenius norm and form, separately in each layer,
We compare the upper triangles of their induced distance matrices by Pearson representational-similarity analysis (RSA), averaging equally over layers. The null independently permutes OV head identity inside every layer. Thus it preserves every QK and OV operator and all marginal within-side geometry; only the correspondence between sides changes. We use Monte Carlo values throughout.
We also inspect Q, K, V, and O support projectors. If is an orthonormal basis for side of head , its head-by-head support similarity is
All six cross-side support geometries are compared under the same head-label permutation logic.
2.3 Separating support, gain, and orientation
For each rank- operator, compute the exact thin SVD from its skinny factors, . We evaluate two controlled reconstructions. The unit-spectrum reconstruction removes relative singular gains while retaining supports and learned left-to-right correspondence. The orientation null independently signed-permutes columns of within every head while leaving , , both support projectors, and rank exactly fixed. We scramble QK and OV separately. This null tests whether the transformation inside fixed read/write supports contributes information beyond choosing the supports themselves.
3. Behavior and counterfactual assignment
3.1 Held-out behavioral geometry
For each head, routing behavior is the flattened causal attention-probability tensor on held-out text. Transport behavior is the centered source response before attention mixing. Head-by-head cosine geometries are constructed for both. We compare static QK with attention, static OV with transport, and attention with transport. The primary Pythia-70M and GPT-2 evaluations use 16 frozen sequences of length 64 and 999 within-layer permutations. The same text windows are decoded and retokenized for GPT-2.
3.2 Exact route–transport reassignment
Let be the complete causal attention pattern of route and let be transport 's centered source-token response. We construct every composition
The trained model uses assignment . A counterfactual assignment is a permutation of the complete transport identities within a layer. This keeps all routes, transports, activation marginals, layer membership, and head count fixed.
Flatten each chosen output and normalize it to . We measure
Above-null means the learned assignment creates more same-signed reinforcement; below-null means more cancellation. We also report the magnitude-sensitive analogue
The principal runs use a fixed signed random projection of residual outputs for speed. Exact full- spot checks preserve the signs and conclusions. The coherence metric and reassignment pipeline were fixed after base-model exploration and then applied unchanged to the complete six-run controlled Pythia-160M panel. These runs are replications, not preregistered tests.
3.3 In-model transpositions
To test causal substitutability, we swap the complete V and O factors of two heads while leaving their QK routes intact. For every noninitial Pythia-70M layer, we evaluate all transpositions on four frozen held-out sequences.
Identity replay matches the unmodified model exactly. The candidate population score is the mean QV and KV distance RSA after reassignment; the outcome is negative language-model loss change. A within-layer rank-shuffle test uses 9,999 draws. We then fit a rank-based partial association controlling for direct cosine similarity of the two swapped QK operators and of the two OV operators. This experiment was a follow-up diagnostic rather than a primary endpoint.
4. Results
4.1 Complete operators have model-dependent population relations
Table 1 shows that corresponding QK and OV geometries are strongly positive throughout the Pythia family and DistilGPT2, but negative in GPT-2. All six Q/K/V/O support relations are positive in each Pythia model; QK and VO are strongest. Hence GPT-2's negative full-operator relation cannot be explained by anti-aligned supports alone.
Table 1. Mean within-layer distance RSA of complete QK and OV operators.
| Model | QK–OV RSA | Positive layers | Assignment test |
|---|---|---|---|
| Pythia-70M | 0.416 | 5/6 | upper |
| Pythia-160M | 0.380 | 12/12 | upper |
| Pythia-410M | 0.207 | 21/24 | upper |
| DistilGPT2 | 0.243 | 5/6 | upper |
| GPT-2 | -0.114 | 4/12 | two-sided |
Six controlled Pythia-160M variants—three changing initialization and three changing data order—are all positive (approximately 0.26–0.37). This makes the Pythia sign robust to the two principal sources of training stochasticity, within the tested family.
4.2 Internal orientation carries the relation
Replacing all singular values by one leaves RSA at 0.484, 0.404, 0.241, 0.296, and -0.068 in Pythia-70M, Pythia-160M, Pythia-410M, DistilGPT2, and GPT-2. In contrast, scrambling only within-support correspondence centers the relation near zero. With 999 draws, independently scrambling either QK or OV gives in 70M, 160M, and DistilGPT2. GPT-2 remains in the opposite tail (OV lower ; QK lower ). Pythia-410M used 99 draws and reaches the corresponding floor .

Figure 1. Learned routing–transport orchestration. (A) Complete QK–OV geometry survives setting singular values to one but collapses under orientation scrambling. (B) Static operators predict their own held-out behavioral geometries. (C) Natural routing–transport similarity peaks early in Pythia-70M and then vanishes despite persistent static structure. (D) Exact within-layer reassignment shows reinforcing Pythia and cancelling GPT-2 regimes. Open circles are six controlled Pythia-160M runs; stars mark .
The effect is therefore not reducible to low rank, support overlap, or singular spectra. The orientation linking input and output directions inside those supports is an essential learned variable.
4.3 Static structure predicts natural behavior
Table 2. Held-out distance RSA. Each static operator predicts its own natural behavior, but cross-side behavioral geometry varies by model.
| Model | QK → attention | OV → transport | attention ↔ transport |
|---|---|---|---|
| Pythia-70M | 0.483 | 0.914 | 0.010 |
| Pythia-160M | 0.406 | 0.916 | -0.065 |
| DistilGPT2 | 0.193 | 0.852 | -0.126 |
| GPT-2 | 0.229 | 0.853 | -0.260 |
Static-to-behavior relations are positive in every document in the two principal models. GPT-2's negative attention–transport relation is negative in all 12 layers and all 16 documents (). Fresh random GPT-2 is null (RSA 0.015, two-sided ). Thus static operator geometry is behaviorally meaningful on each side, but a static QK–OV sign is not itself a theorem about distribution-weighted use.
Pythia-70M makes this distinction especially clear. Natural cross-side RSA is -0.013 at initialization, rises to 0.412 at step 512 and 0.546 at step 1,000, then falls through 0.320, 0.043, and -0.042 to 0.010 at the endpoint. Static QK–OV geometry remains positive. The model passes through an early phase of behavioral co-organization before the two observable geometries decouple.
4.4 Learned assignment creates reinforcement or cancellation
Table 3. Normalized coherence under actual versus permuted route–transport assignment. Positive differences indicate reinforcement; negative differences indicate cancellation.
| Model/checkpoint | Matched–null | Direction | Tail |
|---|---|---|---|
| Pythia-70M, step 1k | 0.059 | reinforcement | .001 |
| Pythia-70M, final | 0.089 | reinforcement | .001 |
| Pythia-160M, final | 0.055 | reinforcement | .002 |
| DistilGPT2, final | 0.013 | null | .382 |
| GPT-2, final | -0.069 | cancellation | .010 |
| Random GPT-2 | 0.001 | null | .218 |
The magnitude-sensitive effect is 0.096 in final Pythia-70M () and -0.033 in GPT-2 (lower ). Full-residual-dimensional spot checks retain both signs. Both reinforcing Pythia and cancelling GPT-2 assignments reduce effective or stable population rank relative to reassignment, showing that the effect is not generic diversity maximization. Instead, the learned assignment organizes linear dependence with different signs.
The Pythia-160M effect replicates in all six controlled runs. Normalized matched-minus-null differences are 0.070, 0.044, 0.062, 0.067, 0.101, and 0.067. Five attain the 199-draw floor and the sixth has . Raw effects range from 0.066 to 0.116, all with . These replications use four held-out sequences each and establish sign robustness rather than a precise population effect size.
Assignment reinforcement is null at Pythia-70M initialization, appears by step 512, and remains significant at every later checkpoint. This contrasts with the transient direct behavioral RSA. The exact composition statistic therefore captures a persistent consequence of binding routes to transports that a pairwise similarity proxy can miss.
4.5 Population geometry does not independently predict causal substitutability
We tested whether preserving abstract population geometry predicts damage when intact OV operators are swapped between QK routes. Complete derangements caused large loss increases, but geometry preservation did not rank their damage (). Across every pairwise two-head transposition in Pythia-70M, geometry preservation correlated weakly with lower loss damage (, ). Direct QK similarity (, ) and direct OV similarity (, ) explained that association; the partial population-geometry effect was null (, ).
This negative result bounds the paper's claim. Population orchestration is a real structural and compositional property, but the present evidence does not show that its scalar summary independently predicts causal head substitutability.
5. Discussion
The central object in this work is neither a discrete head type nor an individual circuit. It is the assignment between two learned collections. A layer can retain every available route and transport while changing its population-level computation by rebinding them. The exact reassignment null makes this statement sharper than a correlation between independently measured head properties.
The orientation control also clarifies where reusable weight structure can hide. Support projectors describe which residual directions are accessible, and singular values describe gain, but neither specifies which input direction maps to which output direction. Preserving both while scrambling only this correspondence destroys the observed model-specific relation. Analyses based only on spectra or subspace overlap therefore omit a learned level of operator organization.
The difference between Pythia and GPT-2 should not yet be read as a quality ranking. Reinforcement and cancellation could support different normalization, redundancy, specialization, or robustness strategies. The current result says that these regimes are learned and reproducible, not why one training recipe selects one regime.
6. Related work
The QK/OV circuit decomposition and residual-stream operator view follow the framework of Elhage et al. [1]. Work on multi-head diversity has measured or regularized subspace, attention-pattern, and output disagreement separately [2, 3]. Pythia and PolyPythias enable controlled checkpoint, initialization, and data-order comparisons [4, 5]. Recent work studies attention superposition, feature-resolved attention, and whether QK and OV should be treated as a single computational unit [6, 7]. Projection-kernel methods establish that query and key subspaces themselves have non-random relations [8].
Our contribution is narrower than a general account of attention: a factorization-invariant population assay on pretrained language models, an orientation null preserving supports and spectra, and an exact assignment null preserving all marginal routes and transports. Together they expose learned reinforcing and cancelling regimes and their training dynamics. We make no priority claim beyond this operational combination pending a complete literature review.
7. Limitations and next tests
The models are small decoder-only transformers from two closely related families. Behavioral validation uses a limited frozen text sample. The random output sketch is unbiased for inner products but adds approximation noise; full-dimensional checks are currently spot checks. Monte Carlo tests address the specified reassignment null, not every alternative explanation. The counterfactual composition assay is exact for fixed observed routes and transports but does not recompute routes after an in-model intervention. The central hypotheses were refined during exploratory analysis; the controlled seed panel tests a frozen follow-up pipeline but is not a preregistration.
The decisive next steps are replication in a modern non-GPT architecture, full-dimensional confirmation across all principal runs, and energy-controlled in-model reassignment or replacement. A useful functional extension would test whether reinforcement or cancellation predicts robustness, specialization, or loss under matched perturbations. Semantic interpretation should follow, not precede, these structural tests.
8. Conclusion
Trained attention layers organize more than a bag of head computations. The complete QK and OV collections possess model-dependent relational structure; that structure depends on learned orientation inside low-rank supports, predicts natural behavior on each side, and changes the coherence of composed head outputs under exact reassignment. Controlled seeds and random initialization distinguish the result from chance architecture. Attention-head populations can therefore be studied as orchestrated assignments of routes to transports, providing a new intermediate scale for unsupervised mechanistic analysis.
Appendix A. Reproducibility details
Pythia-70M-deduped has 6 layers, 8 heads, ; Pythia-160M-deduped has 12 layers, 12 heads, ; Pythia-410M-deduped has 24 layers, 16 heads, ; GPT-2 small has 12 layers and 12 heads with ; DistilGPT2 has 6 layers and 12 heads with . Final Pythia checkpoints use step 143,000. Operator and orientation tests use 999 permutations except the 410M orientation decomposition, which uses 99. Behavioral tests use frozen token arrays stored with the repository results. Every result JSON records model identity, preprocessing, random seed, sample size, and null definition.
Four versioned scripts implement operator geometry, orientation decomposition, behavioral validation, and counterfactual assignment. A separate claim verifier reloads the committed JSON records and fails if manuscript values drift.
Appendix B. Controlled Pythia-160M assignment runs
Table 4. Counterfactual assignment replication. Effects are matched minus within-layer permutation means.
| Run | Normalized effect | Raw effect | ||
|---|---|---|---|---|
| Data seed 1 | 0.0704 | .005 | 0.0753 | .005 |
| Data seed 2 | 0.0438 | .020 | 0.0662 | .005 |
| Data seed 3 | 0.0622 | .005 | 0.1074 | .005 |
| Weight seed 1 | 0.0669 | .005 | 0.0849 | .005 |
| Weight seed 2 | 0.1007 | .005 | 0.1059 | .005 |
| Weight seed 3 | 0.0672 | .005 | 0.1157 | .005 |
References
[1] N. Elhage et al. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread, 2021. Online article.
[2] J. Li et al. Multi-Head Attention with Disagreement Regularization. EMNLP, 2018. Proceedings page.
[3] H. Yun, T. Kang, and K. Jung. Analyzing and Controlling Inter-Head Diversity in Multi-Head Attention. Applied Sciences, 2021.
[4] S. Biderman et al. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. ICML, 2023.
[5] O. van der Wal et al. PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs. ICLR, 2025.
[6] A. Lindsey et al. Progress on Attention. Transformer Circuits, 2025. Online article.
[7] Anonymous. When Is an Attention Head a Computational Unit? OpenReview, 2026. OpenReview page.
[8] H. Yamagiwa, Y. Takase, and H. Shimodaira. Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel. arXiv:2601.10266, 2026.