Abstract

An attention head binds a routing operator (QK) to a transport operator (OV), but these two sides are usually analyzed separately. We test whether trained models organize this binding at the population level, without task labels or named head types. Within each layer, we compare the head-by-head geometry of complete QK and OV operators and construct counterfactual layers that preserve every learned route and transport while permuting only their assignment. Pythia-70M, 160M, and 410M show positive QK–OV population relations, whereas GPT-2 shows a negative relation. An exact thin-SVD control shows that this model-dependent relation survives removing singular gains but disappears when left-to-right correspondence is scrambled while preserving supports, spectra, and ranks. Static QK and OV geometry separately predicts held-out attention and source-response geometry. Most importantly, the learned route–transport assignment makes normalized population outputs more reinforcing than matched permutations in Pythia, but more cancelling in GPT-2; random initialization is null. Reinforcement replicates in all six controlled Pythia-160M initialization and data-order variants. During Pythia-70M training, direct behavioral QK–OV similarity peaks early and disappears, while assignment-level reinforcement persists. These results identify head orchestration as a learned object distinct from isolated-head taxonomy: training organizes not only the available routes and transports, but which route carries which transport.

Keywords: mechanistic interpretability · attention · operator geometry · unsupervised discovery

1. Introduction

Multi-head attention is usually described as parallel computation followed by concatenation. In residual-stream coordinates, however, each head is more naturally split into two low-rank operators: QK decides where information is read, and OV decides what is moved and where it is written [1]. This decomposition motivates a question between single-head interpretation and whole-layer analysis: is the pairing of routes with transports itself learned population structure?

The question is not answered by showing that Q and K, or V and O, are coupled within a head. Nor is it answered by head diversity alone. A trained layer could contain the same multiset of routing maps and transport maps under many different pairings. Those pairings would preserve each side in isolation but could change whether the resulting head outputs reinforce, cancel, or span different directions.

We develop three complementary tests. First, we compare the relational geometries of complete QK and OV operators. Second, we separate support, singular spectrum, and left-to-right orientation using exact low-rank factorization and matched nulls. Third, we compose every observed attention pattern with every observed transport response, allowing the actual route–transport assignment to be compared with exact within-layer permutations. Discovery uses no prompts, task labels, linguistic categories, or known head classes; held-out activations are used only for validation and the counterfactual composition assay.

Across five pretrained decoder-only models, the result is not a universal geometry but a family of learned regimes. Pythia models show positive static QK–OV geometry and reinforcing assignments. GPT-2 shows negative static and behavioral relations and a cancelling assignment. DistilGPT2 is intermediate. This heterogeneity is scientifically useful: it rules out an architectural identity and suggests that head populations admit different orchestration strategies.

2. Operators and population geometry

2.1 Routing and transport

Use row-vector residual states x∈R1×dx\in\mathbb{R}^{1\times d}. For head hh, define

MQKh=WQh(WKh)⊤,MOVh=WVhWOh.(1)M_{\mathrm{QK}}^h = W_Q^h (W_K^h)^\top, \qquad M_{\mathrm{OV}}^h = W_V^h W_O^h. \tag{1}

Ignoring positional transformations for notation, the pre-softmax score from destination ii to source jj is xiMQKhxj⊤x_iM_{\mathrm{QK}}^h x_j^\top, while the source-side transport is xjMOVhx_jM_{\mathrm{OV}}^h. The implemented experiments use the model's exact positional mechanism and softmax when collecting behavior.

Equation 1 removes non-identifiability of the skinny factorization: replacing factors L,RL,R of LR⊤LR^\top by LA,RA−⊤LA,RA^{-\top} leaves the complete operator unchanged. We fold the attention-input LayerNorm scale into readers and remove residual mean directions following standard interpretability preprocessing. Raw-factor audits in Pythia-70M, GPT-2, and DistilGPT2 recover the same signs and similar magnitudes.

2.2 Relational similarity

Normalize each complete operator by its Frobenius norm and form, separately in each layer,

GijQK=⟨M^QKi,M^QKj⟩F,GijOV=⟨M^OVi,M^OVj⟩F.(2)G_{ij}^{\mathrm{QK}} = \langle \widehat M_{\mathrm{QK}}^i,\widehat M_{\mathrm{QK}}^j\rangle_F, \qquad G_{ij}^{\mathrm{OV}} = \langle \widehat M_{\mathrm{OV}}^i,\widehat M_{\mathrm{OV}}^j\rangle_F. \tag{2}

We compare the upper triangles of their induced distance matrices by Pearson representational-similarity analysis (RSA), averaging equally over layers. The null independently permutes OV head identity inside every layer. Thus it preserves every QK and OV operator and all marginal within-side geometry; only the correspondence between sides changes. We use (b+1)/(B+1)(b+1)/(B+1) Monte Carlo pp values throughout.

We also inspect Q, K, V, and O support projectors. If BitB_i^t is an orthonormal basis for side tt of head ii, its head-by-head support similarity is

Hijt=tr⁡(PitPjt)=∥(Bit)⊤Bjt∥F2.(3)H_{ij}^t = \operatorname{tr}(P_i^tP_j^t) = \left\lVert (B_i^t)^\top B_j^t\right\rVert_F^2. \tag{3}

All six cross-side support geometries are compared under the same head-label permutation logic.

2.3 Separating support, gain, and orientation

For each rank-rr operator, compute the exact thin SVD from its skinny factors, M=UΣV⊤M=U\Sigma V^\top. We evaluate two controlled reconstructions. The unit-spectrum reconstruction UV⊤UV^\top removes relative singular gains while retaining supports and learned left-to-right correspondence. The orientation null independently signed-permutes columns of VV within every head while leaving UU, Σ\Sigma, both support projectors, and rank exactly fixed. We scramble QK and OV separately. This null tests whether the transformation inside fixed read/write supports contributes information beyond choosing the supports themselves.

3. Behavior and counterfactual assignment

3.1 Held-out behavioral geometry

For each head, routing behavior is the flattened causal attention-probability tensor on held-out text. Transport behavior is the centered source response xWVhWOhxW_V^hW_O^h before attention mixing. Head-by-head cosine geometries are constructed for both. We compare static QK with attention, static OV with transport, and attention with transport. The primary Pythia-70M and GPT-2 evaluations use 16 frozen sequences of length 64 and 999 within-layer permutations. The same text windows are decoded and retokenized for GPT-2.

3.2 Exact route–transport reassignment

Let AhA_h be the complete causal attention pattern of route hh and let Tp(X)T_p(X) be transport pp's centered source-token response. We construct every composition

Yh,p=AhTp(X).(4)Y_{h,p}=A_hT_p(X). \tag{4}

The trained model uses assignment p(h)=hp(h)=h. A counterfactual assignment is a permutation of the complete transport identities within a layer. This keeps all routes, transports, activation marginals, layer membership, and head count fixed.

Flatten each chosen output and normalize it to y^h\widehat y_h. We measure

C(p)=∥∑hy^h,p(h)∥22H=1+2H∑h<k⟨y^h,p(h),y^k,p(k)⟩.(5)C(p) = \frac{\left\lVert\sum_h \widehat y_{h,p(h)}\right\rVert_2^2}{H} = 1+\frac{2}{H}\sum_{h<k} \left\langle \widehat y_{h,p(h)},\widehat y_{k,p(k)} \right\rangle . \tag{5}

Above-null CC means the learned assignment creates more same-signed reinforcement; below-null CC means more cancellation. We also report the magnitude-sensitive analogue

Craw(p)=∥∑hyh,p(h)∥22∑h∥yh,p(h)∥22.(6)C_{\mathrm{raw}}(p) = \frac{ \left\lVert\sum_h y_{h,p(h)}\right\rVert_2^2 }{ \sum_h \left\lVert y_{h,p(h)}\right\rVert_2^2 }. \tag{6}

The principal runs use a fixed signed random projection of residual outputs for speed. Exact full-dd spot checks preserve the signs and conclusions. The coherence metric and reassignment pipeline were fixed after base-model exploration and then applied unchanged to the complete six-run controlled Pythia-160M panel. These runs are replications, not preregistered tests.

3.3 In-model transpositions

To test causal substitutability, we swap the complete V and O factors of two heads while leaving their QK routes intact. For every noninitial Pythia-70M layer, we evaluate all (82)=28\binom{8}{2}=28 transpositions on four frozen held-out sequences.

Identity replay matches the unmodified model exactly. The candidate population score is the mean QV and KV distance RSA after reassignment; the outcome is negative language-model loss change. A within-layer rank-shuffle test uses 9,999 draws. We then fit a rank-based partial association controlling for direct cosine similarity of the two swapped QK operators and of the two OV operators. This experiment was a follow-up diagnostic rather than a primary endpoint.

4. Results

4.1 Complete operators have model-dependent population relations

Table 1 shows that corresponding QK and OV geometries are strongly positive throughout the Pythia family and DistilGPT2, but negative in GPT-2. All six Q/K/V/O support relations are positive in each Pythia model; QK and VO are strongest. Hence GPT-2's negative full-operator relation cannot be explained by anti-aligned supports alone.

Table 1. Mean within-layer distance RSA of complete QK and OV operators.

ModelQK–OV RSAPositive layersAssignment test
Pythia-70M0.4165/6upper p=.001p=.001
Pythia-160M0.38012/12upper p=.001p=.001
Pythia-410M0.20721/24upper p=.001p=.001
DistilGPT20.2435/6upper p=.001p=.001
GPT-2-0.1144/12two-sided p=.009p=.009

Six controlled Pythia-160M variants—three changing initialization and three changing data order—are all positive (approximately 0.26–0.37). This makes the Pythia sign robust to the two principal sources of training stochasticity, within the tested family.

4.2 Internal orientation carries the relation

Replacing all singular values by one leaves RSA at 0.484, 0.404, 0.241, 0.296, and -0.068 in Pythia-70M, Pythia-160M, Pythia-410M, DistilGPT2, and GPT-2. In contrast, scrambling only within-support correspondence centers the relation near zero. With 999 draws, independently scrambling either QK or OV gives p=.001p=.001 in 70M, 160M, and DistilGPT2. GPT-2 remains in the opposite tail (OV lower p=.002p=.002; QK lower p=.005p=.005). Pythia-410M used 99 draws and reaches the corresponding floor p=.01p=.01.

Figure 1: Learned routing–transport orchestration.

Figure 1. Learned routing–transport orchestration. (A) Complete QK–OV geometry survives setting singular values to one but collapses under orientation scrambling. (B) Static operators predict their own held-out behavioral geometries. (C) Natural routing–transport similarity peaks early in Pythia-70M and then vanishes despite persistent static structure. (D) Exact within-layer reassignment shows reinforcing Pythia and cancelling GPT-2 regimes. Open circles are six controlled Pythia-160M runs; stars mark p≤.05p\leq .05.

The effect is therefore not reducible to low rank, support overlap, or singular spectra. The orientation linking input and output directions inside those supports is an essential learned variable.

4.3 Static structure predicts natural behavior

Table 2. Held-out distance RSA. Each static operator predicts its own natural behavior, but cross-side behavioral geometry varies by model.

ModelQK → attentionOV → transportattention ↔ transport
Pythia-70M0.4830.9140.010
Pythia-160M0.4060.916-0.065
DistilGPT20.1930.852-0.126
GPT-20.2290.853-0.260

Static-to-behavior relations are positive in every document in the two principal models. GPT-2's negative attention–transport relation is negative in all 12 layers and all 16 documents (p=.001p=.001). Fresh random GPT-2 is null (RSA 0.015, two-sided p=.661p=.661). Thus static operator geometry is behaviorally meaningful on each side, but a static QK–OV sign is not itself a theorem about distribution-weighted use.

Pythia-70M makes this distinction especially clear. Natural cross-side RSA is -0.013 at initialization, rises to 0.412 at step 512 and 0.546 at step 1,000, then falls through 0.320, 0.043, and -0.042 to 0.010 at the endpoint. Static QK–OV geometry remains positive. The model passes through an early phase of behavioral co-organization before the two observable geometries decouple.

4.4 Learned assignment creates reinforcement or cancellation

Table 3. Normalized coherence under actual versus permuted route–transport assignment. Positive differences indicate reinforcement; negative differences indicate cancellation.

Model/checkpointMatched–null CCDirectionTail pp
Pythia-70M, step 1k0.059reinforcement.001
Pythia-70M, final0.089reinforcement.001
Pythia-160M, final0.055reinforcement.002
DistilGPT2, final0.013null.382
GPT-2, final-0.069cancellation.010
Random GPT-20.001null.218

The magnitude-sensitive effect is 0.096 in final Pythia-70M (p=.001p=.001) and -0.033 in GPT-2 (lower p=.045p=.045). Full-residual-dimensional spot checks retain both signs. Both reinforcing Pythia and cancelling GPT-2 assignments reduce effective or stable population rank relative to reassignment, showing that the effect is not generic diversity maximization. Instead, the learned assignment organizes linear dependence with different signs.

The Pythia-160M effect replicates in all six controlled runs. Normalized matched-minus-null differences are 0.070, 0.044, 0.062, 0.067, 0.101, and 0.067. Five attain the 199-draw floor p=.005p=.005 and the sixth has p=.020p=.020. Raw effects range from 0.066 to 0.116, all with p=.005p=.005. These replications use four held-out sequences each and establish sign robustness rather than a precise population effect size.

Assignment reinforcement is null at Pythia-70M initialization, appears by step 512, and remains significant at every later checkpoint. This contrasts with the transient direct behavioral RSA. The exact composition statistic therefore captures a persistent consequence of binding routes to transports that a pairwise similarity proxy can miss.

4.5 Population geometry does not independently predict causal substitutability

We tested whether preserving abstract population geometry predicts damage when intact OV operators are swapped between QK routes. Complete derangements caused large loss increases, but geometry preservation did not rank their damage (ρ=−.076\rho=-.076). Across every pairwise two-head transposition in Pythia-70M, geometry preservation correlated weakly with lower loss damage (ρ=.169\rho=.169, p=.026p=.026). Direct QK similarity (ρ=.402\rho=.402, p=.0001p=.0001) and direct OV similarity (ρ=.318\rho=.318, p=.0002p=.0002) explained that association; the partial population-geometry effect was null (ρ=−.054\rho=-.054, p=.721p=.721).

This negative result bounds the paper's claim. Population orchestration is a real structural and compositional property, but the present evidence does not show that its scalar summary independently predicts causal head substitutability.

5. Discussion

The central object in this work is neither a discrete head type nor an individual circuit. It is the assignment between two learned collections. A layer can retain every available route and transport while changing its population-level computation by rebinding them. The exact reassignment null makes this statement sharper than a correlation between independently measured head properties.

The orientation control also clarifies where reusable weight structure can hide. Support projectors describe which residual directions are accessible, and singular values describe gain, but neither specifies which input direction maps to which output direction. Preserving both while scrambling only this correspondence destroys the observed model-specific relation. Analyses based only on spectra or subspace overlap therefore omit a learned level of operator organization.

The difference between Pythia and GPT-2 should not yet be read as a quality ranking. Reinforcement and cancellation could support different normalization, redundancy, specialization, or robustness strategies. The current result says that these regimes are learned and reproducible, not why one training recipe selects one regime.

The QK/OV circuit decomposition and residual-stream operator view follow the framework of Elhage et al. [1]. Work on multi-head diversity has measured or regularized subspace, attention-pattern, and output disagreement separately [2, 3]. Pythia and PolyPythias enable controlled checkpoint, initialization, and data-order comparisons [4, 5]. Recent work studies attention superposition, feature-resolved attention, and whether QK and OV should be treated as a single computational unit [6, 7]. Projection-kernel methods establish that query and key subspaces themselves have non-random relations [8].

Our contribution is narrower than a general account of attention: a factorization-invariant population assay on pretrained language models, an orientation null preserving supports and spectra, and an exact assignment null preserving all marginal routes and transports. Together they expose learned reinforcing and cancelling regimes and their training dynamics. We make no priority claim beyond this operational combination pending a complete literature review.

7. Limitations and next tests

The models are small decoder-only transformers from two closely related families. Behavioral validation uses a limited frozen text sample. The random output sketch is unbiased for inner products but adds approximation noise; full-dimensional checks are currently spot checks. Monte Carlo tests address the specified reassignment null, not every alternative explanation. The counterfactual composition assay is exact for fixed observed routes and transports but does not recompute routes after an in-model intervention. The central hypotheses were refined during exploratory analysis; the controlled seed panel tests a frozen follow-up pipeline but is not a preregistration.

The decisive next steps are replication in a modern non-GPT architecture, full-dimensional confirmation across all principal runs, and energy-controlled in-model reassignment or replacement. A useful functional extension would test whether reinforcement or cancellation predicts robustness, specialization, or loss under matched perturbations. Semantic interpretation should follow, not precede, these structural tests.

8. Conclusion

Trained attention layers organize more than a bag of head computations. The complete QK and OV collections possess model-dependent relational structure; that structure depends on learned orientation inside low-rank supports, predicts natural behavior on each side, and changes the coherence of composed head outputs under exact reassignment. Controlled seeds and random initialization distinguish the result from chance architecture. Attention-head populations can therefore be studied as orchestrated assignments of routes to transports, providing a new intermediate scale for unsupervised mechanistic analysis.

Appendix A. Reproducibility details

Pythia-70M-deduped has 6 layers, 8 heads, d=512d=512; Pythia-160M-deduped has 12 layers, 12 heads, d=768d=768; Pythia-410M-deduped has 24 layers, 16 heads, d=1024d=1024; GPT-2 small has 12 layers and 12 heads with d=768d=768; DistilGPT2 has 6 layers and 12 heads with d=768d=768. Final Pythia checkpoints use step 143,000. Operator and orientation tests use 999 permutations except the 410M orientation decomposition, which uses 99. Behavioral tests use frozen token arrays stored with the repository results. Every result JSON records model identity, preprocessing, random seed, sample size, and null definition.

Four versioned scripts implement operator geometry, orientation decomposition, behavioral validation, and counterfactual assignment. A separate claim verifier reloads the committed JSON records and fails if manuscript values drift.

Appendix B. Controlled Pythia-160M assignment runs

Table 4. Counterfactual assignment replication. Effects are matched minus within-layer permutation means.

RunNormalized effectppRaw effectpp
Data seed 10.0704.0050.0753.005
Data seed 20.0438.0200.0662.005
Data seed 30.0622.0050.1074.005
Weight seed 10.0669.0050.0849.005
Weight seed 20.1007.0050.1059.005
Weight seed 30.0672.0050.1157.005

References

[1] N. Elhage et al. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread, 2021. Online article.

[2] J. Li et al. Multi-Head Attention with Disagreement Regularization. EMNLP, 2018. Proceedings page.

[3] H. Yun, T. Kang, and K. Jung. Analyzing and Controlling Inter-Head Diversity in Multi-Head Attention. Applied Sciences, 2021.

[4] S. Biderman et al. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. ICML, 2023.

[5] O. van der Wal et al. PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs. ICLR, 2025.

[6] A. Lindsey et al. Progress on Attention. Transformer Circuits, 2025. Online article.

[7] Anonymous. When Is an Attention Head a Computational Unit? OpenReview, 2026. OpenReview page.

[8] H. Yamagiwa, Y. Takase, and H. Shimodaira. Measuring Affinity between Attention-Head Weight Subspaces via the Projection Kernel. arXiv:2601.10266, 2026.