Abstract

Can a transformer circuit be discovered from weights before choosing a task or prompt? We identify populations of GELU MLP neurons in particular Pythia layers whose read and write vectors are jointly near-antipodal. Exact antipodality would cancel GELU's even component and produce an affine low-rank operator; the learned populations approximate this construction, with fidelity depending on direction, magnitude, and bias agreement. The prospectively frozen direction score predicts held-out response fidelity across sixteen Pythia-160M training runs (Spearman ρ=0.959\rho=0.959; group-stratified permutation p=0.00075p=0.00075); a post-hoc magnitude-and-bias exactness diagnostic improves the association to ρ=0.974\rho=0.974. The population's read channel identifies its dominant upstream attention heads from weights: across two models, two corpora, and twenty windows, the selected heads are always causal ranks one and two. Frozen linear operators accurately predict selected-population response variation; we separately report uncentered fidelity and mean-response error. Clamping the selected neurons during head ablation worsens loss in every tested window, consistent with partial compensation, while response injection alone does not improve baseline loss. Training trajectories show negative operator organization before accurate affine response. The mechanism is not universal across architectures or seeds. The result is a prompt-independent bridge from static weight relations to a natural, causally active multi-component computation.

Keywords: mechanistic interpretability · transformer weights · GELU · self-repair

1. Introduction

Mechanistic interpretability usually begins with a behavior and searches for components that cause it. We study the inverse problem: whether trained weights alone can suggest a falsifiable computation, with natural inputs reserved for subsequent validation. This separates discovery from semantic interpretation and can expose computations that do not align with a chosen prompt family.

We find a specific relation inside GELU MLPs. In selected Pythia layers, disjoint neurons have nearly opposite read directions and nearly opposite write directions. GELU has an exact odd-part identity, so an ideally antipodal pair implements an affine rank-one map despite the nonlinear activation. Multiple pairs form a low-dimensional residual operator whose dominant symmetric component is negative. We call its intervention response compensating when it moves the ablated model's output distribution toward the unablated baseline; this is not a claim of a recurrent control loop.

Directional antipodality alone is insufficient: read and write magnitudes and effective biases also matter. Correct pairing reveals an interpretable algebraic parameterization, but once the neuron set is known, a pair-free population operator predicts at least as well. Static channel overlap localizes upstream writers, although singular-value weighting and natural source activity explain part of the same ranking signal. Finally, a path-specific clamp establishes that the selected response improves the ablated model, not that it explains all self-repair.

Our contributions are:

  • A weight-only construction of candidate affine GELU populations and a post-hoc magnitude-and-bias exactness diagnostic.
  • A sixteen-run prospective test in which the frozen direction score predicts unseen causal response fidelity.
  • Exhaustive upstream-head ranking across two models, two corpora, and five windows, including norm, activity, singular-value-aware, and random-subspace baselines.
  • Matched-neuron and exact-neuron re-pairing controls, plus the natural pair-free population operator.
  • Four-condition mediation experiments reporting loss, KL divergence, and logit-space outcomes for every window.
  • Checkpoint trajectories showing that negative feedback-like geometry precedes accurate antipodal linearization.

2. Algebraic mechanism

2.1 Exact and approximate GELU cancellation

We use row-vector residual states. Specifically, xx is 1×d1\times d, rir_i and wiw_i are d×1d\times1 columns, bib_i is scalar, and AA is d×dd\times d; residual updates therefore act as xAxA. Pythia's activation is the exact erf-based GELU 10,

ϕ(z)=zΦ(z),\phi(z)=z\Phi(z),

where Φ\Phi is the standard normal cumulative distribution function. Therefore

ϕ(z)−ϕ(−z)=z.\phi(z)-\phi(-z)=z.

For normalized residual input xx, write the MLP output as

f(x)=∑iϕ(xri+bi)wi⊤,f(x)=\sum_i \phi(xr_i+b_i)w_i^\top,

where rir_i and wiw_i are read and write vectors under the row-vector convention. If

rj=−ri,bj=−bi,wj=−wi,r_j=-r_i,\qquad b_j=-b_i,\qquad w_j=-w_i,

then the pair contributes

ϕ(z)wi⊤+ϕ(−z)(−wi⊤)=zwi⊤,\phi(z)w_i^\top+\phi(-z)(-w_i^\top)=zw_i^\top,

an affine rank-one map. Negative eigenvalues do not by themselves imply erasure: for a symmetric residual update x↦x+xAx\mapsto x+xA, a mode with eigenvalue λ\lambda contracts only when −2<λ<0-2<\lambda<0, with near-erasure around λ=−1\lambda=-1. We therefore measure natural channel variance and causal response instead of inferring function from sign alone.

Real pairs violate the ideal conditions. Define relative read and write sum defects and a bias defect scaled by the pair's natural read scale:

dr=∥ri+rj∥∥ri∥+∥rj∥,dw=∥wi+wj∥∥wi∥+∥wj∥,d_r=\frac{\lVert r_i+r_j\rVert}{\lVert r_i\rVert+\lVert r_j\rVert},\qquad d_w=\frac{\lVert w_i+w_j\rVert}{\lVert w_i\rVert+\lVert w_j\rVert}, db=∣bi+bj∣(∥ri∥2+∥rj∥2)/2.d_b=\frac{|b_i+b_j|}{\sqrt{(\lVert r_i\rVert^2+\lVert r_j\rVert^2)/2}}.

Our joint exactness score is

eij=min⁡{1−dr,  1−dw,  exp⁡(−db)}.e_{ij}=\min\{1-d_r,\;1-d_w,\;\exp(-d_b)\}.

This is a diagnostic rather than a fitted probabilistic model. It rewards the three algebraic conditions required by the GELU identity and uses no activations.

2.2 LayerNorm coordinates

Let a LayerNorm have scale γ\gamma, shift β\beta, and centering projector PP. Ignoring the input-dependent scalar standard deviation only for geometric discovery, the folded read is

r~i=Pdiag⁡(γ)ri,\tilde r_i=P\operatorname{diag}(\gamma)r_i,

and the effective preactivation bias is

b~i=bi+β⊤ri.\tilde b_i=b_i+\beta^\top r_i.

Pair discovery uses r~i\tilde r_i and b~i\tilde b_i. Causal prediction uses the unfurled read weights because the measured perturbation is taken after the actual LayerNorm. Thus AfoldA_{\mathrm{fold}} denotes the residual-coordinate geometric operator, while AnormA_{\mathrm{norm}} denotes the operator acting on observed normalized-input differences. Biases affect whether the population is globally affine but cancel from the first-order prediction of a response difference when the same frozen affine intercept is subtracted.

All reported discovery direction scores and read/bias exactness defects substitute r~i\tilde r_i and b~i\tilde b_i for rir_i and bib_i in the preceding formulas. Write vectors are unaffected by this fold.

2.3 Weight-only discovery and operators

For every neuron pair we first compute a direction-only score

sij=min⁡{−cos⁡(r~i,r~j),  −cos⁡(wi,wj)}.s_{ij}=\min\{-\cos(\tilde r_i,\tilde r_j),\;-\cos(w_i,w_j)\}.

Unless stated otherwise, we greedily select the twenty highest-scoring disjoint pairs, with deterministic score and index tie order. For p=(i,j)p=(i,j), distinguish folded and normalized-coordinate pair reads:

rfold,p=r~i−r~j2,rnorm,p=ri−rj2,wp=wi−wj2.r_{\mathrm{fold},p}=\frac{\tilde r_i-\tilde r_j}{2},\qquad r_{\mathrm{norm},p}=\frac{r_i-r_j}{2},\qquad w_p=\frac{w_i-w_j}{2}.

The pair-derived operators are

Afold=∑prfold,pwp⊤,Anorm=∑prnorm,pwp⊤.A_{\mathrm{fold}}=\sum_p r_{\mathrm{fold},p}w_p^\top,\qquad A_{\mathrm{norm}}=\sum_p r_{\mathrm{norm},p}w_p^\top.

The pair-free baseline over the identical selected neuron set GG is

Aodd,norm=12∑i∈Griwi⊤.A_{\mathrm{odd,norm}}=\frac12\sum_{i\in G}r_iw_i^\top.

For an exact collection of antipodal pairs the normalized-coordinate pair and pair-free operators agree. AnormA_{\mathrm{norm}} tests the proposed pair construction; Aodd,normA_{\mathrm{odd,norm}} tests how much prediction follows from neuron selection without using a correspondence. AfoldA_{\mathrm{fold}} and its pair-free analogue are used only for residual-coordinate geometry. The channel UU is an orthonormal basis, obtained by thin QR, for the span of the folded pair-read differences.

For each proposed pair their difference is exactly

Aodd,p−Apair,p=14(ri+rj)(wi+wj)⊤.A_{\mathrm{odd},p}-A_{\mathrm{pair},p} =\frac14(r_i+r_j)(w_i+w_j)^\top.

Thus the pair-free and pair-derived operators agree under exact antipodality, while their discrepancy is jointly controlled by the read and write sum defects. This identity explains why the pair-free operator can slightly outperform when learned pairs are only approximate.

2.4 Static architectural wiring

For upstream attention head hh, let BhB_h be a thin-QR orthonormal basis for the columns of its slice of the attention output projection. We define

overlap⁡(h,U)=∥Bh⊤U∥F2min⁡{dim⁡(Bh),dim⁡(U)}.\operatorname{overlap}(h,U)=\frac{\lVert B_h^\top U\rVert_F^2}{\min\{\dim(B_h),\dim(U)\}}.

This basis-invariant score asks what fraction of the smaller subspace is shared. It is invariant to orthogonal basis changes within either represented subspace; it is not invariant to arbitrary reparameterizations of the residual stream. We select the two heads with largest overlap and reserve activations for validation.

We compare this score with raw output-projection norm, the singular-value-weighted channel fraction, natural removed-head energy, and overlap with 99 matched-dimensional Haar-random channels. The weighted fraction is the direct projection-energy ratio and does not discard the head's singular values.

3. Experimental design

3.1 Models, runs, and activation implementation

Exploratory screening includes Pythia-31M, Pythia-70M, Pythia-160M, Pythia-410M, GPT-2, and DistilGPT2. The focused studies use EleutherAI/pythia-70m-deduped layer 3 and EleutherAI/pythia-160m-deduped layer 4 11. The sixteen-run study uses the reference Pythia-160M model and fifteen official data-order and initialization variants released with PolyPythias 12. The releases are not independent biological-style replicates or a complete factorial design; grouped analyses respect the four release categories.

Experiments use Transformers 5.15.0. Its Pythia configuration maps gelu to the exact Torch GELU implementation with no tanh approximation. Models are evaluated without dropout.

3.2 Corpus and window protocol

Validation uses WikiText-2 raw test and TinyStories validation. We concatenate each file's nonempty raw text, tokenize it once with the model tokenizer, and take contiguous 512-token slices without cross-window state. Source-ranking offsets are 0, 512, 1024, 2048, and 4096 on each corpus. The seed study uses WikiText tokens 512–1023. Matched-neuron and re-pairing controls use offset 8704. Mediation uses offsets 9216, 9728, 10240, 10752, and 11264. Every exhaustive head in a model-corpus study is evaluated on the same five windows. Discovery of neurons, channels, and source heads remains weight-only.

3.3 Response and contraction metrics

Let CC subtract the tokenwise sample mean from every response dimension. For observed response YY and frozen prediction Y^\widehat Y, centered response-variation fidelity is

Rcentered2=1−∥C(Y−Y^)∥F2∥CY∥F2.R_{\mathrm{centered}}^2 =1-\frac{\lVert C(Y-\widehat Y)\rVert_F^2}{\lVert CY\rVert_F^2}.

We additionally report uncentered explained energy and relative mean-response error:

Runcentered2=1−∥Y−Y^∥F2∥Y∥F2,R_{\mathrm{uncentered}}^2 =1-\frac{\lVert Y-\widehat Y\rVert_F^2}{\lVert Y\rVert_F^2}, Emean=∥Y‾−Y^‾∥2∥Y‾∥2.E_{\mathrm{mean}} =\frac{\lVert \overline Y-\overline{\widehat Y}\rVert_2}{\lVert \overline Y\rVert_2}.

No slope or intercept is fitted. The centered score ignores a constant vector error, whereas the uncentered score and EmeanE_{\mathrm{mean}} expose it. Response energy is ∥CY∥F2\lVert CY\rVert_F^2 divided by the number of token observations. For channel basis UU and pre-update residual XX, channel variance ratio is

q=∥C[(X+Y)U]∥F2∥C[XU]∥F2.q=\frac{\lVert C[(X+Y)U]\rVert_F^2}{\lVert C[XU]\rVert_F^2}.

Values below one indicate contraction on sampled states. Negative symmetric energy is the fraction of squared eigenvalue mass below zero in (A+A⊤)/2(A+A^\top)/2; it describes the operator's symmetric part and is not alone evidence of erasure.

3.4 Interventions

Pythia uses parallel attention and MLP branches. We remove a head at the previous layer's attention output projection by subtracting the exact contribution of its head slice before the projection bias. The target layer's post-attention LayerNorm output, selected-neuron GELU outputs, full MLP output, and attention output are captured. All downstream operations recompute normally.

For mediation, four forward passes share each token window: baseline; source-head ablation with natural selected-neuron response; source-head ablation while the selected GELU activations are clamped to their baseline values; and baseline source heads with only the selected activations replaced by their values from the ablated run. The final condition is response-only injection. Loss is next-token cross entropy. KL is the mean tokenwise KL from baseline predictions. Logit displacement is centered across vocabulary before squared Euclidean energy and cosine are computed.

3.5 Statistical scope and multiplicity

The direction score, top-twenty greedy selection, layer-selection rule, two-head source selection, seed window, and causal outcomes were fixed before the sixteen-run sweep. The chosen layer maximizes the mean top-twenty direction score, with lower layer index breaking a tie. The magnitude-and-bias exactness score was developed after inspecting those outcomes and is therefore exploratory; its grouped permutation p-value does not correct for diagnostic selection. The broader model and layer screen was also exploratory and involved many candidate structures, so its p-values are descriptive. For the seed study we report pooled Spearman correlations, within-category correlations, leave-one-category-out sensitivity, and a 19,999-draw permutation that shuffles outcomes only within run category. Even the grouped permutation assumes within-category exchangeability.

4. Results

4.1 Direction, magnitude, and bias all matter

The focal populations are approximately rather than exactly antipodal. On isotropic mean-zero unit-RMS normalized-coordinate probes, the twenty-pair operator explains 83.6% of Pythia-70M and 84.3% of Pythia-160M population variance. Mean read/write defects are 0.168/0.075 in 70M and 0.253/0.236 in 160M. Effective bias mismatch is much larger in 70M: mean dbd_b is 1.298 versus 0.107.

Within the focal 70M population, the direction score does not order pair-level synthetic affine fidelity (ρ=−0.256\rho=-0.256), whereas joint exactness does (ρ=0.794\rho=0.794). In 160M the corresponding correlations are 0.862 and 0.949. Across the sixteen runs, the prospectively fixed direction score predicts natural centered response fidelity with ρ=0.959\rho=0.959. The exploratory exactness diagnostic gives ρ=0.974\rho=0.974, with descriptive grouped-permutation p=0.00015p=0.00015. These results suggest that direction finds candidates while magnitude and bias determine whether curvature cancellation is accurate; confirmation requires untouched runs.

Table 1. Static defects and synthetic affine fidelity in the focal populations.

ModelRead defectWrite defectBias defectSynthetic R02R_0^2Direction ρ\rhoExactness ρ\rho
Pythia-70M L30.1680.0751.2980.836-0.2560.794
Pythia-160M L40.2530.2360.1070.8430.8620.949

The finding is not specific to exactly twenty pairs. For 5, 10, 20, and 40 pairs, greedy selection exactly matches a mixed-integer optimized disjoint matching within the top-500 direction-score candidate graph in both models. Synthetic pair-operator R02R_0^2 is 0.85, 0.85, 0.84, and 0.79 in 70M and 0.94, 0.91, 0.84, and 0.72 in 160M. Thus 5–20 pairs lie on a broad high-fidelity plateau; expanding to 40 adds weaker candidates and degrades affinity. Pair-free prediction remains slightly higher at every size. This sensitivity audit was performed after the focal result and is diagnostic, not a new confirmatory dataset.

Figure 1: Pythia-160M run study on WikiText-2 tokens 512-1023. Each point is one official release; color denotes reference, data-only, weights-only, or both-varied category. Weight-only pair score orders held-out response fidelity and complete-block channel contraction.

Figure 1. Pythia-160M run study on WikiText-2 tokens 512–1023. Each point is one official release; color denotes reference, data-only, weights-only, or both-varied category. Weight-only pair score orders held-out response fidelity and complete-block channel contraction.

4.2 The run study survives dependence-aware checks

Across all sixteen Pythia-160M releases, direction score predicts centered response fidelity (ρ=0.959\rho=0.959), uncentered fidelity (ρ=0.944\rho=0.944), and complete-block channel ratio (ρ=−0.971\rho=-0.971). It also anticorrelates with relative mean-response error (ρ=−0.762\rho=-0.762). Within the nine both-varied runs, the prospectively specified centered-fidelity and block-ratio correlations are 0.817 and -0.883. Three-run one-factor categories are too small for reliable inference. Group-stratified permutation p-values for the specified outcomes are 0.00075 and 0.00015. Excluding the reference gives 0.961 and -0.968; every leave-one-category-out analysis remains at absolute ρ\rho at least 0.893. The exploratory joint exactness diagnostic gives ρ=0.974\rho=0.974 with centered response and -0.968 with block ratio.

These results are evidence that static geometry predicts functional fidelity, not an estimate of a population-level seed effect. The striking separation between controlled and both-varied release categories is descriptive because model histories share data and initialization factors.

4.3 Weights localize the dominant upstream heads

Static overlap versus natural response energy has median ρ=0.881\rho=0.881 on WikiText and 0.762 on TinyStories for Pythia-70M, and 0.902 and 0.755 for Pythia-160M. The two preselected heads are causal ranks one and two in all twenty windows. Against 99 matched-dimensional random channels per window, empirical two-sided pp is at most 0.05 in 19 of 20 windows and 0.08 in the remaining window.

The baseline audit qualifies uniqueness. Singular-value-weighted channel fraction is nearly as predictive: median ρ\rho ranges from 0.748 to 0.895. Natural removed-head energy ranges from 0.524 to 0.860. Raw projection norm is much weaker and reverses sign in 160M. The robust claim is therefore that the discovered channel identifies a specific high-impact interface; unweighted subspace overlap is not uniquely superior to all plausible wiring metrics.

Table 2. Upstream-head ranking compared with diagnostic baselines.

Model-corpusSubspace ρ\rhoWeighted ρ\rhoSource activity ρ\rhoProjection norm ρ\rhoRandom-channel pp range
70M WikiText0.8810.8570.8330.4520.01–0.02
70M TinyStories0.7620.7860.5240.4520.01–0.08
160M WikiText0.9020.8950.860-0.3570.01–0.03
160M TinyStories0.7550.7480.832-0.3990.02–0.02

Figure 2: Exhaustive Pythia-160M layer-3 head ablations feeding layer 4. Five 512-token windows are shown for WikiText-2; static overlap is computed before activations, and heads L3H0 and L3H8 are ranks one and two in each window.

Figure 2. Exhaustive Pythia-160M layer-3 head ablations feeding layer 4. Five 512-token windows are shown for WikiText-2; static overlap is computed before activations, and heads L3H0 and L3H8 are ranks one and two in each window.

4.4 Pairing explains parameterization, not extra prediction after selection

Matched ordinary-neuron controls remain strong. At WikiText offset 8704, true centered response fidelity is 0.984 versus control maximum 0.691 in 70M, and 0.916 versus 0.641 in 160M; true complete-block ratio is lower than all 99 controls. Randomly re-pairing the exact selected neurons degrades the pair-derived operator to maximum centered fidelity 0.663 and 0.640. This establishes that the true correspondence is special within the pair construction.

However, the natural pair-free operator AoddA_{\mathrm{odd}} is at least as predictive on source-ranking windows. Table 3 separates centered variation, uncentered total energy, and relative mean-response error for both operators. Pythia-70M remains extremely accurate without centering: pair-derived median uncentered fidelity is 0.981 on WikiText and 0.973 on TinyStories. Pythia-160M has a substantial constant-vector mismatch: centered pair-derived fidelity is 0.914–0.917, uncentered fidelity is 0.805–0.839, and relative mean error is 0.772–0.852. The pair-free operator improves 160M uncentered fidelity to about 0.85 but does not solve its mean mismatch. Thus the frozen operator predicts most total response energy as well as response variation, but the strongest 91–93% figures apply only after centering.

Table 3. Centered, uncentered, and mean-response fidelity by frozen operator.

Model-corpusPair centeredPair uncenteredPair mean errorPair-free centeredPair-free uncenteredPair-free mean error
70M WikiText0.9850.9810.2360.9860.9830.202
70M TinyStories0.9750.9730.2010.9760.9740.172
160M WikiText0.9140.8050.8520.9340.8510.762
160M TinyStories0.9170.8390.7720.9330.8500.774

Therefore the selected neuron set contains the predictive linear map without needing pair labels. Pairing is informative as a weight-only discovery signature and a mechanistic explanation for why a nonlinear population becomes affine, but it adds no observed predictive advantage after neuron selection. All source-ranking headline metrics above identify the operator explicitly. Matched-control and training-dynamics scores are centered response-fidelity measurements.

The re-pairing contraction statistic must also be read carefully. Re-pairing changes the measurement subspace constructed from pair differences, not the physical MLP block. It tests whether the proposed pairing identifies a more contracting channel.

4.5 The selected response contributes to compensation

Clamping the selected neurons to their baseline activations during source-head ablation worsens cross entropy in all ten windows. Median loss saved by the natural selected response is 0.129 nats in 70M and 0.084 in 160M. Median squared logit displacement repaired is 19.5% and 9.4%; the 160M distribution includes one negative window. KL from baseline is lower under natural ablation than under the clamp in every window.

Response-only injection does not improve the unablated model. It raises median loss from 4.234 to 4.268 in 70M and from 3.577 to 3.604 in 160M. Its logit displacement is only partially aligned with the response observed under ablation (median cosine 0.665 and 0.277). The supported interpretation is conditional, path-specific compensation: the response helps after its source is removed, but it is not a generally beneficial additive direction.

Table 4. Four-condition mediation summary.

ModelWindows helpedBaseline lossAblated lossClamped lossInjected lossLoss savedLogit repair
Pythia-70M5/54.2344.9455.0734.2680.12919.5%
Pythia-160M5/53.5774.0874.2253.6040.0849.4%

Figure 3: De-novo discovery across Pythia-70M checkpoints at layer 3. The initialization point is shown separately from log-scaled positive training steps; every checkpoint reselects pairs and source heads.

Figure 3. De-novo discovery across Pythia-70M checkpoints at layer 3. The initialization point is shown separately from log-scaled positive training steps; every checkpoint reselects pairs and source heads.

4.6 Negative organization precedes affine fidelity

In both checkpoint trajectories, negative operator energy and partial source organization appear by roughly step 8k while direction scores remain weak and natural response prediction fails. Between 8k and 32k, source heads stabilize and fidelity rises. By 64k, Pythia-70M reaches direction score 0.946 and centered response fidelity 0.982; Pythia-160M reaches 0.800 and 0.801. Both remain mature through step 143k.

This repeated ordering supports a developmental account in which broad negative organization precedes accurate affine parameterization. It does not prove that optimization first learns a functional controller and then deliberately reparameterizes it; the trajectories are observational.

Figure 4: Independent de-novo trajectory for Pythia-160M layer 4. Negative organization appears before high affine-response fidelity.

Figure 4. Independent de-novo trajectory for Pythia-160M layer 4. Negative organization appears before high affine-response fidelity.

4.7 Scope boundaries

The complete conjunction is localized. Pythia-31M has a weaker case. GPT-2 and DistilGPT2 contain negative pair operators without the same net block contraction. Pythia-410M exceeds the alignment-preserving geometry null but has no disjoint pair above 0.8. Several both-varied Pythia-160M runs show weak exactness and poor affine prediction. Pair geometry, an affine population, upstream wiring, channel contraction, and loss compensation are separate empirical stages.

Rushing and Nanda 1 identify sparse MLP erasure and anti-erasure as mechanisms of self-repair and separate LayerNorm scaling. McDougall et al. 2 show that withdrawing copy suppression can produce apparent attention-head self-repair. Our contribution is not the discovery that compensation exists, but a microscopic weight-only signature, its affine operator, and a static route to its upstream source.

Transcoders 3, sparse feature circuits 4, and neuron-basis analyses 5 provide complementary decompositions for MLP computation. We learn no auxiliary dictionary. Gurnee et al. 7 report universal neurons and antipodal pairs across GPT-2 runs; Gerstner and Schuetze 8 classify gated neurons by within-neuron read-write alignment; Nandan et al. 9 trace universal neurons through training. Our construction instead uses simultaneous within-model cross-neuron opposition on both sides and tests the induced multi-neuron map.

The mathematical framework for transformer circuits 6 motivates composing weight-defined read and write maps. Pearce et al. 13 show that bilinear MLPs admit exact weight-space tensor decompositions and recover circuits directly from weights. Our result identifies a restricted population inside an ordinary GELU MLP whose odd component behaves approximately linearly, making a similarly direct analysis possible without changing the trained architecture.

Pythia 11 and PolyPythias 12 make the checkpoint and run-level analyses possible. We use these resources as structured observational panels, while explicitly avoiding an independence claim unsupported by their shared training factors.

6. Limitations and future work

Discovery involved a broad exploratory program before the focused protocols were frozen. The two focal layers should therefore be treated as discovered cases, not an unbiased prevalence estimate. Our optimized matching audit is exact only within the top-500 candidate graph, and pair-count sensitivity has synthetic rather than natural causal outcomes. A direct perturbation sweep along channel and orthogonal directions would better map the local region over which the affine approximation holds.

The random-channel comparison has only 99 draws and is evaluated after channel discovery. Natural source activity is not weight-only and is used only as a diagnostic baseline, but it explains much of head-response ordering. A stronger future model should predict response magnitude by combining source covariance, projection singular values, LayerNorm geometry, and the frozen population operator.

Zero ablation is off-distribution. The clamp is path-specific but can alter downstream normalization and interactions; it does not identify all compensation in the block. Response-only injection confirms context dependence but is not a sufficiency test for a semantic behavior. Windows share weights and contiguous corpus context, so they are repeated measurements rather than independent model replicates.

Finally, we identify a mathematical transformation and its local wiring but not the channel's semantic content. Token-level interpretation, resampling interventions, and tests on larger independently trained model families are needed before making a general claim about language-model computation.

7. Conclusion

Static weight relations can generate a precise causal hypothesis even in an ordinary nonlinear MLP. In particular Pythia layers, a selected neuron population is organized around near-antipodal GELU pairs and acts approximately as a low-rank negative operator. Magnitude and bias agreement, not direction alone, determine how faithfully the algebra works. The same weight-derived channel localizes its dominant upstream heads, predicts a high-dimensional natural response, and contributes conditionally to reducing ablation damage.

Correct pairs provide a special and interpretable parameterization, but predictive power after neuron selection belongs to the population operator. Together, the results connect an unsupervised parameter relation to a frozen vector prediction, an architectural source, a dependence-aware multi-run association, and a path-specific causal effect.

Appendix A. Complete run audit

The table reports all sixteen official Pythia-160M runs. Channel dimension is twenty in every row. Layer is selected using the frozen maximum-mean-top-twenty rule. Pair score is direction-only; the response column is centered variation fidelity and block ratio uses WikiText-2 tokens 512–1023. Head indices are within the immediately previous attention layer. Machine-readable records additionally contain uncentered fidelity and relative mean-response error for every run.

Table 5. Complete audit of the sixteen Pythia-160M releases.

ModelGroupLayerPair scoreCentered responseBlock ratioHeads
pythia-160mreference40.8250.7580.4330,2
pythia-160m-data-seed1data only40.8210.8620.3762,4
pythia-160m-data-seed2data only40.7920.7810.5530,8
pythia-160m-data-seed3data only40.8750.9590.2702,4
pythia-160m-weight-seed1weights only40.8970.9670.2700,8
pythia-160m-weight-seed2weights only40.8960.9600.2532,4
pythia-160m-weight-seed3weights only40.7860.7390.4402,4
pythia-160m-seed1both varied50.5900.0580.8669,10
pythia-160m-seed2both varied60.6100.2560.7681,3
pythia-160m-seed3both varied50.6300.2880.7424,5
pythia-160m-seed4both varied50.602-0.1400.8312,6
pythia-160m-seed5both varied50.5990.2340.7910,9
pythia-160m-seed6both varied60.6480.5710.7741,10
pythia-160m-seed7both varied50.605-0.0110.8563,9
pythia-160m-seed8both varied50.6660.6520.6186,11
pythia-160m-seed9both varied40.6820.4830.6075,8

Appendix B. Complete mediation audit

Every row is one 512-token WikiText-2 window. Base, ablated, clamped, and injected are cross-entropy losses. Saved is clamped minus naturally ablated loss. Repair is one minus natural-over-clamped centered squared logit displacement; it can be negative.

Table 6. Per-window four-condition mediation outcomes.

ModelOffsetBaseAblatedClampedInjectedSavedRepair
70M92164.4804.9455.0734.5400.1290.257
70M97284.2344.5294.6114.2680.0820.289
70M102404.3084.9385.0674.3940.1290.195
70M107524.0835.0275.1504.1420.1230.184
70M112644.1755.3065.4694.2600.1630.134
160M92163.8424.2734.3293.8680.056-0.918
160M97283.6343.9263.9733.6560.0470.040
160M102403.5774.0644.1483.6040.0840.105
160M107523.4664.0874.2253.4930.1390.121
160M112643.3964.2814.4463.4490.1650.094

Complete KL values and logit cosines are recorded in results/revision_audit_v1/revision_audit.json; all ten natural-ablation KL values are lower than their clamped counterparts. The response-only condition has nonzero KL and slightly worse loss than baseline in every window.

Appendix C. Reproducibility map

The repository contains frozen version-one protocols for the seed study, mediation, and matched populations. It also contains every per-window source ranking and a consolidated audit record with operator defects, exactness values, and mediation rows. Pair selection is deterministic given weights and index tie order. Corpus files are excluded from version control, while source metadata and hashes are retained.

An independent claim-verification script recomputes all headline quantities from machine-readable records. The exactness audit uses isotropic mean-zero unit-RMS probes only to test the static affine identity; natural response and mediation use untouched corpus windows. No semantic labels, task scores, or prompt categories enter discovery.

References

  1. C. Rushing and N. Nanda. Explorations of Self-Repair in Language Models. ICML, PMLR 235, 2024. ↩
  2. C. McDougall, A. Conmy, C. Rushing, T. McGrath, and N. Nanda. Copy Suppression: Comprehensively Understanding an Attention Head. arXiv:2310.04625, 2023. ↩
  3. J. Dunefsky, P. Chlenski, and N. Nanda. Transcoders Find Interpretable LLM Feature Circuits. NeurIPS, 2024. ↩
  4. S. Marks, C. Rager, E. J. Michaud, Y. Belinkov, D. Bau, and A. Mueller. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. ICLR, 2025. ↩
  5. A. Arora, Z. Wu, J. Steinhardt, and S. Schwettmann. Language Model Circuits Are Sparse in the Neuron Basis. arXiv:2601.22594, 2026. ↩
  6. N. Elhage et al. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread, 2021. ↩
  7. W. Gurnee et al. Universal Neurons in GPT2 Language Models. Transactions on Machine Learning Research, 2024. arXiv:2401.12181. ↩
  8. S. Gerstner and H. Schuetze. Understanding Gated Neurons in Transformers from Their Input-Output Functionality. arXiv:2505.17936, 2025. ↩
  9. A. Nandan et al. Universal Neurons in GPT-2: Emergence, Persistence, and Functional Impact. arXiv:2508.00903, 2025. ↩
  10. D. Hendrycks and K. Gimpel. Gaussian Error Linear Units. arXiv:1606.08415, 2016. ↩
  11. S. Biderman et al. Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling. ICML, PMLR 202, 2023. ↩↩2
  12. O. van der Wal, P. Lesci, M. Mueller-Eberstein, N. Saphra, H. Schoelkopf, W. Zuidema, and S. Biderman. PolyPythias: Stability and Outliers across Fifty Language Model Pre-Training Runs. ICLR, 2025. ↩↩2
  13. M. T. Pearce, T. Dooms, A. Rigg, J. M. Oramas, and L. Sharkey. Bilinear MLPs Enable Weight-Based Mechanistic Interpretability. arXiv:2410.08417, 2024. ↩