A research program for recovering learned computational architecture from weights
Abstract
We design the outer architecture of a transformer, but we do not directly design the computational system that exists after training. We specify attention, MLPs, normalization, residual connections, embeddings, and an objective. Gradient descent then produces a large collection of parameters whose internal organization was not explicitly engineered by us. In that sense, there are two architectures: the architecture humans wrote down, and the architecture learning created inside it.
This research program is about recovering the second one.
The long-term goal is not merely to find interesting neurons, attention heads, sparse features, or prompt-specific circuits. It is to understand the trained model as a whole system: what its natural internal state variables are, what transformations act on them, how information is routed, what conditions govern those transformations, which structures are preserved or destroyed, how local operations compose into larger subsystems, and whether the entire trained checkpoint admits a much smaller mechanistic description than its raw parameter count suggests.
If the trained parameters are denoted by W, we would like to discover a much smaller description Z together with some generative theory G such that
W≈G(Z).
The important point is that Z should not merely be a compressed code for the weights. It should describe how the machine works. A schematic target might look like
W≈G(S,O,R,C,H)+ε,
where S represents learned state spaces, O primitive transformations, R routing and gating structure, C composition rules, H hierarchical organization, and ε whatever remains unexplained.
The central question is whether trained transformers are mechanistically much simpler than their raw parameter counts make them appear, and if so, what the smaller machine is.
Status of the proposal. This article describes a research program, not a report of established empirical results. The broader claims about learned operator systems, state spaces, circuit grammars, and hierarchical compression are hypotheses to be tested. The goal is to make those hypotheses precise enough that they can be systematically searched, falsified, compared against null models, and evaluated for causal and predictive fidelity.
1. Introduction
1.1 Why start from the weights?
The main reason for focusing on weights is not that activations are unimportant. It is that weights and activations answer different kinds of questions.
An activation is a state produced when a particular input interacts with the model. If a network has layers
hl+1=Fl(hl;Wl),
then for a particular input x,
hl=Hl(x;W).
The activation depends on both the prompt and the model. It is one realized trajectory through the system.
The weights are different. They specify the transformation laws that generate all such trajectories. In a simple linear system,
h=Wx,
the matrix W does not contain the particular value of h independently of x. But it does determine which input directions are preserved, annihilated, amplified, merged, rotated, or mapped into particular output subspaces. The same distinction holds in a transformer, except that the transformations are conditional, nonlinear, relational, and composed across many layers.
The activations are executions of the machinery. The weights specify the machinery itself.
For whole-model reverse engineering, this distinction matters because the space of possible prompts is enormous. Behavioral evaluation samples
f(x1),f(x2),…,f(xn),
but no finite prompt set can obviously cover every rare conditional strategy, dormant pathway, unusual internal state, or context-dependent computation a large model might contain. This becomes particularly relevant if we care about rare or strategically hidden behaviors. A model that behaves differently in evaluation-like and deployment-like conditions could systematically expose one branch of its computation while concealing another.
The prompt space may be effectively unbounded, but the trained mechanism is finite. This motivates a different ambition: rather than attempting to enumerate executions, reverse engineer the generator of those executions.
The central scalability hypothesis is therefore
complexity of possible behavior≫complexity of the mechanism generating that behavior.
If this inequality is strong, model-level analysis could scale very differently from prompt-level analysis. An enormous range of inputs might ultimately route through the same finite collection of state spaces, operators, gates, and reusable subsystems.
1.2 Intrinsic structure before human meaning
There is a second motivation for beginning at the level of weights. Interpretability often begins by asking semantic questions: whether a neuron represents a concept, whether a direction corresponds to truthfulness, whether a head implements induction, or whether a sparse feature corresponds to a recognizable category.
Those questions are useful, but they impose a human ontology very early. There is no reason to assume that the model’s natural computational variables correspond cleanly to concepts we already know how to name. A twenty-dimensional state space might be crucial to routing or control without corresponding to any single linguistic concept. A semantic feature may be only one coordinate inside a larger computational subspace. A control variable may matter enormously while having no simple verbal interpretation at all.
A lower-level mechanistic program should therefore permit an intermediate stage where we discover objects such as V17, T42, or G9 without knowing what they mean. We may know that V17 is written by several components, preserved for three layers, transformed by T42, and later annihilated. We may know that G9 gates a family of downstream operations. That is already real knowledge of the machine.
First recover the mathematical organization. Then determine which parts of that organization are actually exercised by valid inputs. Only afterward attach human meaning.
1.3 Relation to current mechanistic interpretability
Mechanistic interpretability has already shown that meaningful internal computation can sometimes be recovered from neural networks. Circuit-level work demonstrated that attention heads, MLP components, and combinations of components can participate in compact and causally important mechanisms. More recent work has focused increasingly on finding better units of analysis than the architecture itself provides.
This shift is motivated by the fact that neurons and heads are not necessarily the natural computational units of the trained system. Individual neurons can be polysemantic, heads can support several functions, and representations can be distributed through superposition. Sparse autoencoders and related methods attempt to decompose activations into more interpretable latent features, while transcoders and attribution methods attempt to expose how those features interact during particular computations.
These approaches represent substantial progress, but they do not yet provide a complete theory of an ordinary trained language model. A decomposition that reconstructs activations is not automatically the decomposition in which the model’s transformations are simplest. Prompt-conditioned circuit tracing still describes particular executions, and scaling from many such execution-specific explanations to a compact global theory remains difficult. Human-readable feature interpretation may itself become a bottleneck if large models contain millions of useful latent variables. Replacement models also introduce approximation error, and the field still lacks a convincing measure of what fraction of an entire model has actually been explained.
The weights-first program is intended as a complementary lower-level approach. Its aim is to recover the computational anatomy of the whole trained system before requiring every internal object to be human interpretable.
1.4 Separating discovery from validation
A useful methodological principle is to separate structural discovery from behavioral validation.
During discovery, the algorithm would ideally receive only the architecture and trained parameters. It would not select directions because they activate on a particular semantic category or choose pathways because they already matter on a known prompt. Once a structural hypothesis has been discovered and frozen, activations, interventions, natural text, and behavioral tests can be used to determine whether the hypothesized mechanism is actually used and whether it predicts causal behavior.
The scientific pipeline becomes
W⟶H⟶P⟶E,
where H is a structural hypothesis discovered from weights, P is a prediction derived from that hypothesis, and E is independent evidence from model execution.
Weights-only discovery should not become an ideological restriction. Reachability, prevalence, semantics, and behavioral significance ultimately require activations. The point is to preserve a clean distinction between what was inferred from the machine itself and what was learned by watching it run.
2. Conceptual framework
2.1 Organizing principle
A transformer is not an arbitrary mathematical object. It is built from a small number of elementary operations. This gives us a principled way to organize the entire research program from first principles.
The architecture provides a finite mathematical alphabet. Training chooses weights that combine those architectural primitives into learned operations. Learned operations act on some internal state organization. Those operations compose into circuits, circuits into modules, and modules into a global computational architecture.
The resulting conceptual hierarchy is:
Architectural primitives — the fixed mathematical operations supplied by the transformer architecture.
Learned operations — reusable effective transformations instantiated by training.
State architecture — the internal spaces on which those transformations act.
Composition rules — the ways learned operations combine across components and depth.
Circuit families — recurring compositional structures built from those operations.
Modules — larger subsystems assembled from related circuits.
Global mechanistic theory — a compact account of how the learned system works as a whole.
This hierarchy is not assumed to describe the model perfectly. Each transition is an empirical question. The model might have clean primitive families but no simple global state decomposition. It might have clear state spaces but no small operator algebra. Different subsystems may require different mathematical descriptions. The purpose of the hierarchy is not to force a particular ontology, but to organize the space of possible discoveries.
2.2 The transformer as a typed mathematical language
A stronger formulation is to treat the architecture as a typed mathematical language rather than merely a list of layer types. Let Σ denote an architectural signature containing the relevant vector spaces and primitive operations: residual, query, key, value, MLP, embedding, token-position, and logit spaces, together with linear maps, addition, Hadamard products, nonlinearities, inner products, softmax, normalization, masking, and positional transformations.
Let L(Σ) denote the set of legal typed expressions generated from those primitives. Not every syntactically different expression represents a genuinely different computation. Before training there are already architectural identities and function-preserving symmetries: parameter permutations, attention-basis changes, normalization invariances, softmax shift invariance, exact activation identities, and other gauge freedoms. Call these architectural relations Earch. The meaningful architectural search space is therefore closer to the quotient
March=L(Σ)/Earch.
Training selects a highly nongeneric point in this space and may introduce additional low-complexity relations EW: approximate idempotence, commutation, closure, shared invariant subspaces, coordinated nonlinear identities, low tensor rank, repeated conjugacies across depth, or other structure not forced by the architecture. Schematically, the learned machine can be viewed as
L(Σ)/(Earch∪EW).
On this view, a central form of reverse engineering is the recovery of EW: the extra relations introduced by optimization that make this particular trained program simpler than a generic member of the architectural family.
2.3 A generated hypothesis library and a bounded complexity frontier
The hypothesis library should be generated as much as possible from the mathematical type of the object being studied, rather than from a hand-written catalog of familiar mechanisms. For vectors, natural low-complexity relations include equality, proportionality, orthogonality, linear dependence, span, and subspace membership. For operators, they include rank, kernels, images, spectra, invariant spaces, minimal polynomials, and factorizations. Families of operators introduce generated algebras, commutants, simultaneous block structure, and common invariant spaces. Bilinear and multilinear objects introduce tensor rank and separability. Typed graphs of spaces and maps introduce legal paths and relations among paths.
This gives the search a principled complexity filtration. Let
R(1)⊆R(2)⊆⋯
be nested classes of admissible relations ordered by description complexity. Complexity may depend on expression length, number of participating objects, polynomial degree, rank, tensor order, latent dimension, number of free coefficients, or composition-word length. For a checkpoint W, define
RW(k)={r∈R(k):r(W)≈0}.
The question is then how much structural and functional compression is obtained as k increases. This yields a mechanistic complexity frontier: after a sufficiently systematic search through R(k), an unexplained residual is known to lie beyond the searched simplicity class without claiming that it is fundamentally irreducible.
Research priority should not be confused with abstraction depth. A high-level hypothesis such as a shared layer basis or repeated conjugacy may be cheap to test and explain millions of parameters, while a lower-level hypothesis may be expensive and narrow. Priority should therefore be governed by expected explanatory leverage: mathematical naturalness, potential generality, compression, testability, coordination cost, and search cost.
2.4 The recursive decompilation cycle and evidence grades
The same research cycle should be applied recursively to every object class that emerges:
Define the objects exactly. Specify their types, domains and codomains, legal compositions, gauge freedoms, and invariants.
Generate low-complexity structure. Derive identities, symmetries, rank constraints, factorizations, decompositions, and short relations natural for that type.
Order the search. Enumerate hypotheses by mathematical complexity and expected explanatory value.
Compile each hypothesis into a native signature. Derive an efficiently testable condition on the weights.
Scan systematically. Search every compatible location rather than selecting examples for semantic interest.
Establish surprise. Compare against matched nulls that preserve lower-order geometry while destroying the target relation.
Compress. Replace recurring low-level collections with a shorter effective description.
Validate independently. Use held-out weights, new seeds, activations, perturbations, or behavioral tests that were not used in discovery.
Promote. Treat a robust compressed structure as a new primitive object.
Residualize and repeat. Analyze what remains unexplained and enrich the mathematical language only when simpler classes fail.
while the outer loop promotes recurring structure into new object types and reruns the appropriate structure theory.
Every claimed simplification should also state its evidence grade. Useful categories are: exact; uniformly approximate with an explicit error bound; conditionally exact or approximate under a specified state or gate condition; distributionally approximate on reachable states; locally approximate through a Jacobian or Taylor description; and interventionally stable under held-out perturbations. This prevents a prompt-specific local approximation from being mistaken for a persistent global mechanism.
3. Mathematical primitives of the transformer
3.1 Architectural primitive set
A simplified transformer block computes
QKV=N(X)WQ,=N(X)WK,=N(X)WV.
followed by
SAYatt=dkQK⊤+M,=softmax(S),=AVWO.
An MLP computes something of the form
YMLP=ϕ(N(X)Win+b)Wout+c,
and residual connections perform updates such as
X′=X+Y.
Modern gated MLPs additionally use multiplicative structure such as
ϕ(Wgx)⊙Wux.
At the architectural level, therefore, the model is built from a small set of fundamental mathematical operations:
linear and affine maps;
additive residual updates;
scalar nonlinear functions;
bilinear interactions;
multiplicative gating in gated architectures;
softmax normalization and competition;
LayerNorm or RMSNorm;
positional and causal structure;
discrete embedding and output interfaces;
parallel composition across heads.
Everything a transformer learns must ultimately be expressed through compositions of these.
The first research domain is therefore to understand the full mathematical consequences of each primitive and the characteristic structures training can instantiate using it.
3.2 Linear and affine structure
Every linear map admits a rank-one decomposition
W=i∑σiuivi⊤.
Each term computes
x↦σiui(vi⊤x),
which can be interpreted as reading a scalar coordinate and writing it into another direction. From these elementary operations come projections, amplification, attenuation, rotations, reflections, shears, low-rank translations, invariant spaces, kernels, images, and rank collapse.
The most promising research program here is not simply to compute singular values independently for every matrix, but to search jointly for recurrence. Do different components repeatedly read from the same subspaces? Do many matrices share output spaces? Can large families of matrices be expressed using a small operator dictionary? Is there a common coordinate system in which many apparently dense transformations become sparse or block structured?
A particularly high-value direction is simultaneous factorization across all compatible weight-derived operators. If hundreds of matrices can be represented as different sparse combinations of a small number of shared bases or primitive transforms, that would provide direct mechanistic compression.
3.3 Residual addition and rewrite operations
The transformer’s computation is fundamentally additive:
x′=x+F(x).
This means the natural object is often not F alone, but the combined update I+F. If
F(x)=−PVx,
then
x′=(I−PV)x,
which deletes the component inside V. If
F(x)=PVx,
then the same state is amplified. More generally,
F(x)=(R−I)PVx
replaces the representation inside V with a transformed version.
Residual addition makes suppression, amplification, cancellation, correction, restoration, overwrite, accumulation, and interference mathematically natural. A major research direction is therefore to construct a census of residual rewrite operators across the model, looking especially for repeated suppressive, corrective, restorative, or amplifying structures. Another is to study multi-component interactions in which one write only becomes simple when considered together with another, rather than analyzing every component independently.
3.4 Scalar nonlinearities and MLP computation
An ordinary MLP neuron computes
dϕ(a⊤x+b).
This is a nonlinear ridge function: one affine coordinate is measured, passed through a scalar nonlinearity, and written along another direction. Groups of neurons can combine into thresholds, bumps, soft switches, conditional affine transformations, gain-control mechanisms, local basis functions, or computations much simpler than the individual neurons suggest.
The most promising research direction here is to abandon the assumption that the neuron is the computational primitive. Search systematically for pairs and larger populations of neurons whose joint function admits a substantially shorter description than the collection of units considered separately. This includes discovering shared read directions, shared write directions, repeated gating surfaces, low-dimensional nonlinear subnetworks, and algebraic identities specific to the activation function.
A second major program is architectural comparison. Different nonlinearities make different algebraic identities and gating geometries available. Comparing GELU, ReLU, SiLU, and gated MLP architectures could reveal how architectural choices constrain the learned primitive vocabulary.
3.5 Multiplicative gating
Gated architectures introduce direct multiplicative interactions of the form
g(x)⊙v(x).
This allows one learned quantity to modulate another, making conditional computation especially natural. Gate and payload may be partly separable: one set of directions determines whether an operation occurs, while another determines what information is transformed when it does.
The most important research direction is to test whether gated MLPs factor into reusable gate families and reusable payload families. If many neurons combine a small set of gate predicates with a small set of payload transformations, a large MLP may admit a much smaller combinatorial description.
3.6 Bilinear interaction and relational computation
Attention introduces a qualitatively different primitive:
xi⊤Mxj.
With
M=k∑σkqkkk⊤,
the interaction becomes
xi⊤Mxj=k∑σk(qk⊤xi)(kk⊤xj).
This allows the model to condition communication on relationships between states at different positions. Similarity, compatibility, asymmetric relations, positional relations, and content-position conjunctions are all natural possibilities.
The central research direction here is to build a global dictionary of relational primitives. Rather than treating every head’s QK matrix independently, search for query and key factors that recur across many heads and layers. Determine whether apparently different attention heads are using the same underlying relation in different contexts.
3.7 Softmax and routing
Softmax turns relational scores into a normalized distribution,
aj=∑keskesj,
which creates competition, selection, weighted averaging, aggregation, and routing.
This makes attention naturally decomposable into two parts. Query-key structure determines routing; value-output structure determines payload. A head may therefore be better described as a combination of a routing predicate and a payload transformation rather than as one indivisible unit.
One of the highest-value attention research programs is to test whether heads factor combinatorially into a small routing vocabulary and a small payload vocabulary. If this is true, hundreds of heads may compress into a much smaller set of reusable primitives.
3.8 Normalization as part of the mechanism
LayerNorm and RMSNorm introduce state-dependent normalization. In simplified form, LayerNorm behaves approximately like
x↦∥Px∥2/d+ϵPx,
where P removes the mean component.
Normalization therefore removes information, induces approximate scale invariance, couples coordinates through a shared norm, and changes downstream effective gains. A component can influence later computation not only through the semantic content it writes, but by changing the normalization regime in which later components operate.
A dedicated research program should characterize the geometry induced by normalization and distinguish direct residual communication from normalization-mediated interaction. This is necessary for any complete structural theory of transformer computation.
3.9 Positional structure and symbolic interfaces
Causal masks and positional encodings constrain which token positions may interact and how relative or absolute position is represented. Embeddings and unembeddings connect discrete symbols to the internal continuous state.
The most useful directions here are to separate positional and content subspaces, test whether relational operations factor into content and position components, and identify which internal spaces are directly connected to input token identity or output logits. These interfaces may provide the boundary conditions for the deeper learned computational system.
4. Learned computational architecture
4.1 Learned operation vocabulary
The architectural primitives above are fixed by design. The next level of the program asks what simple effective operations training repeatedly constructs from them.
Plausible learned operation families include reading, writing, projection, amplification, suppression, rotation, translation between subspaces, gating, matching, selection, aggregation, copying, residual rewriting, correction, and routing.
The important scientific task is not to assume that all of these exist in every model. It is to discover which families actually recur and how much of the checkpoint they explain.
The highest-value research direction at this level is a model-wide primitive census. Instead of reporting isolated examples, search every layer and component for recurring operator types using common metrics and matched null models. The result should be a candidate instruction vocabulary for the network.
A second direction is dictionary learning over effective operators. If many learned transformations can be approximated by
Ti≈k∑cikGk
with a small set of Gk, then a large portion of the model’s apparent complexity may reduce to sparse combinations of reusable primitives.
A useful default is to increase nonlinear complexity only when simpler descriptions fail. A generic effective-operator hierarchy is
F(x)≈A0x+k∑gk(x)Akx+j∑Bj(x,x)+⋯,
where A0 captures unconditional linear structure, gk(x)Akx captures state-dependent operators, and Bj captures low-rank bilinear structure. The purpose of this expansion is not to force every component into this form, but to make the progression from linear to gated to multilinear descriptions explicit and testable.
4.2 Learned state architecture
An operator vocabulary is incomplete unless we know what those operators act on.
The implementation gives the residual stream an ambient vector space, but there is no reason to assume its coordinate axes are the natural state variables learned by the model. The relevant objects may instead be subspaces
V1,V2,…,VK,
which may overlap, nest, persist across depth, or transform into one another.
A useful global model might look like
Wi≈a,b∑BbTba(i)Ba⊤,
where the columns of Ba span candidate state spaces and the matrices Tba(i) describe interactions among them.
The single most important research direction at this level is joint reader-writer factorization across the entire model. Search for state spaces that simultaneously simplify many attention and MLP operators. If the original matrices are dense but the latent interaction maps become sparse, the network begins to look like a communication architecture rather than an opaque collection of matrices.
Once candidate spaces exist, the next major direction is lifecycle analysis. Determine where each space is first written, who reads it, whether it persists, how it transforms, whether it branches or merges, and where it disappears. This could reveal persistent registers, temporary workspaces, shared buses, private channels, or other communication structures without requiring any semantic label.
A further high-upside direction is to test whether the residual state admits deeper factorization, such as approximate tensor-product structure or redundant encodings. If transformations often act independently on distinct latent factors, this could provide an enormous reduction in complexity.
A particularly principled version of joint state discovery treats the residual stream as a representation of the algebra generated by the effective operators. Given T1,…,Tm, form
A=alg(T1,…,Tm)
and study common invariant or reducing subspaces, simultaneous block decompositions, and the commutant
A′={X:XA=AXfor everyA∈A}.
Nontrivial commutant structure can expose latent decompositions without specifying in advance what those decompositions are supposed to mean. Exact block structure would suggest independent sectors; approximate block structure suggests weak coupling; block-triangular structure can reveal directional dependence.
Once candidate spaces Va and Vb exist, the natural objects become restricted typed maps
Tba(i)=PbTiPa:Va→Vb.
Relations that fail globally may hold after restriction. This is important because a single head or MLP may contain several mechanically distinct sub-operations acting on different latent sectors. Communication labels such as “bus,” “private channel,” “register,” “workspace,” or “control signal” should therefore be conclusions drawn from the morphology of these typed maps, not assumptions imposed beforehand.
4.3 Composition rules
Primitive operations and state spaces are still not a complete computation. The next level asks how operations compose.
For compatible transformations,
TkTk−1⋯T1
may simplify in ways that none of the individual factors reveal. Products may collapse in rank, become projections, approximate inverses, or cancel one another.
A major research program should therefore search for algebraically exceptional compositions across the whole network. Particularly valuable targets include
A2A2ABABAB≈A,≈I,≈I,≈BA,≈C.
These correspond to projectors, involutions, inverses, commuting operations, and approximate closure.
The highest-upside version of this program is to search for a small operator algebra generated by reusable transformations
G1,…,GK.
If much of the network can be described using a small set of generators and approximate multiplication rules, the model would have something resembling a learned instruction algebra.
A complementary direction is to build the global read-write compatibility graph and search for paths whose products become unusually simple. This allows circuits to emerge from weight-space structure rather than being chosen because they already matter for some prompt.
A discovered family of spaces and typed maps can also be represented as a quiver: nodes are state spaces and directed edges are legal maps between them. Legal compositions are then paths. This prunes circuit search by type: only paths whose domains and codomains match are meaningful. One can enumerate short paths and search for relations such as
BA≈0,BA≈C,BA≈I,DC≈BA.
Path equivalence, cancellation, inverse sequences, rank collapse, and closure provide a principled route to circuit discovery. At the operator level, the desired compressed description may resemble a presentation by generators and relations,
⟨G1,…,Gr∣R1,…,Rs⟩,
interpreted in whichever algebraic category—associative algebra, semigroup, path algebra, or related typed system—best fits the discovered objects.
4.4 Nonlinear structure in increasing mathematical order
Linear and operator decompositions will eventually fail because transformer computation is state dependent. The response should be to enrich the description language in the least complex mathematically natural way.
For a nonlinear map F, begin with the Jacobian
JF(x)=DFx.
Across many states, test whether the Jacobian family occupies a low-dimensional operator span,
JF(x)≈k=1∑rgk(x)Ak,r≪d2.
If so, an apparently complicated nonlinear subsystem may be modulating a small operator vocabulary with state-dependent coefficients. If first-order structure is insufficient, move to Hessians and higher derivative tensors, asking about multilinear rank, sparsity, separability, and low-degree closure. The controlled progression is
The description language should become richer only when simpler classes fail to explain the residual.
4.5 Circuit families
In this framework, a circuit is not defined primarily as the set of components active on one example. It is a recurring compositional structure in the learned machine.
Possible circuit classes include feedforward transformations, routing systems, memory mechanisms, error-correction systems, and feedback structures. Residual depth makes approximate negative and positive feedback possible even in a feedforward architecture.
The most useful research direction is to discover circuit families rather than individual circuits. Once state spaces and primitive operations are available, cluster structurally equivalent pathways and determine whether many superficially different component chains implement the same abstract computation.
This could transform circuit analysis from an expanding catalog into a compact grammar.
4.6 Modules and larger computational organizations
Circuits may themselves form larger coherent subsystems. At this point it becomes important not to assume a single global ontology.
Several mathematically plausible organizations should compete. Some subsystems may resemble message-passing graphs, where latent state spaces act as nodes and learned transformations as typed edges. Others may resemble control systems, with controlled variables, feedback signals, and regulators. Some may implement iterative dynamical processes across depth. Others may look like finite-state systems, constraint-solving procedures, error-correcting codes, tensor-factorized computations, or interpreter-like structures separating control from payload.
These should be treated as candidate module classes rather than assumptions.
The most promising research program here is unsupervised module discovery from the previously recovered state and operator graph. Search for regions with dense internal interaction, sparse external interaction, repeated internal motifs, or shared dynamical structure. Then compare alternative descriptions using predictive accuracy and description length.
Another high-value direction is repeated computation across depth. Different layers may implement approximately the same operation in changing coordinates:
Fl≈RlFRl−1.
If this occurs broadly, a nominally deep feedforward transformer may contain something much closer to a repeated refinement algorithm.
4.7 Global learned architecture
The next level asks whether all these local and intermediate structures admit a compact whole-model description.
The main research directions are shared coordinate systems, global operator dictionaries, sparse interaction graphs, hierarchical module discovery, and multiscale coarse-graining. The goal is to move from raw parameters to primitives, from primitives to circuits, from circuits to modules, and from modules to global processes.
At this stage different ontologies should explicitly compete. A bus-like description, operator algebra, tensor-factor model, dynamical system, sparse graph, or hierarchical hybrid may explain different amounts of the network. The right description is whichever provides the strongest combination of compression, prediction, and causal fidelity.
This is where minimum description length becomes especially useful. For a mechanistic theory H, one can compare
L(H)+L(W∣H).
The best theory is not the one with the most evocative terminology. It is the one that explains the most structure with the least additional complexity.
5. Whole-model mechanistic compression
5.1 Generative reconstruction and completeness
Eventually the research program should attempt to build an alternative specification of the model.
Let
Z={S,O,R,C,H}.
From this, construct
W=G(Z).
A strong mechanistic theory should do more than compress the observed checkpoint. It should predict structured parts of the weights that were hidden during discovery. Hiding whole heads, neuron groups, operator blocks, or layers and reconstructing them is a particularly strong test because it distinguishes genuine generative structure from retrospective description.
The reconstructed model should also preserve function. Compare losses, logits, KL divergence, intervention responses, and downstream behavior between W and W.
Finally, analyze the residual
ε=W−W.
If it still contains repeated subspaces, spectral outliers, cross-layer correspondences, or functionally important structure, the theory is incomplete. The analysis should continue recursively until the remaining residual is close to an appropriate random ensemble and contributes little unique function.
This gives the project a possible notion of completeness.
5.2 Existing mathematical machinery to reuse
A large part of this program should reuse existing mathematics and compiler/program-synthesis machinery rather than inventing new tools from scratch. The important frameworks have distinct roles:
Automated theory exploration and equality saturation can generate bounded equational theories from transformer primitives and maintain many equivalent descriptions while searching for low-cost forms. Systems such as Ruler, Enumo, egg, TASO, and TENSAT illustrate the relevant machinery.
Library learning and anti-unification can promote recurring simplified structures into new primitives, as in BABBLE- or DreamCoder-like systems. This is the computational analogue of the promotion step in the recursive decompiler.
Matrix algebras, commutants, and simultaneous block decomposition can let common latent subspaces fall out of operator families rather than being specified semantically in advance.
Quiver representations and path algebras provide a natural formalism for typed latent spaces, maps, legal compositions, and relations among paths.
Noncommutative computational algebra can provide normal forms and generator-relation methods once approximate operator or path algebras emerge.
Operads, wiring diagrams, and categorical syntax can encode typed composition rules and prevent invalid compositions by construction.
Invariant theory and weight-space symmetry analysis are needed to quotient function-preserving gauge freedoms before treating coordinate patterns as mechanisms.
Tensor decomposition and multilinear algebra are natural for bilinear attention, multiplicative gating, position-feature factorization, Hessians, and higher derivative tensors.
Minimum description length should act as a referee among competing abstractions by balancing compression against structural and causal fidelity.
System identification methods such as sparse governing-equation discovery and Koopman-style operator analysis provide a complementary top-down route once weight-derived hypotheses have been frozen.
Causal abstraction, abstract interpretation, reachability analysis, and model checking belong late in the pipeline, after a compact subsystem has been discovered, where they can certify or exploit the recovered abstraction.
The most convincing discoveries should be those for which bottom-up weight structure and top-down functional compression converge on the same object. Pure bottom-up identity enumeration may miss large functional simplifications; pure top-down fitting may invent convenient surrogates with weak native correspondence. Agreement between the two routes is much stronger evidence.
6. Validation and scientific controls
Several research programs should run across every level of the hierarchy.
6.1 Gauge and symmetry
One is gauge and symmetry analysis. A claimed structure should either be invariant under function-preserving reparameterizations or be formulated in a canonicalized coordinate system.
6.2 Matched null models
Another is matched null modeling. Large weight spaces contain accidental alignments, so every structural claim should be compared against random ensembles that preserve the relevant lower-order statistics while destroying only the organization under investigation.
6.3 Cross-seed replication
Cross-seed replication is equally important. If the same state spaces, primitive families, or modules emerge independently after appropriate alignment, that is powerful evidence that they reflect something necessary rather than accidental.
6.4 Training dynamics
Training checkpoints add a developmental dimension. Tracking when primitives, state spaces, circuits, and modules emerge could reveal how gradient descent constructs internal mechanisms and whether there are structural phase transitions during learning.
6.5 Scaling and architecture comparison
Scaling and architecture comparison provide another axis. Different model sizes, nonlinearities, normalization schemes, and positional encodings may produce different learned primitive vocabularies. This can distinguish structures forced by the task from structures induced by the architecture.
6.6 Causal validation
Finally, static hypotheses should be tested causally after discovery. The strongest structural theories should quantitatively predict what happens when the corresponding source states or operators are perturbed.
6.7 Held-out structured weight prediction
A genuinely generative mechanistic theory should predict unseen structured parts of the checkpoint. Instead of masking random scalar weights, hide coherent objects such as whole heads, neuron populations, operator blocks, or layers. Reconstructing them from the discovered dictionary, state-space decomposition, or module theory distinguishes explanatory structure from retrospective curve fitting.
6.8 Functional reconstruction
For a compressed reconstruction W or a high-level simulator F, compare not only weight error but logits, loss, KL divergence, downstream behavior, intervention responses, and rare-condition behavior. Functional fidelity is more important than Frobenius reconstruction error because a small residual in parameter norm can still carry large causal importance.
6.9 Residual-directed discovery and three notions of completeness
After fitting one mechanistic component F1, define the residual
R1=F−F1,
then apply the same search to R1:
R1≈F2+R2,
and continue. The residual should determine the next research cycle. This prevents repeatedly rediscovering already-understood motifs while a large unexplained remainder persists.
Three different notions of completeness should be tracked:
Search completeness: within a bounded grammar, have all relations below complexity k been enumerated or approximated sufficiently well?
Structural completeness: how much of the native checkpoint is explained by the current generative description, measured in gauge-invariant quantities or after canonicalization?
Functional completeness: how much causal computation is preserved by the abstraction, including responses to interventions and rare conditions?
The appropriate endpoint is not a claim that every irreducible nonlinearity has been discovered. It is a bounded statement: no structure within a mathematically generated and systematically searched family of descriptions up to complexity k explains the remaining computation to tolerance ϵ. As the searched language becomes richer, the residual receives a progressively stronger certificate of being bespoke relative to the tested simplicity class.
7. Reachability and semantic interpretation
7.1 Reachability and actual usage
A static structural map tells us what computation is available. It does not automatically tell us which internal states are reachable from valid inputs.
If structural analysis discovers a mechanism that activates in a state-space region R, we then face the reachability question
∃xsuch thathl(x;W)∈R?
This is where activations become essential. The difference is that they are now being used to answer a specific question about an already mapped machine rather than serving as the starting point for discovering the ontology itself.
The main research directions here are reachable-set approximation, usage-frequency estimation, trigger discovery, and the study of dormant mechanisms. Rare but structurally coherent pathways deserve particular attention because prompt-based analysis may systematically miss them.
A structurally available pathway should not automatically be treated as a behavioral risk. Reachability analysis should distinguish frequently used mechanisms, rare but coherent branches, and mathematically available operators that valid inputs cannot reach. Conversely, a rare conditional mechanism that is structurally present but absent from ordinary evaluation traces is a legitimate target for targeted state search. The route to formal assurance is therefore decomposition ightarrow abstraction ightarrow reachability or verification, not raw weights ightarrow proof.
If the mechanistic abstraction eventually becomes sufficiently compact, this level may connect to formal methods such as abstract interpretation, symbolic execution, reachability analysis, model checking, and control-theoretic verification.
7.2 Semantic interpretation
Human semantics belongs at the top of the hierarchy rather than the bottom.
Once the model’s state spaces, operators, circuits, and modules have been mechanically identified, we can ask what they mean. Token projections, natural activations, behavioral probes, causal interventions, and sparse representation methods can all help attach semantic interpretations to the recovered machinery.
Sparse autoencoders become especially interesting here because they provide a different decomposition of the model. They ask which sparse latent variables explain activations; the weights-first program asks which state variables and operators make the model’s computation simplest.
These decompositions need not coincide. One computational state space may contain many semantic features. One semantic feature may participate in several mechanically distinct pathways. Some computational variables may function primarily as routing or control states and resist simple semantic labeling.
Understanding the relationship between representational features and computational variables could therefore become a major research program in its own right.
8. Research workflow and success criteria
8.1 Research workflow
The research program can be summarized as ten successive analytical stages:
Architectural primitives. Characterize the mathematical operations made available by the transformer architecture.
Learned operation vocabulary. Identify the effective transformations that training repeatedly constructs from those primitives.
Learned state architecture. Recover the internal state spaces on which those operations act.
Composition rules. Determine how learned operations combine, cancel, commute, invert, or close under composition.
Circuit families. Group recurring compositions into reusable circuit classes.
Modules. Discover larger subsystems formed from interacting circuit families.
Global learned architecture. Find the most compact whole-model organization supported by the recovered states and operators.
Whole-model mechanistic compression. Reconstruct the checkpoint from that mechanistic description and analyze the unexplained residual.
Reachability and usage. Determine which structurally available mechanisms are actually reachable and how often they are used.
Semantic interpretation. Attach human meaning only after the computational organization has been established.
The important thing is that this hierarchy is generated from the mathematics of the transformer rather than from a predetermined list of human concepts.
At each level, the project asks whether the model actually possesses a simpler organization than the level below. The existence of clean primitives, state spaces, circuits, modules, or global algebraic structure is not guaranteed. Each is an empirical hypothesis.
8.2 Operational decompiler loop
The ten analytical stages describe what levels of structure we hope to recover. The operational loop describes how to move from one level to the next:
A recurring structure is not merely recorded; it is compressed into a new higher-level object, inserted into the working library, and used to generate the next round of hypotheses. The decompiler therefore learns its own vocabulary as it proceeds.
8.3 A principled priority program
The first wave should emphasize low-complexity searches with high explanatory leverage, even when they sit at different conceptual depths:
Build a complete bounded atlas of architectural identities, invariances, and gauge equivalences.
Extract gauge-aware native operators such as QK/OV and effective MLP maps before searching for structure.
Run a model-wide census of rank collapse, shared singular spaces, common kernels and images, proportional or antipodal factors, and low-dimensional operator spans.
Search for joint latent-state decompositions using commutants, simultaneous block decomposition, and related methods.
Enumerate short legal typed paths and search for cancellation, closure, inverse paths, and path equivalence.
Test repeated layer structure, shared bases, conjugate copies, recurrent depth templates, and global operator dictionaries.
Prototype recursive library learning that promotes repeated compressed structures into new primitives.
Require structured held-out weight prediction as a central test of generative explanatory power.
Only after simpler searches support them should the program prioritize richer candidate organizations such as tensor-product state factors, sophisticated error-correcting structure, finite-state or interpreter-like systems, constraint-solving modules, or higher-order noncommutative organization.
8.4 Concrete early research cycles
A practical first implementation can be organized into mutually reinforcing cycles:
Cycle A — transformer equational theory: define a bounded typed grammar and automatically enumerate exact identities and rewrites.
Cycle B — trained-versus-null enrichment: scan trained models, initialization checkpoints, training trajectories, and matched nulls for each generated relation.
Cycle C — joint operator decomposition: construct invariant effective operators and search common spans, commutants, and simultaneous block or triangular structure.
Cycle D — subspace-resolved algebra: restrict operators to robust latent spaces and search low-degree relations again inside the discovered sectors.
Cycle E — quiver and path compression: construct the typed latent-map graph, enumerate bounded paths, and compress equivalent composites into a path grammar.
Cycle F — library promotion: promote recurring operator and path patterns into new primitives and test whether this yields real MDL savings and better held-out prediction.
Cycle G — top-down convergence: on a separate validation corpus, estimate activation/Jacobian operator families and test whether they converge on the same operator library recovered from weights.
These cycles turn the program from a hierarchy of ideas into an executable research agenda.
8.5 What success would look like
A successful endpoint would not necessarily be an English label for every neuron. A much stronger outcome would be a compact alternative specification of the trained model.
Imagine that a network containing hundreds of millions of raw parameters could instead be described using a much smaller collection of learned state spaces, a reusable dictionary of primitive transformations, sparse routing rules, a finite collection of circuit families, and a hierarchy of larger modules. Suppose this description predicted substantial portions of held-out weight structure, reconstructed the model’s function with high fidelity, recurred across independent training runs, and left little additional functionally important structure unexplained.
Many internal objects might still have abstract names such as V41, G13, or M7. But we would know where V41 comes from, who reads it, how long it persists, which operator transforms it, under what conditions that operator applies, how M7 composes several lower-level mechanisms, and what happens when the pathway is perturbed.
At that point the model would no longer be an opaque collection of parameters. It would be a mechanically deciphered system whose semantic interpretation could then be layered on top.
A useful mechanism-level scorecard should report, at minimum: compression, coverage, null-model surprise, causal fidelity, cross-seed recurrence, training-time emergence, held-out prediction, and gauge robustness. A mathematically elegant relation affecting two negligible channels should rank below a less picturesque decomposition that explains a large fraction of a block's causal computation.
8.6 Broader scientific program
If the same learned primitives, state organizations, composition laws, or modules recur across architectures, random seeds, scales, and tasks, this research becomes larger than interpretability.
It becomes a study of the mathematics of learned systems.
Which primitive operators does gradient descent repeatedly discover? Which internal state organizations are universal? Which mechanisms arise because of the transformer architecture, and which arise because of language itself? Do different nonlinearities produce different computational vocabularies? Do larger models introduce qualitatively new mechanism classes or mostly recombine existing ones? Do mechanisms form in a consistent developmental order?
These questions point toward an experimental mathematics of trained neural networks.
9. Conclusion
9.1 The central empirical hypothesis: hierarchical compressibility
The project succeeds most dramatically if trained transformers are hierarchically compressible:
many parameters→few local relations→few operators→few state systems→few circuit schemas→few modules→compact global machine.
This is an empirical hypothesis, not an assumption. Compression may stop at any stage. Some levels may admit elegant algebra while others remain distributed and bespoke. Mapping where simplicity persists—and where it fails—would itself be a substantive scientific result.
If the hypothesis is correct, algebraic compressibility should strengthen during training, recur across independent seeds at least at some abstraction level, transfer across prompts, and predict intervention fidelity. Comparative studies should track how operator vocabulary size, latent-state dimensions, module count, redundancy, and compression ratio change with model scale and architecture.
The governing question is therefore: At what scale does the trained transformer's mechanistic description stop getting simpler?
9.2 The deepest formulation
A transformer begins with a small architectural alphabet. Let that alphabet be A. Training turns it into a much richer learned computational language LW.
The aim of this research program is to infer the grammar of LW directly from the finite mathematical object W.
First recover the effective operations training has constructed from the architectural primitives. Then recover the state spaces on which those operations act. Discover the rules by which they compose, the circuits formed by those compositions, the modules formed by those circuits, and the global architecture formed by those modules. Build the shortest mechanistically faithful description that reproduces the trained system. Only then ask what its internal variables and processes mean in human terms.
The ultimate objective can be written as
Z⋆=argZmin[L(Z)+L(W∣Z)]
subject to the requirement that the reconstruction generated from Z remains structurally and functionally faithful to the original model.
In ordinary language, the objective is to find the shortest faithful explanation of the trained machine.
Current mechanistic interpretability is increasingly capable of identifying meaningful representations and tracing sophisticated computations on selected examples. What it does not yet provide is a complete theory of an ordinary trained language model. The possibility motivating this program is that completeness may require attacking the problem from a lower level: reverse engineer the finite mechanism that generates the executions, recover its intrinsic mathematical organization before forcing it into human semantic categories, and determine whether the enormous behavioral complexity of a language model is generated by substantially less mechanistic complexity underneath.
Finding out whether that is true is the research program.
Appendix A. A bounded transformer grammar
A first prototype can make the search space explicit with a typed expression grammar such as
e::=x∣c∣e+e∣e⊙e∣Ae∣ϕ(e)∣softmax(e)∣Norm(e),
augmented with typed tensor contraction, positional maps, masking, and residual composition. Candidate parameter relations can include bounded forms such as
B=A,B=−A,B=cA,B=AP,B=G−1A,
where P ranges over restricted permutations or orthogonal transforms and G over an allowed gauge family.
A finite completeness claim requires explicit bounds: maximum expression size, number of nonlinearities, number of participating channels or components, polynomial degree, operator-word length, tensor order, rank or latent dimension, and allowed constant or parameter families. Exact semantics should be treated first, followed by controlled approximate matching.
Appendix B. Framework map
Framework
Role in the decompilation program
Theory exploration / equality saturation
Generate and simplify bounded equational theories from transformer primitives.
Library learning / anti-unification
Promote recurring simplified structures into reusable learned primitives.