← blog field notes — mechanistic interpretability

J-space

A small, verbalizable subset of everything a model is computing — and the one fixed matrix that lets you read it out mid-layer, before the model ever finishes its forward pass.

layer ℓ J_ℓ final coords "kindness" "perhaps" "river"
h_ℓ, uncorrectedtransportedread by unembedding
01 — the concept

Not everything the model computes gets said

A transformer holds far more in its residual stream than it ever puts into words. J-space is Anthropic's name for the sliver that's different: representations the model is poised to verbalize, whether or not it happens to in that exact context.

Small

A few dozen concepts active at a time — under a tenth of the model's total internal activity. Most computation never surfaces here.

Privileged

Concepts injected into J-space tend to get verbalized. Injections outside it mostly don't. It behaves like a global workspace.

The question J-space answers isn't "what will the model say next" — it's "what is this activation disposed to make the model say, if the occasion arises."

02 — two lenses

Why you can't just read the residual stream

Every layer rotates, rescales, and reshuffles its activations. A concept living along one axis at layer ℓ may live along an entirely different axis by the final layer. Decoding early activations with the model's own vocabulary decoder — built for final-layer coordinates — reads the wrong basis.

Logit lens

Applies the unembedding matrix directly to h_ℓ. Assumes every layer shares the final layer's coordinates. Wrong assumption — output is noise in early-to-mid layers.

Jacobian lens (J-lens)

First transports h_ℓ into final-layer coordinates using J_ℓ, then unembeds. Corrects for the coordinate drift the logit lens ignores.

03 — building the lens

Where J_ℓ actually comes from

J_ℓ is a fixed matrix, built once, offline — before it's ever used to read anything. It's assembled from real gradients, computed the ordinary way, then averaged until the noise washes out.

Forward pass, cache everything

Run a prompt through the model. Save the residual-stream vector at every layer — h_ℓ, h_ℓ+1, ..., h_final. Nothing unusual: a normal forward pass with the intermediates kept instead of discarded.

Backward pass, one layer at a time

Walk backward from h_final to h_ℓ. At each layer, evaluate that layer's local derivative — how sensitive its output is to its input — using the specific cached numbers from the forward pass. This is why the cache exists.

Chain the local derivatives

Multiply the local derivatives together, layer by layer, walking backward. The product is one concrete matrix: the sensitivity of h_final to h_ℓ, for this one example.

Repeat over thousands of examples

Different prompts push different activations positive or negative, flipping different ReLUs on and off — each example's gradient differs slightly. None is more "correct" than another.

Average

Average the gradient matrices across the corpus. The result, J_ℓ, is no longer tied to any one example — it's a stable, reusable transport matrix. This is also what separates verbalizable representations from ones that only mattered in one odd sentence.

h_ℓ [1, -0.5] h_ℓ+1 [2, 0.5] h_final [2.5, 1] local deriv. → W1 [[2,0],[1,1]] local deriv. → W2 [[1,1],[0,2]] J_ℓ (example 1) W2·W1 = [[3,1],[2,2]] J_ℓ (example 2) [[1,1],[2,2]] … average → J_ℓ = [[2,1],[2,2]]

One example (teal → coral → brass) plus a faded second example, averaged into the final, stable J_ℓ.

04 — worked by hand

The arithmetic, in full

Two tiny 2×2 layers, small enough to check by hand. This is the exact computation behind the diagram above.

h_ℓ = [1, -0.5]
W1 = [[2,0],[1,1]]
h_pre = W1·h_ℓ = [2, 0.5] — both entries positive, ReLU passes through unchanged
h_ℓ+1 = [2, 0.5]
W2 = [[1,1],[0,2]]
h_final = W2·h_ℓ+1 = [2.5, 1]
local deriv. ℓ→ℓ+1 = diag(1,1)·W1 = [[2,0],[1,1]]
local deriv. ℓ+1→final = W2 = [[1,1],[0,2]]
J_ℓ (this example) = W2 · [[2,0],[1,1]] = [[3,1],[2,2]]
— a second example, h_ℓ = [-1, 2], flips the first ReLU off —
J_ℓ (example 2) = [[1,1],[2,2]]
J_ℓ (averaged) = [[2,1],[2,2]]
new h_ℓ = [0.5, 0.5]
transported = J_ℓ·h_ℓ = [1.5, 2.0] → hand to the unembedding matrix
05 — what it is not

Four things easy to conflate it with

The mechanics borrow pieces from training, from attribution patching, from ordinary linear algebra — which makes it easy to import the wrong mental model.

Not gradient descent

Nothing is minimized and nothing gets updated. Each example contributes one measurement; you average measurements, you don't step toward a minimum.

Not ablation

No activation is zeroed or patched. J_ℓ transports the real, untouched vector into a new coordinate system — a simulation, not an intervention.

Not next-token prediction

The real forward pass, run in full, still produces the actual output. J-lens is a diagnostic read-out of an intermediate activation, applied after the fact.

Not an exact change of basis

A textbook basis change is exact and invertible. J_ℓ is a learned, averaged, first-order approximation — the best linear stand-in for a nonlinear computation, nothing more.

06 — the payoff

Reading intent against what gets said

This is the piece that connects back to faithfulness work: if a concept sits stably in J-space at a given reasoning step, it was available to be reported. Whether the model's written chain of thought actually mentions it is a separate, checkable question — a stated reason can now be held up against what the model was actually in a position to say.

A chain of thought is a claim about what happened inside. J-space is a way to check the claim.

discussion