“The eye wink at the hand; yet let that be / Which the eye fears, when it is done, to see.” ~ Macbeth
Over the past year, I got really interested in a problem that I originally encountered in protein LLMs. Massive activation and outlier features: certain feature coordinates in transformers are outliers with massive values, irrespective of the token. In the original incarnation of the problem, these massive values let to rather ill-behaved latent space for protein design, where the putative latent space was created by some aggregation statistics over the layer-wise encodings of the protein.
But this problem exists in traditional LLM transformers too. My amazing collabroators Mausum, Vishal, Vraj and I have put out a recent preprint attacking this problem from a mechanistic stance.
The puzzle
A Transformer keeps a residual stream: one vector per token, of size . Each layer reads this vector and adds an update to it.
In trained language models, a few coordinates of this vector take huge values. In the models we studied, the largest value is between hundreds to thousands of times the median coordinate size. Much more extreme values have been reported in the literature! These are called massive activations (MAs). They have practical costs. They make low-bit quantization hard. They shape how KV caches can be compressed.
The odd thing is how long they last. An MA usually appears in an early layer and stays massive until the end of the network. Every later attention head and FFN block could, in principle, learn a linear update that brings the value back to a normal range. None of them does. Why?
Earlier explanations focus on where MAs come from: attention sinks on the BOS token, low-cost storage on delimiter tokens, and FFN blocks that act as directional amplifiers (Sun et al., 2026). These explain how an MA is created. They do not explain why later layers do not correct them.
Reading and writing
Each block in a Pre-LN Transformer does two things with the residual stream. With the layer input and tildes for RMS-normalized inputs:
Some weight matrices read from the residual stream:
- attention:
,
,
- FFN (SwiGLU):
,
Other matrices write back to it:
- attention:
- FFN:
For a linear map , there are two ways a coordinate can escape correction.
- Read-blindness. The coordinate lies in the right null space of
. Changing its value does not change the output. The block cannot see it.
- Write-blindness. The coordinate lies in the left null space of
. No input produces an output along it. The block cannot change it.
To correct an MA, a block must be able to do both: see the value, and write an update that pushes against it. If either side is blocked, the correction cannot happen.
How we measure blindness
First we fix the set of massive coordinates. Let be coordinate
of the residual stream after layer
, at token
, for sequence
. A coordinate is massive if, somewhere in the data, it is much larger than the typical coordinate:
The results are robust to the choice of .
Next, for a read matrix , we take its Gram matrix
. It has the same null space as
and is positive semi-definite. The diagonal entry
says how strongly
responds to coordinate
. We turn this into a score in
:
If every coordinate were treated the same, each score would be . A score near
means
barely responds to coordinate
. A score near
means
responds strongly.
To compare MAs with the other coordinates, we use a one-sided Mann–Whitney test and report the rank-biserial effect . A value near
means almost every massive coordinate has a higher blindness score than almost every other coordinate.
We look at each block in two ways:
- Weight view. Use the trained weights alone.
- Operator view. Use second moments of real inputs and real residual updates on held-out text. This shows what the blocks do in practice.
Result 1: The read side is blind, the write side is open
We tested four models: Llama at 135M, 1.28B, and 2.56B parameters, and Qwen3 1.7B.
On the read side, ,
, and the FFN gate/up projections separate MAs from the other coordinates almost perfectly. The effect is
for
and the FFN, and
for
, in all four models.
shows a weaker blindness score (
from 0.11 to 0.31). A second measure, which checks how much each coordinate sits in the approximate null space of
, gives
from 0.86 to 1.00.
On the write side, and
have scores near the uniform value of
. MAs are inside their write range. The operator view agrees: in every model, both attention and the FFN put more energy into massive coordinates than into the other coordinates.
The MAs are also present at the block inputs after normalization. So the blindness comes from the learned read projections. Normalization does not remove the signal.
We propose the following asymmetry: later blocks do not read the MA value, so they cannot condition a correction on it, however, they still write to the MA coordinates, so the value can keep amplifying!
We postulate that token-independent feature values may be necessary to model language, but thier amplification is a unfortunate side effect of how the FFN and Attention operator works.
Result 2: The model adjusts if you close read blindness of an operator
Is any single read matrix necessary? We trained three variants of the Llama 1.28B model with the same hyperparameters:
| variant | change | perplexity | # MA coords | max MA |
|---|---|---|---|---|
| baseline | none | 14.21 | 10 | |
| 14.67 | 7 | |||
| 14.24 | 10 | |||
| FFN Frozen | 13.67 | 11 | 681 |
An orthogonal has no null space, so it cannot be blind to anything. Both
interventions work: the
blindness score drops to 0.50 with no difference between MAs and other coordinates. MAs remain.
and the FFN gate/up projections still separate MAs from the rest with
.
Freezing the FFN read projections gives the same pattern in the other direction. The FFN scores return to 0.50. Attention stays blind ( for
,
for
). The largest MA shrinks to 681, but there are still 11 massive coordinates. Part of this drop may come from the smaller number of trainable parameters.
In all three variants, the write side stays open.
So no single read matrix carries the asymmetry. When one is forced open, the others take over. This result surprised us!
Result 3: what the FFN amplifier adds
Sun et al. (2026) describe the FFN write to coordinate as a quadratic form in the normalized input
:
Under their approximation, . Here
is the top eigenvector of
and
is its eigenvalue. The direction
need not line up with the coordinate axis
. So the FFN can be blind to the value stored in
and still write a large update to
.
We asked what sets massive coordinates apart in this picture (Llama 1.28B):
| diagnostic | median on | |
|---|---|---|
| 7.74 | +0.78 | |
| 2.19 | +0.98 | |
| 0.39 | +1.00 | |
| within-set | 0.53 vs 0.46 | — |
The amplifier directions of different massive coordinates are only a little more aligned than those of other coordinates (0.53 versus 0.46). They do not share one direction. What sets them apart is gain, and above all how much of the quadratic form sits in its top direction. That last ratio separates the two groups completely in this sample.
Why does the gain concentrate? We measured how concentrated each write row is, using the inverse participation ratio:
Among massive coordinates, rows with higher IPR have larger . Our reading: the FFN read-blindness is partial (FFN scores are weaker than attention scores, 0.63 versus 0.75 at 1.28B). A few FFN channels can still read MA-related directions. When the write row for
uses mainly those channels,
has one strong mode, and the FFN amplifies.
As I mentioned before, unfortunate side-effect!
Result 4: blindness comes first
We tracked 21 training checkpoints, from step 0 to step 20k, for the 1.28B model. We fixed the massive and control coordinates at the final checkpoint and traced them backward. For each signal, we recorded the first checkpoint where the two groups separate with .
- FFN read-blindness separates at 2.3k steps.
- FFN amplifier gain separates at 4.5k steps.
Caveat: The time ordering does not prove that blindness causes the amplifier, but provides circumstantial evidence that one dominates the other. The two parts play different roles. The amplifier explains how some coordinates receive very large FFN writes. Read-blindness explains why later blocks do not correct the values already there.
Result 5: training keeps it this way
Is the asymmetry an accident, or does training hold it in place? We used two local measures at the trained 1.28B checkpoint.
Curvature. For parameter slices tied to each residual coordinate, we sampled random unit directions and measured the Hessian curvature
. High curvature means moving those weights raises the loss quickly (“stiff”). Low curvature means the loss barely changes (“sloppy”), following Transtrum et al.(2015).
We observe that read-side weights are stiffer on MAs. Write-side weights are sloppier. The loss holds the read-blindness in place more firmly than it holds the writes.
The next optimizer step. We applied one synthetic AdamW step at the trained checkpoint and computed its first-order effect on the size of each MA. The step increases MA size at every layer, with a larger effect in later layers. On matched non-massive coordinates the effect is near zero. Most of the increase comes from AdamW’s stored moment state. The new batch gradient has a much smaller effect. Weight decay pushes the other way, but the total effect is still an increase.
So at the end of training, the optimizer is still pushing massive activations to grow.
What this means
Massive activations persist because Transformers learn to ignore them when reading and keep adding to them when writing. This holds across scale (135M to 2.56B) and across two model families (Llama and Qwen3). The blindness is spread over several read matrices. Blocking it in one place moves it to another. The loss landscape and the optimizer state both support it.
Our guess about why: read-blindness is a cheap way for a model to keep token-independent features in the residual stream. Language modeling may need such features. The FFN quadratic amplifier and the lack of feedback control in write operations then push some of these features to massive values.
For anyone trying to remove MAs, for example to make quantization easier: changing one projection is not enough. Our interventions show the model routes around it.
More to come! Stay tuned! We are just getting warmed up!


Leave a Reply