Over the last three articles we built one attention layer from scratch: queries, keys and values, the scaling, a hand-run example with three tokens, then multi-head attention and positional encoding to make that layer usable on real sentences. That is the hard part of the Transformer. What remains is assembly.
An actual Transformer is a small number of distinct parts, repeated many times. The original paper stacks identical encoder layers and identical decoder layers 1. Modern models stack dozens, sometimes around a hundred. The question this article answers is why that repetition is possible at all: a stack of six layers and a stack of ninety-six layers have the same shape, so mechanically stacking them is trivial — the difficulty is that deep stacks are hard to train, and two small design choices are what made depth work.
The unit that gets repeated
An encoder layer contains two sublayers:
- Multi-head self-attention — the layer from the previous articles. Every token produces a query, key and value, and gets back a weighted blend of all the values.
- A position-wise feed-forward network — two linear layers with a ReLU in between, applied to each position separately and identically. In the original design the inner layer is 2048 wide while the input and output are 512 wide 1.
A decoder layer has those two plus a third: encoder–decoder attention (often called cross-attention), which runs the same attention computation but takes its queries from the decoder's own stream and its keys and values from the encoder's output 1. The decoder's self-attention is also masked — each position can only look at earlier positions, so predictions for token can only depend on tokens before .
Two structural facts are worth holding onto:
- Only the attention sublayers move information between positions. The feed-forward network sees one token at a time; it has no way to look at its neighbours. Attention is the only place where "the cat" can find out about "mat".
- Every sublayer is wrapped in "Add & Norm":
LayerNorm(x + Sublayer(x))1. Both the addition and the normalization are applied once per sublayer, not once per layer.
Because every sublayer outputs the same width it receives ( in the original), a layer maps an matrix to an matrix, and so does a stack of ninety-six of them. The paper is explicit that this constant width exists to facilitate these residual connections 1 — the + needs both sides to line up.
Why depth is hard in the first place
Training a network means nudging weights to reduce a loss. The direction and size of each nudge comes from backpropagation: the local derivative of every operation on the path from the loss back to that weight, multiplied together. That multiplication is where depth bites. If a stack of layers each contributes a factor smaller than 1, the product shrinks geometrically and the signal reaching the early layers underflows — those layers effectively stop learning. If the factors are larger than 1, the product explodes instead, and training diverges. Either way, the problem gets worse as grows, which is exactly why "just add more layers" was not a viable strategy for years.
Residual connections: a layer writes an edit instead of a replacement
The fix is to change what a layer computes:
The layer no longer produces a replacement for its input; it produces a modification and adds it. The Transformer borrowed this from ResNet, where it was introduced in 2015 and made 100-plus-layer convolutional networks trainable 14.
What it does for gradients is visible directly. Write a stack as , where is what sublayer adds. Then
When you expand that product, one of its terms is the identity alone. In plain words: there is a route from the loss back to layer that passes through no weight matrix and no nonlinearity at all — every step of that route is a plain . The gradient does not have to survive multiplications; it has a way down that survives zero of them. The other terms are still attenuated products, but the stack no longer depends on them.
The same property applies to information moving forward. A feature written into the stream at layer 1 can still be read at layer 20, because nothing replaced it — layers only add. This is the "residual stream" picture: one vector per token, of width , running from the embedding at the bottom to the output projection at the top. Each sublayer reads a projection of it, computes something, and adds the result back. One more consequence of all those additions: the width never changes, so any sublayer can read from any other.
You can watch the gradient effect with a dozen lines. This is a toy, not a Transformer — the sublayer is a linear map plus ReLU — but the pattern is the real one:
1import torch
2
3def gradient_by_depth(n_layers=24, d=256, residual=True, seed=0):
4 torch.manual_seed(seed)
5 weights = [(torch.randn(d, d) * d**-0.5).requires_grad_(True) for _ in range(n_layers)]
6 h = torch.randn(8, 16, d).requires_grad_(True) # batch 8, 16 tokens, d features
7 states = [h]
8 for w in weights:
9 update = 0.5 * torch.relu(states[-1] @ w) # one sublayer's contribution
10 states.append(states[-1] + update if residual else update)
11 states[-1].sum().backward()
12 return [s.grad.norm().item() for s in states] # gradient reaching each depthRun it with residual=True and residual=False. Without the +, the gradient reaching the bottom is orders of magnitude smaller than the one at the top, and the gap widens with every layer you add. With the +, the values stay within a small factor of each other. The exact numbers depend on d, the seed, and that 0.5 — that factor is what makes the sublayer's contribution the same order of magnitude as the state it is added to, which is what a freshly initialized sublayer looks like.
Layer normalization: keeping the numbers at a workable scale
There is a second problem, created by the first. If every layer adds a contribution of roughly the same size, the stream's magnitude drifts as the stack gets deeper — the norm of a sum of roughly independent contributions grows with . Later sublayers would then be reading numbers on a scale that depends on how deep they happen to sit.
Layer normalization fixes this per token. For one token's vector of features, it computes
with learned scale and shift , one per feature. The mean and variance come from the features of that single vector for that single example — no other token, no other item in the batch contributes 2. That matters for sequences: batch normalization estimates statistics across the batch, so it behaves differently at different batch sizes and needs separate bookkeeping per time step, whereas layer normalization does exactly the same computation at training and inference time 2. For a model that has to generate one token at a time, and that sees sentences of wildly different lengths, that is the property you want.
So the division of labour is: the addition carries information and gradients past every layer; the normalization resets the scale of what each sublayer reads, so weights learned at layer 3 are not being asked to work on numbers that only existed at layer 3. The output of a LayerNorm sits at a fixed scale (before and scale and shift it), so a sublayer's input distribution does not depend on how many layers precede it.
Where you put the normalization decides whether you can go deep
The original Transformer applies the norm after the residual addition 1:
This is the layout shown in Figure 1 of the paper as "Add & Norm". It was inherited from the ResNet convention and not separately justified. It works — but only with a carefully tuned recipe. The original model uses a learning-rate warmup of 4000 steps and Adam with ; skip either and the 6-layer base model fails to converge.
Xiong and colleagues explained why in 2020 3. With this placement, the gradient at initialization near the output layer has an expected magnitude that grows with depth, so the first optimizer steps overshoot badly unless the learning rate is held small at the start. Warmup is a workaround for that. Their analysis also shows the fix: move the normalization inside the residual branch,
so the addition acts on the unnormalized stream. Now the identity term in the Jacobian is not divided by the norm's derivative, gradients pass through depth nearly unchanged, and the expected gradient magnitude actually decreases slowly with depth () rather than growing 3. In experiments, this "pre-norm" version trains to comparable quality with no warmup at all 3. A companion study trained 30-layer encoders stably where post-norm versions diverged around 20 layers 4.
Two honest qualifications, because this is a design trade-off rather than a solved problem:
- Deeper is not automatically better. When warmup is tuned, post-norm slightly beats pre-norm at shallow depth (6 layers) 5. The reason is that pre-norm's identity path makes adjacent layers' outputs more similar to each other, so the upper layers end up doing less. The identity path that makes gradients flow is the same path that lets layers pass information through unprocessed.
- It interacts with initialization. A post-norm variant that scales each residual branch by a depth-dependent factor has been trained to depths in the hundreds, so post-norm's instability is an engineering problem rather than a fundamental barrier 6.
In practice, plain pre-norm with RMSNorm (a normalization that drops the mean subtraction) is the default across open large language models; a few designs use a "sandwich" with a norm both before and after each sublayer, and at least one recent family deliberately returns to a reordered post-norm paired with extra stabilizers 6. The placement is still a knob people turn.
What to take away
The encoder and decoder are the same block with a different set of sublayers: self-attention, feed-forward, plus cross-attention in the decoder, each wrapped in add-and-normalize. You can stack that block to any depth shape-wise because every sublayer preserves . What actually makes depth trainable is the combination: the addition gives gradients and information a route that does not have to survive every operation in between, and the normalization keeps the scale each sublayer reads independent of how far down the stack it sits. Putting the norm inside the branch rather than on the residual sum is what lets the identity route run unmediated — and that is precisely the change that removed the original architecture's dependence on a painstaking learning-rate schedule.
Depth buys sequential computation: each layer reads what earlier layers wrote into the shared stream and adds its own contribution, so more layers means more steps of that kind. It is not free, and it is not the same thing as better — the same identity path that enables it also lets layers underuse themselves.