# Encoder and Decoder Stacks: Why Residual Connections and LayerNorm Let Transformers Go Deep

The repeated block behind encoder and decoder, how the residual addition gives gradients an unmediated path through depth, what layer normalization computes per token, and why moving the norm inside the residual branch removed the original warmup requirement

> transformer architecture · layer normalization · residual connections · About 8 min · Oct 6

## Key points

1. An encoder layer is multi-head self-attention plus a position-wise feed-forward network, and a decoder layer adds a third sublayer that attends to the encoder's output; each sublayer is wrapped in `x + Sublayer(x)` followed by layer normalization.
2. The original Transformer stacks N=6 encoder and N=6 decoder layers, with every sublayer and embedding producing width d_model = 512 so that the residual additions line up.
3. Backpropagation multiplies a local derivative for every operation on the path, so in a plain deep stack the gradient reaching early layers shrinks (or explodes) geometrically with depth.
4. With $x_{\ell+1} = x_\ell + \Delta_\ell$ the total derivative contains $I$ at every step of the product, giving gradients a path to any layer that passes through no weight matrix at all; this is the same mechanism that made ResNet-scale depth trainable for convolutional networks.
5. Layer normalization computes mean and variance across the features of a single token vector for a single example, so it behaves identically at training and inference time and does not depend on batch size or sequence length; batch normalization does not have that property.
6. LayerNorm keeps the scale of each sublayer's input independent of depth, which matters because a long chain of additions makes the residual stream's magnitude drift.
7. The original post-norm placement $\mathrm{LN}(x + F(x))$ leaves the gradient near the output growing with depth at initialization, which is why the original recipe needs a 4000-step learning-rate warmup and Adam with $\beta_2=0.98$.
8. Pre-norm placement $x + F(\mathrm{LN}(x))$ keeps the identity route unmediated, so gradients are well behaved at initialization and warmup is no longer necessary; experiments show 30-layer encoders training stably where post-norm diverged around 20 layers.
9. Post-norm still slightly outperforms pre-norm at shallow depth (about 6 layers) when warmup is tuned, because pre-norm's identity path makes upper layers less effective — depth is a knob, not a monotonic improvement, and the placement choice today is often combined with extra stabilizers.

---

Over the last three articles we built one attention layer from scratch: queries, keys and values, the $\sqrt{d_k}$ scaling, a hand-run example with three tokens, then multi-head attention and positional encoding to make that layer usable on real sentences. That is the hard part of the Transformer. What remains is assembly.

An actual Transformer is a small number of distinct parts, repeated many times. The original paper stacks $N=6$ identical encoder layers and $N=6$ identical decoder layers [1]. Modern models stack dozens, sometimes around a hundred. The question this article answers is why that repetition is possible at all: a stack of six layers and a stack of ninety-six layers have the same shape, so mechanically stacking them is trivial — the difficulty is that deep stacks are hard to *train*, and two small design choices are what made depth work.

## The unit that gets repeated

An encoder layer contains two sublayers:

1. **Multi-head self-attention** — the layer from the previous articles. Every token produces a query, key and value, and gets back a weighted blend of all the values.
2. **A position-wise feed-forward network** — two linear layers with a ReLU in between, applied to each position separately and identically. In the original design the inner layer is 2048 wide while the input and output are 512 wide [1].

A decoder layer has those two plus a third: **encoder–decoder attention** (often called cross-attention), which runs the same attention computation but takes its queries from the decoder's own stream and its keys and values from the encoder's output [1]. The decoder's self-attention is also *masked* — each position can only look at earlier positions, so predictions for token $i$ can only depend on tokens before $i$.

Two structural facts are worth holding onto:

- **Only the attention sublayers move information between positions.** The feed-forward network sees one token at a time; it has no way to look at its neighbours. Attention is the only place where "the cat" can find out about "mat".
- **Every sublayer is wrapped in "Add & Norm"**: `LayerNorm(x + Sublayer(x))` [1]. Both the addition and the normalization are applied once per sublayer, not once per layer.

Because every sublayer outputs the same width it receives ($d_{\text{model}} = 512$ in the original), a layer maps an $n \times 512$ matrix to an $n \times 512$ matrix, and so does a stack of ninety-six of them. The paper is explicit that this constant width exists *to facilitate these residual connections* [1] — the `+` needs both sides to line up.

## Why depth is hard in the first place

Training a network means nudging weights to reduce a loss. The direction and size of each nudge comes from backpropagation: the local derivative of every operation on the path from the loss back to that weight, multiplied together. That multiplication is where depth bites. If a stack of $L$ layers each contributes a factor smaller than 1, the product shrinks geometrically and the signal reaching the early layers underflows — those layers effectively stop learning. If the factors are larger than 1, the product explodes instead, and training diverges. Either way, the problem gets worse as $L$ grows, which is exactly why "just add more layers" was not a viable strategy for years.

## Residual connections: a layer writes an edit instead of a replacement

The fix is to change what a layer computes:

$$y = x + \mathrm{Sublayer}(x)$$

The layer no longer produces a replacement for its input; it produces a *modification* and adds it. The Transformer borrowed this from ResNet, where it was introduced in 2015 and made 100-plus-layer convolutional networks trainable [1][4].

What it does for gradients is visible directly. Write a stack as $x_{\ell+1} = x_\ell + \Delta_\ell$, where $\Delta_\ell$ is what sublayer $\ell$ adds. Then

$$\frac{\partial x_L}{\partial x_\ell}=\prod_{k=\ell}^{L-1}\left(I+\frac{\partial \Delta_k}{\partial x_k}\right)$$

When you expand that product, one of its terms is the identity $I$ alone. In plain words: there is a route from the loss back to layer $\ell$ that passes through no weight matrix and no nonlinearity at all — every step of that route is a plain $+$. The gradient does not have to survive $L$ multiplications; it has a way down that survives zero of them. The other terms are still attenuated products, but the stack no longer *depends* on them.

The same property applies to information moving forward. A feature written into the stream at layer 1 can still be read at layer 20, because nothing replaced it — layers only add. This is the "residual stream" picture: one vector per token, of width $d_{\text{model}}$, running from the embedding at the bottom to the output projection at the top. Each sublayer reads a projection of it, computes something, and adds the result back. One more consequence of all those additions: the width never changes, so any sublayer can read from any other.

You can watch the gradient effect with a dozen lines. This is a toy, not a Transformer — the sublayer is a linear map plus ReLU — but the pattern is the real one:

```python
import torch

def gradient_by_depth(n_layers=24, d=256, residual=True, seed=0):
    torch.manual_seed(seed)
    weights = [(torch.randn(d, d) * d**-0.5).requires_grad_(True) for _ in range(n_layers)]
    h = torch.randn(8, 16, d).requires_grad_(True)   # batch 8, 16 tokens, d features
    states = [h]
    for w in weights:
        update = 0.5 * torch.relu(states[-1] @ w)     # one sublayer's contribution
        states.append(states[-1] + update if residual else update)
    states[-1].sum().backward()
    return [s.grad.norm().item() for s in states]     # gradient reaching each depth
```

Run it with `residual=True` and `residual=False`. Without the `+`, the gradient reaching the bottom is orders of magnitude smaller than the one at the top, and the gap widens with every layer you add. With the `+`, the values stay within a small factor of each other. The exact numbers depend on `d`, the seed, and that `0.5` — that factor is what makes the sublayer's contribution the same order of magnitude as the state it is added to, which is what a freshly initialized sublayer looks like.

## Layer normalization: keeping the numbers at a workable scale

There is a second problem, created by the first. If every layer adds a contribution of roughly the same size, the stream's magnitude drifts as the stack gets deeper — the norm of a sum of $L$ roughly independent contributions grows with $L$. Later sublayers would then be reading numbers on a scale that depends on how deep they happen to sit.

Layer normalization fixes this per token. For one token's vector $x$ of $d$ features, it computes

$$\mu=\frac{1}{d}\sum_i x_i,\qquad \sigma^2=\frac{1}{d}\sum_i (x_i-\mu)^2,\qquad \mathrm{LN}(x)=\gamma\odot\frac{x-\mu}{\sqrt{\sigma^2+\epsilon}}+\beta$$

with learned scale $\gamma$ and shift $\beta$, one per feature. The mean and variance come from the features of that single vector for that single example — no other token, no other item in the batch contributes [2]. That matters for sequences: batch normalization estimates statistics across the batch, so it behaves differently at different batch sizes and needs separate bookkeeping per time step, whereas layer normalization does exactly the same computation at training and inference time [2]. For a model that has to generate one token at a time, and that sees sentences of wildly different lengths, that is the property you want.

So the division of labour is: the addition carries information and gradients past every layer; the normalization resets the scale of what each sublayer reads, so weights learned at layer 3 are not being asked to work on numbers that only existed at layer 3. The output of a LayerNorm sits at a fixed scale (before $\gamma$ and $\beta$ scale and shift it), so a sublayer's input distribution does not depend on how many layers precede it.

## Where you put the normalization decides whether you can go deep

The original Transformer applies the norm *after* the residual addition [1]:

$$x_{\ell+1}=\mathrm{LN}\big(x_\ell + F_\ell(x_\ell)\big)$$

This is the layout shown in Figure 1 of the paper as "Add & Norm". It was inherited from the ResNet convention and not separately justified. It works — but only with a carefully tuned recipe. The original model uses a learning-rate warmup of 4000 steps and Adam with $\beta_2 = 0.98$; skip either and the 6-layer base model fails to converge.

Xiong and colleagues explained why in 2020 [3]. With this placement, the gradient at initialization near the output layer has an expected magnitude that grows with depth, so the first optimizer steps overshoot badly unless the learning rate is held small at the start. Warmup is a workaround for that. Their analysis also shows the fix: move the normalization *inside* the residual branch,

$$x_{\ell+1}=x_\ell + F_\ell\big(\mathrm{LN}(x_\ell)\big)$$

so the addition acts on the unnormalized stream. Now the identity term in the Jacobian is not divided by the norm's derivative, gradients pass through depth nearly unchanged, and the expected gradient magnitude actually *decreases* slowly with depth ($\Theta(d\sqrt{\ln d}/\sqrt{L})$) rather than growing [3]. In experiments, this "pre-norm" version trains to comparable quality with no warmup at all [3]. A companion study trained 30-layer encoders stably where post-norm versions diverged around 20 layers [4].

Two honest qualifications, because this is a design trade-off rather than a solved problem:

- **Deeper is not automatically better.** When warmup *is* tuned, post-norm slightly beats pre-norm at shallow depth (6 layers) [5]. The reason is that pre-norm's identity path makes adjacent layers' outputs more similar to each other, so the upper layers end up doing less. The identity path that makes gradients flow is the same path that lets layers pass information through unprocessed.
- **It interacts with initialization.** A post-norm variant that scales each residual branch by a depth-dependent factor has been trained to depths in the hundreds, so post-norm's instability is an engineering problem rather than a fundamental barrier [6].

In practice, plain pre-norm with RMSNorm (a normalization that drops the mean subtraction) is the default across open large language models; a few designs use a "sandwich" with a norm both before and after each sublayer, and at least one recent family deliberately returns to a reordered post-norm paired with extra stabilizers [6]. The placement is still a knob people turn.

## What to take away

The encoder and decoder are the same block with a different set of sublayers: self-attention, feed-forward, plus cross-attention in the decoder, each wrapped in add-and-normalize. You can stack that block to any depth shape-wise because every sublayer preserves $d_{\text{model}}$. What actually makes depth *trainable* is the combination: the addition gives gradients and information a route that does not have to survive every operation in between, and the normalization keeps the scale each sublayer reads independent of how far down the stack it sits. Putting the norm inside the branch rather than on the residual sum is what lets the identity route run unmediated — and that is precisely the change that removed the original architecture's dependence on a painstaking learning-rate schedule.

Depth buys sequential computation: each layer reads what earlier layers wrote into the shared stream and adds its own contribution, so more layers means more steps of that kind. It is not free, and it is not the same thing as better — the same identity path that enables it also lets layers underuse themselves.

## Sources

1. [Vaswani et al., Attention Is All You Need — N=6 encoder/decoder layers, the sublayer list, LayerNorm(x + Sublayer(x)), cross-attention and masking](https://arxiv.org/abs/1706.03762)
2. [Ba, Kiros, Hinton, Layer Normalization — per-example statistics over a layer's hidden units, identical at train and test time](https://arxiv.org/abs/1607.06450)
3. [Xiong et al., On Layer Normalization in the Transformer Architecture — why post-norm needs warmup and pre-norm does not](https://proceedings.mlr.press/v119/xiong20b.html)
4. [Wang et al., Learning Deep Transformer Models for Machine Translation — pre-norm enabling 30-layer encoders where post-norm diverged](https://aclanthology.org/2019.iwslt-1.17.pdf)
5. [Takase et al., B2T Connection — post-norm's better shallow-depth performance and pre-norm's higher inter-layer output similarity](https://aclanthology.org/2023.findings-acl.192.pdf)
6. [Norm-placement overview — the original 4000-step warmup and β₂ = 0.98, DeepNet's deep post-norm, and current pre-norm/sandwich practice](https://lizeman.github.io/llm-arch-kb/normalization/norm-placement/)
7. [He et al., Deep Residual Learning for Image Recognition — the residual connection the Transformer's Add & Norm was inherited from](https://arxiv.org/abs/1512.03385)

---

Original article: https://eulore.ai/articles/transformer-residual-connections-layer-normalization-a21e4574

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
