# Multi-Head Attention and Positional Encoding: Making Self-Attention Work on Real Language

Why one attention head has to compromise, how h parallel heads split the work and glue it back with W^O, and why self-attention without positional information sees language as a bag of tokens

> self-attention · transformer architecture · positional encoding · About 9 min · Oct 6

## Key points

1. A single attention head produces exactly one weight distribution and one blended vector per token, so different relations a token participates in have to share that distribution and get summed together; this is what the original paper calls averaging.
2. Multi-head attention runs h attention computations in parallel with separate learned projections per head, concatenates the outputs along the feature axis, and mixes them with a learned output matrix W^O back to d_model.
3. With h = 8 and d_k = d_v = d_model/h = 64, the concatenated output has exactly the input width again (8 × 64 = 512), and the total arithmetic cost is about the same as one full-width head — the extra cost is memory for h separate n × n score matrices.
4. The division into subspaces is designed; the roles of individual heads are learned. Studies find some heads with recognizable roles (positional offsets, delimiters, specific syntactic relations) and show many heads can be pruned with little loss, so h heads should not be read as h clean jobs.
5. Bidirectional self-attention is permutation equivariant: permuting input tokens permutes output rows identically and leaves the values inside each row unchanged, so without positional information the model effectively sees a bag of tokens.
6. Sinusoidal positional encoding assigns each position a bounded, unique vector built from sine/cosine pairs with geometrically increasing wavelengths from about 2π to 10000·2π; fast dimensions separate neighbours, slow dimensions separate distant blocks.
7. Pairing sine with cosine at each frequency makes each frequency a rotation, so PE at pos + k is a fixed rotation of PE at pos depending only on k — the paper's stated reason for this form, since it lets linear layers read relative offsets.
8. The encoding is added elementwise to the token embedding (with embeddings scaled by √d_model in the original paper); learned positional embeddings performed nearly identically, and the sinusoidal version was chosen partly for possible extrapolation to longer sequences.
9. Modern decoder-only LLMs keep multi-head attention, add a causal mask that leaks some order information, and often use RoPE, which rotates queries and keys by position so attention scores depend on relative offset.

---

The last article built one attention head and ran it by hand: three tokens in, three new vectors out, each one a weighted blend of all the values. That computation is correct and complete, but it is not yet usable on language. Two things are missing, and both are structural rather than numerical.

First, forcing every token to produce exactly one blend means every relation that token participates in has to be squeezed into a single set of weights. Second, nothing in the attention formula refers to where a token sits in the sequence — so the model as written cannot tell "the cat sat on the mat" from "the mat sat on the cat". This article fixes both: **multi-head attention** and **positional encoding**.

## One head has to compromise

Recall what a single head produces. Collect the $n$ token vectors as rows of an $n \times d_{\text{model}}$ matrix $X$. Then

$$\mathrm{head} = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V$$

Row $i$ of that $n \times n$ matrix is one probability distribution over the whole sequence, and the output vector for token $i$ is one weighted sum of all value vectors.

Consider a concrete illustration, not a claim about any particular trained model. Take "the cat sat on the mat" and look at the token "sat". A useful thing for "sat" to retrieve is something from "cat", because the verb has to agree with its subject. Another useful thing is to keep track of its immediate neighbour "the", because local word order matters. With one head, "sat" gets **one** distribution to spend. Putting more mass on "cat" means less on "the", and — more importantly — the two retrieved value vectors get added into a single output vector. If a later layer needs to know *which part* of that vector came from the subject, that distinction has already been blended away. The original paper's own phrasing for this is that with a single head, "averaging inhibits" [1].

## Run h heads in parallel, then glue them back

The fix is to stop trying to fix one head and instead run several at once, each with its own projections:

$$\mathrm{MultiHead}(X)=\mathrm{Concat}(\mathrm{head}_1,\dots,\mathrm{head}_h)\,W^O,\qquad \mathrm{head}_i=\mathrm{Attention}(XW_i^Q,\;XW_i^K,\;XW_i^V)$$

Read the shapes carefully, because this is where the design becomes clear. In the original Transformer, $d_{\text{model}}=512$ and $h=8$ heads, each with $d_k=d_v=d_{\text{model}}/h=64$ [1].

- $X$ is $n \times 512$.
- Head $i$ has its own $W_i^Q$ and $W_i^K$, both $512 \times 64$, and its own $W_i^V$, of shape $512 \times 64$. So head $i$ produces its own $Q_i$, $K_i$, $V_i$, each $n \times 64$.
- Head $i$'s output is $n \times 64$.
- Concatenating the eight outputs side by side along the feature axis gives $n \times (8 \times 64) = n \times 512$ — the same width you started with.
- $W^O$ is a learned $512 \times 512$ matrix that mixes those eight blocks into the final $n \times 512$ output.

```mermaid
flowchart LR
  X["X: n x 512"] --> PE["add positional encoding"]
  PE --> H1["head 1: own Q, K, V -> n x 64"]
  PE --> H2["head 2: own Q, K, V -> n x 64"]
  PE --> H8["head 8: own Q, K, V -> n x 64"]
  H1 --> C["concat along features: n x 512"]
  H2 --> C
  H8 --> C
  C --> O["W^O: mix heads -> n x 512"]
```

Why shrink each head instead of running eight full-width heads? Cost. One head costs about $n^2 d_k$ operations, dominated by the $n \times n$ score matrix. Eight heads of width 64 cost $8 n^2 \cdot 64 = n^2 \cdot 512$ — the same as a single head of width 512. That is what the paper means when it says the total computational cost is similar to single-head attention at full dimensionality [1]. The trade is not free in memory: you now hold $h$ separate $n \times n$ score matrices, not one, which is $h$ times the storage.

What this actually buys you: each token now produces $h$ independent weight rows and $h$ separate blends, all computed from *different* projections. One head can spend its mass on the immediately preceding token while another spends it on a distant subject, and because each head has its own $W^V$, what each one copies back is different content too. The heads do not compete for one distribution. Finally, $W^O$ lets the next layer treat the heads' outputs as one combined vector rather than a stack of separate results.

## What heads actually do once trained

The split into subspaces is designed; the *roles* are not. Nothing in the loss says head 3 should track syntax. They are learned end-to-end along with everything else, and interpretability studies find a mixed picture: some heads in trained models do line up with recognizable functions — attending to fixed positional offsets, to separator tokens, or to specific syntactic relations such as the direct object of a verb — while many others do not have a clean description [5].

Pruning studies put a number on how little is essential. In one analysis of a translation model, removing 38 of 48 encoder heads cost only 0.15 BLEU, and the heads that survived pruning longest were exactly the ones with identifiable roles like tracking adjacent tokens or specific dependency relations [4]. So the right mental model is: multi-head attention creates $h$ parallel subspaces and lets gradient descent decide what to put in them. Do not assume $h$ clean jobs.

## Attention does not know word order

Now the second gap, and it is provable rather than empirical. Look again at the pieces of a single head. The score $s_{ij}$ depends only on tokens $i$ and $j$. Softmax normalizes along a row. The output for token $i$ sums over the value vectors of all tokens. No term anywhere refers to the index $i$ or $j$ as a *position* — indices appear only as labels for which token is which.

Make that precise. Let $P$ be a permutation matrix that reorders the tokens of $X$. Each token's query, key and value move with it, so $Q \to PQ$, $K \to PK$, $V \to PV$. The score matrix becomes

$$(PQ)(PK)^\top = P\,QK^\top P^\top$$

and since softmax acts row by row, permuting rows and columns just permutes the normalized weights the same way, giving $P A P^\top$. The output is

$$(P A P^\top)(P V) = P A (P^\top P) V = P A V$$

The output is the original output with its rows reordered, and the numbers inside each row are identical. This property is called **permutation equivariance**, and it means a bidirectionally-attending Transformer is order-blind: shuffle the input tokens and you get the same vectors back in shuffled order [6]. In practice the model's later layers would then be fitting a function of a bag of words. Positional information has to be injected deliberately.

## Positional encoding: a fingerprint for each position

What we need is a function that maps a position, an integer like 0, 1, 2, …, to a vector of length $d_{\text{model}}$ that gets added elementwise to that token's embedding. It should give every position a distinct signature, keep values bounded (so position doesn't swamp the token's meaning), and make it easy for the model to compare positions.

The original Transformer uses:

$$PE_{(pos,\,2i)}=\sin\!\left(\frac{pos}{10000^{\,2i/d_{\text{model}}}}\right),\qquad PE_{(pos,\,2i+1)}=\cos\!\left(\frac{pos}{10000^{\,2i/d_{\text{model}}}}\right)$$

Here $pos$ is the token's position, and $i$ ranges from $0$ to $d_{\text{model}}/2-1$, so each $i$ supplies a *pair* of dimensions — one sine, one cosine, sharing the same frequency. The constant 10000 is a fixed choice from the paper, not something derived; what matters is that the wavelengths form a geometric progression, from $2\pi$ at $i=0$ up to roughly $10000 \cdot 2\pi$ at the top [3][7].

Make it concrete with a tiny case, $d_{\text{model}}=4$, so $i \in \{0,1\}$. For $i=0$ the denominator is $10000^0=1$, giving $\sin(pos)$ and $\cos(pos)$. For $i=1$ the denominator is $10000^{2/4}=100$, giving $\sin(pos/100)$ and $\cos(pos/100)$.

| $pos$ | dim 0: $\sin(pos)$ | dim 1: $\cos(pos)$ | dim 2: $\sin(pos/100)$ | dim 3: $\cos(pos/100)$ |
| --- | --- | --- | --- | --- |
| 0 | 0 | 1 | 0 | 1 |
| 1 | 0.841 | 0.540 | 0.010 | 0.99995 |
| 2 | 0.909 | −0.416 | 0.020 | 0.99980 |

Two rows are clearly different, and every value stays inside $[-1,1]$. Notice how the two frequency bands behave differently. Dimensions 0 and 1 swing fast: positions 1 and 2 already differ a lot there. Dimensions 2 and 3 barely move — $\cos(0.01)$ and $\cos(0.02)$ differ only in the fifth decimal — so they stay nearly constant across short spans and only distinguish positions that are far apart. Stack enough frequency bands and you get the behaviour of a binary counter or a set of clock hands: a fast hand separates neighbours, a slow hand tells you which large block you are in, and the combination is unique per position.

There is a second reason for pairing sine with cosine at each frequency. The pair $(\sin\theta, \cos\theta)$ is a point on the unit circle, so a frequency is an angle you can rotate. Applying the angle-addition identities to one frequency $\omega$:

$$\sin(\omega(pos+k))=\sin(\omega\,pos)\cos(\omega k)+\cos(\omega\,pos)\sin(\omega k)$$
$$\cos(\omega(pos+k))=\cos(\omega\,pos)\cos(\omega k)-\sin(\omega\,pos)\sin(\omega k)$$

The pair at $pos+k$ is the pair at $pos$ rotated by the fixed angle $\omega k$, and $k$ — not $pos$ — is what the rotation depends on [2][7]. This is exactly the property the paper cites as its motivation: since $PE_{pos+k}$ is a linear function of $PE_{pos}$ for any fixed offset $k$, a linear layer such as $W^Q$ or $W^K$ can in principle learn to read off *how far apart* two tokens are, not just where each one is [1]. Treat that as a design hypothesis with a clean algebraic justification, not as a guarantee about what a trained network does internally.

Two practical details. The encoding is **added**, not concatenated: $x_i \leftarrow \text{embedding}(\text{token}_i) + PE(i)$, which works because both vectors have length $d_{\text{model}}$. And in the original paper the token embeddings are first scaled by $\sqrt{d_{\text{model}}}$ before the addition [1]. The paper also tried *learned* positional embeddings — a table of trainable vectors, one per position — and got nearly identical results on translation quality; the sinusoidal version was kept partly because it might extrapolate to sequences longer than those seen in training [1].

## Where this leaves a modern language model

Both ideas survive in today's models, though the details move. Decoder-only LLMs use the same multi-head machinery but apply a **causal mask** that sets the scores of future positions to $-\infty$ before softmax, so each token can only attend backwards. That mask is fixed by position, which is why such models are no longer fully permutation-equivariant — shuffling tokens changes what each one is allowed to see, and work on position-free training suggests causal models can pick up some ordering information implicitly [6]. Positional encodings are still added in most designs, but a popular modern variant, **RoPE**, skips the addition: it rotates each query and key vector by an angle proportional to that token's position, so that the dot product between a query and a key depends on their *relative* offset. Same trigonometry, applied inside the attention computation instead of to the input [2].

So the two pieces fit together like this. Multi-head attention gives each token several independent ways to look at the sequence and several independent things to retrieve from it. Positional encoding tells it which tokens are where, so that "cat sat on mat" and "mat sat on cat" are different inputs. Neither is sufficient alone: heads without position cannot use order, and position without heads still forces every relation through one blended vector.

## Sources

1. [Vaswani et al., Attention Is All You Need — multi-head formulas, h=8, d_k=d_v=64, W^O shape, sinusoidal positional encoding and its relative-position motivation](https://arxiv.org/html/1706.03762v4)
2. [Hugging Face blog: Designing Positional Encoding — deriving the sinusoidal rotation matrix and how RoPE rotates queries and keys](https://github.com/huggingface/blog/blob/main/designing-positional-encoding.md)
3. [Understanding Positional Encoding in Transformers — worked walk-through of the sin/cos formula and the geometric progression of wavelengths](https://erdem.pl/2021/05/understanding-positional-encoding-in-transformers/)
4. [Voita et al., Analyzing Multi-Head Self-Attention — pruning 38 of 48 encoder heads for a 0.15 BLEU drop, and the roles of surviving heads](https://aclanthology.org/P19-1580.pdf)
5. [Clark et al., What Does BERT Look At? — heads attending to delimiters, fixed positional offsets, and specific dependency relations](https://nlp.stanford.edu/pubs/clark2019what.pdf)
6. [Position Information Emerges in Causal Transformers Without Positional Encodings — causal attention is not permutation-equivariant](https://aclanthology.org/2025.coling-main.632.pdf)
7. [Linearity of Sinusoidal Positional Encodings — verification that PE(pos+k) is a linear function of PE(pos)](https://cs.brown.edu/courses/cs146/assets/files/linearity.pdf)

---

Original article: https://eulore.ai/articles/multi-head-attention-and-positional-encoding-40cee95e

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
