The last article built one attention head and ran it by hand: three tokens in, three new vectors out, each one a weighted blend of all the values. That computation is correct and complete, but it is not yet usable on language. Two things are missing, and both are structural rather than numerical.

First, forcing every token to produce exactly one blend means every relation that token participates in has to be squeezed into a single set of weights. Second, nothing in the attention formula refers to where a token sits in the sequence — so the model as written cannot tell "the cat sat on the mat" from "the mat sat on the cat". This article fixes both: multi-head attention and positional encoding.

One head has to compromise

Recall what a single head produces. Collect the nn token vectors as rows of an n×dmodeln \times d_{\text{model}} matrix XX. Then

head=softmax ⁣(QK⊤dk)V\mathrm{head} = \mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V

Row ii of that n×nn \times n matrix is one probability distribution over the whole sequence, and the output vector for token ii is one weighted sum of all value vectors.

Consider a concrete illustration, not a claim about any particular trained model. Take "the cat sat on the mat" and look at the token "sat". A useful thing for "sat" to retrieve is something from "cat", because the verb has to agree with its subject. Another useful thing is to keep track of its immediate neighbour "the", because local word order matters. With one head, "sat" gets one distribution to spend. Putting more mass on "cat" means less on "the", and — more importantly — the two retrieved value vectors get added into a single output vector. If a later layer needs to know which part of that vector came from the subject, that distinction has already been blended away. The original paper's own phrasing for this is that with a single head, "averaging inhibits" 1.

Run h heads in parallel, then glue them back

The fix is to stop trying to fix one head and instead run several at once, each with its own projections:

MultiHead(X)=Concat(head1,…,headh) WO,headi=Attention(XWiQ,  XWiK,  XWiV)\mathrm{MultiHead}(X)=\mathrm{Concat}(\mathrm{head}_1,\dots,\mathrm{head}_h)\,W^O,\qquad \mathrm{head}_i=\mathrm{Attention}(XW_i^Q,\;XW_i^K,\;XW_i^V)

Read the shapes carefully, because this is where the design becomes clear. In the original Transformer, dmodel=512d_{\text{model}}=512 and h=8h=8 heads, each with dk=dv=dmodel/h=64d_k=d_v=d_{\text{model}}/h=64 1.

  • XX is n×512n \times 512.
  • Head ii has its own WiQW_i^Q and WiKW_i^K, both 512×64512 \times 64, and its own WiVW_i^V, of shape 512×64512 \times 64. So head ii produces its own QiQ_i, KiK_i, ViV_i, each n×64n \times 64.
  • Head ii's output is n×64n \times 64.
  • Concatenating the eight outputs side by side along the feature axis gives n×(8×64)=n×512n \times (8 \times 64) = n \times 512 — the same width you started with.
  • WOW^O is a learned 512×512512 \times 512 matrix that mixes those eight blocks into the final n×512n \times 512 output.
Rendering

Why shrink each head instead of running eight full-width heads? Cost. One head costs about n2dkn^2 d_k operations, dominated by the n×nn \times n score matrix. Eight heads of width 64 cost 8n2⋅64=n2⋅5128 n^2 \cdot 64 = n^2 \cdot 512 — the same as a single head of width 512. That is what the paper means when it says the total computational cost is similar to single-head attention at full dimensionality 1. The trade is not free in memory: you now hold hh separate n×nn \times n score matrices, not one, which is hh times the storage.

What this actually buys you: each token now produces hh independent weight rows and hh separate blends, all computed from different projections. One head can spend its mass on the immediately preceding token while another spends it on a distant subject, and because each head has its own WVW^V, what each one copies back is different content too. The heads do not compete for one distribution. Finally, WOW^O lets the next layer treat the heads' outputs as one combined vector rather than a stack of separate results.

What heads actually do once trained

The split into subspaces is designed; the roles are not. Nothing in the loss says head 3 should track syntax. They are learned end-to-end along with everything else, and interpretability studies find a mixed picture: some heads in trained models do line up with recognizable functions — attending to fixed positional offsets, to separator tokens, or to specific syntactic relations such as the direct object of a verb — while many others do not have a clean description 5.

Pruning studies put a number on how little is essential. In one analysis of a translation model, removing 38 of 48 encoder heads cost only 0.15 BLEU, and the heads that survived pruning longest were exactly the ones with identifiable roles like tracking adjacent tokens or specific dependency relations 4. So the right mental model is: multi-head attention creates hh parallel subspaces and lets gradient descent decide what to put in them. Do not assume hh clean jobs.

Attention does not know word order

Now the second gap, and it is provable rather than empirical. Look again at the pieces of a single head. The score sijs_{ij} depends only on tokens ii and jj. Softmax normalizes along a row. The output for token ii sums over the value vectors of all tokens. No term anywhere refers to the index ii or jj as a position — indices appear only as labels for which token is which.

Make that precise. Let PP be a permutation matrix that reorders the tokens of XX. Each token's query, key and value move with it, so Q→PQQ \to PQ, K→PKK \to PK, V→PVV \to PV. The score matrix becomes

(PQ)(PK)⊤=P QK⊤P⊤(PQ)(PK)^\top = P\,QK^\top P^\top

and since softmax acts row by row, permuting rows and columns just permutes the normalized weights the same way, giving PAP⊤P A P^\top. The output is

(PAP⊤)(PV)=PA(P⊤P)V=PAV(P A P^\top)(P V) = P A (P^\top P) V = P A V

The output is the original output with its rows reordered, and the numbers inside each row are identical. This property is called permutation equivariance, and it means a bidirectionally-attending Transformer is order-blind: shuffle the input tokens and you get the same vectors back in shuffled order 6. In practice the model's later layers would then be fitting a function of a bag of words. Positional information has to be injected deliberately.

Positional encoding: a fingerprint for each position

What we need is a function that maps a position, an integer like 0, 1, 2, …, to a vector of length dmodeld_{\text{model}} that gets added elementwise to that token's embedding. It should give every position a distinct signature, keep values bounded (so position doesn't swamp the token's meaning), and make it easy for the model to compare positions.

The original Transformer uses:

PE(pos, 2i)=sin⁡ ⁣(pos10000 2i/dmodel),PE(pos, 2i+1)=cos⁡ ⁣(pos10000 2i/dmodel)PE_{(pos,\,2i)}=\sin\!\left(\frac{pos}{10000^{\,2i/d_{\text{model}}}}\right),\qquad PE_{(pos,\,2i+1)}=\cos\!\left(\frac{pos}{10000^{\,2i/d_{\text{model}}}}\right)

Here pospos is the token's position, and ii ranges from 00 to dmodel/2−1d_{\text{model}}/2-1, so each ii supplies a pair of dimensions — one sine, one cosine, sharing the same frequency. The constant 10000 is a fixed choice from the paper, not something derived; what matters is that the wavelengths form a geometric progression, from 2π2\pi at i=0i=0 up to roughly 10000⋅2π10000 \cdot 2\pi at the top 37.

Make it concrete with a tiny case, dmodel=4d_{\text{model}}=4, so i∈{0,1}i \in \{0,1\}. For i=0i=0 the denominator is 100000=110000^0=1, giving sin⁡(pos)\sin(pos) and cos⁡(pos)\cos(pos). For i=1i=1 the denominator is 100002/4=10010000^{2/4}=100, giving sin⁡(pos/100)\sin(pos/100) and cos⁡(pos/100)\cos(pos/100).

posposdim 0: sin⁡(pos)\sin(pos)dim 1: cos⁡(pos)\cos(pos)dim 2: sin⁡(pos/100)\sin(pos/100)dim 3: cos⁡(pos/100)\cos(pos/100)
00101
10.8410.5400.0100.99995
20.909−0.4160.0200.99980

Two rows are clearly different, and every value stays inside [−1,1][-1,1]. Notice how the two frequency bands behave differently. Dimensions 0 and 1 swing fast: positions 1 and 2 already differ a lot there. Dimensions 2 and 3 barely move — cos⁡(0.01)\cos(0.01) and cos⁡(0.02)\cos(0.02) differ only in the fifth decimal — so they stay nearly constant across short spans and only distinguish positions that are far apart. Stack enough frequency bands and you get the behaviour of a binary counter or a set of clock hands: a fast hand separates neighbours, a slow hand tells you which large block you are in, and the combination is unique per position.

There is a second reason for pairing sine with cosine at each frequency. The pair (sin⁡θ,cos⁡θ)(\sin\theta, \cos\theta) is a point on the unit circle, so a frequency is an angle you can rotate. Applying the angle-addition identities to one frequency ω\omega:

sin⁡(ω(pos+k))=sin⁡(ω pos)cos⁡(ωk)+cos⁡(ω pos)sin⁡(ωk)\sin(\omega(pos+k))=\sin(\omega\,pos)\cos(\omega k)+\cos(\omega\,pos)\sin(\omega k) cos⁡(ω(pos+k))=cos⁡(ω pos)cos⁡(ωk)−sin⁡(ω pos)sin⁡(ωk)\cos(\omega(pos+k))=\cos(\omega\,pos)\cos(\omega k)-\sin(\omega\,pos)\sin(\omega k)

The pair at pos+kpos+k is the pair at pospos rotated by the fixed angle ωk\omega k, and kk — not pospos — is what the rotation depends on 27. This is exactly the property the paper cites as its motivation: since PEpos+kPE_{pos+k} is a linear function of PEposPE_{pos} for any fixed offset kk, a linear layer such as WQW^Q or WKW^K can in principle learn to read off how far apart two tokens are, not just where each one is 1. Treat that as a design hypothesis with a clean algebraic justification, not as a guarantee about what a trained network does internally.

Two practical details. The encoding is added, not concatenated: xi←embedding(tokeni)+PE(i)x_i \leftarrow \text{embedding}(\text{token}_i) + PE(i), which works because both vectors have length dmodeld_{\text{model}}. And in the original paper the token embeddings are first scaled by dmodel\sqrt{d_{\text{model}}} before the addition 1. The paper also tried learned positional embeddings — a table of trainable vectors, one per position — and got nearly identical results on translation quality; the sinusoidal version was kept partly because it might extrapolate to sequences longer than those seen in training 1.

Where this leaves a modern language model

Both ideas survive in today's models, though the details move. Decoder-only LLMs use the same multi-head machinery but apply a causal mask that sets the scores of future positions to −∞-\infty before softmax, so each token can only attend backwards. That mask is fixed by position, which is why such models are no longer fully permutation-equivariant — shuffling tokens changes what each one is allowed to see, and work on position-free training suggests causal models can pick up some ordering information implicitly 6. Positional encodings are still added in most designs, but a popular modern variant, RoPE, skips the addition: it rotates each query and key vector by an angle proportional to that token's position, so that the dot product between a query and a key depends on their relative offset. Same trigonometry, applied inside the attention computation instead of to the input 2.

So the two pieces fit together like this. Multi-head attention gives each token several independent ways to look at the sequence and several independent things to retrieve from it. Positional encoding tells it which tokens are where, so that "cat sat on mat" and "mat sat on cat" are different inputs. Neither is sufficient alone: heads without position cannot use order, and position without heads still forces every relation through one blended vector.