The last article built one attention head and ran it by hand: three tokens in, three new vectors out, each one a weighted blend of all the values. That computation is correct and complete, but it is not yet usable on language. Two things are missing, and both are structural rather than numerical.
First, forcing every token to produce exactly one blend means every relation that token participates in has to be squeezed into a single set of weights. Second, nothing in the attention formula refers to where a token sits in the sequence — so the model as written cannot tell "the cat sat on the mat" from "the mat sat on the cat". This article fixes both: multi-head attention and positional encoding.
One head has to compromise
Recall what a single head produces. Collect the token vectors as rows of an matrix . Then
Row of that matrix is one probability distribution over the whole sequence, and the output vector for token is one weighted sum of all value vectors.
Consider a concrete illustration, not a claim about any particular trained model. Take "the cat sat on the mat" and look at the token "sat". A useful thing for "sat" to retrieve is something from "cat", because the verb has to agree with its subject. Another useful thing is to keep track of its immediate neighbour "the", because local word order matters. With one head, "sat" gets one distribution to spend. Putting more mass on "cat" means less on "the", and — more importantly — the two retrieved value vectors get added into a single output vector. If a later layer needs to know which part of that vector came from the subject, that distinction has already been blended away. The original paper's own phrasing for this is that with a single head, "averaging inhibits" 1.
Run h heads in parallel, then glue them back
The fix is to stop trying to fix one head and instead run several at once, each with its own projections:
Read the shapes carefully, because this is where the design becomes clear. In the original Transformer, and heads, each with 1.
- is .
- Head has its own and , both , and its own , of shape . So head produces its own , , , each .
- Head 's output is .
- Concatenating the eight outputs side by side along the feature axis gives — the same width you started with.
- is a learned matrix that mixes those eight blocks into the final output.
Why shrink each head instead of running eight full-width heads? Cost. One head costs about operations, dominated by the score matrix. Eight heads of width 64 cost — the same as a single head of width 512. That is what the paper means when it says the total computational cost is similar to single-head attention at full dimensionality 1. The trade is not free in memory: you now hold separate score matrices, not one, which is times the storage.
What this actually buys you: each token now produces independent weight rows and separate blends, all computed from different projections. One head can spend its mass on the immediately preceding token while another spends it on a distant subject, and because each head has its own , what each one copies back is different content too. The heads do not compete for one distribution. Finally, lets the next layer treat the heads' outputs as one combined vector rather than a stack of separate results.
What heads actually do once trained
The split into subspaces is designed; the roles are not. Nothing in the loss says head 3 should track syntax. They are learned end-to-end along with everything else, and interpretability studies find a mixed picture: some heads in trained models do line up with recognizable functions — attending to fixed positional offsets, to separator tokens, or to specific syntactic relations such as the direct object of a verb — while many others do not have a clean description 5.
Pruning studies put a number on how little is essential. In one analysis of a translation model, removing 38 of 48 encoder heads cost only 0.15 BLEU, and the heads that survived pruning longest were exactly the ones with identifiable roles like tracking adjacent tokens or specific dependency relations 4. So the right mental model is: multi-head attention creates parallel subspaces and lets gradient descent decide what to put in them. Do not assume clean jobs.
Attention does not know word order
Now the second gap, and it is provable rather than empirical. Look again at the pieces of a single head. The score depends only on tokens and . Softmax normalizes along a row. The output for token sums over the value vectors of all tokens. No term anywhere refers to the index or as a position — indices appear only as labels for which token is which.
Make that precise. Let be a permutation matrix that reorders the tokens of . Each token's query, key and value move with it, so , , . The score matrix becomes
and since softmax acts row by row, permuting rows and columns just permutes the normalized weights the same way, giving . The output is
The output is the original output with its rows reordered, and the numbers inside each row are identical. This property is called permutation equivariance, and it means a bidirectionally-attending Transformer is order-blind: shuffle the input tokens and you get the same vectors back in shuffled order 6. In practice the model's later layers would then be fitting a function of a bag of words. Positional information has to be injected deliberately.
Positional encoding: a fingerprint for each position
What we need is a function that maps a position, an integer like 0, 1, 2, …, to a vector of length that gets added elementwise to that token's embedding. It should give every position a distinct signature, keep values bounded (so position doesn't swamp the token's meaning), and make it easy for the model to compare positions.
The original Transformer uses:
Here is the token's position, and ranges from to , so each supplies a pair of dimensions — one sine, one cosine, sharing the same frequency. The constant 10000 is a fixed choice from the paper, not something derived; what matters is that the wavelengths form a geometric progression, from at up to roughly at the top 37.
Make it concrete with a tiny case, , so . For the denominator is , giving and . For the denominator is , giving and .
| dim 0: | dim 1: | dim 2: | dim 3: | |
|---|---|---|---|---|
| 0 | 0 | 1 | 0 | 1 |
| 1 | 0.841 | 0.540 | 0.010 | 0.99995 |
| 2 | 0.909 | −0.416 | 0.020 | 0.99980 |
Two rows are clearly different, and every value stays inside . Notice how the two frequency bands behave differently. Dimensions 0 and 1 swing fast: positions 1 and 2 already differ a lot there. Dimensions 2 and 3 barely move — and differ only in the fifth decimal — so they stay nearly constant across short spans and only distinguish positions that are far apart. Stack enough frequency bands and you get the behaviour of a binary counter or a set of clock hands: a fast hand separates neighbours, a slow hand tells you which large block you are in, and the combination is unique per position.
There is a second reason for pairing sine with cosine at each frequency. The pair is a point on the unit circle, so a frequency is an angle you can rotate. Applying the angle-addition identities to one frequency :
The pair at is the pair at rotated by the fixed angle , and — not — is what the rotation depends on 27. This is exactly the property the paper cites as its motivation: since is a linear function of for any fixed offset , a linear layer such as or can in principle learn to read off how far apart two tokens are, not just where each one is 1. Treat that as a design hypothesis with a clean algebraic justification, not as a guarantee about what a trained network does internally.
Two practical details. The encoding is added, not concatenated: , which works because both vectors have length . And in the original paper the token embeddings are first scaled by before the addition 1. The paper also tried learned positional embeddings — a table of trainable vectors, one per position — and got nearly identical results on translation quality; the sinusoidal version was kept partly because it might extrapolate to sequences longer than those seen in training 1.
Where this leaves a modern language model
Both ideas survive in today's models, though the details move. Decoder-only LLMs use the same multi-head machinery but apply a causal mask that sets the scores of future positions to before softmax, so each token can only attend backwards. That mask is fixed by position, which is why such models are no longer fully permutation-equivariant — shuffling tokens changes what each one is allowed to see, and work on position-free training suggests causal models can pick up some ordering information implicitly 6. Positional encodings are still added in most designs, but a popular modern variant, RoPE, skips the addition: it rotates each query and key vector by an angle proportional to that token's position, so that the dot product between a query and a key depends on their relative offset. Same trigonometry, applied inside the attention computation instead of to the input 2.
So the two pieces fit together like this. Multi-head attention gives each token several independent ways to look at the sequence and several independent things to retrieve from it. Positional encoding tells it which tokens are where, so that "cat sat on mat" and "mat sat on cat" are different inputs. Neither is sufficient alone: heads without position cannot use order, and position without heads still forces every relation through one blended vector.