The last article ended on a claim rather than a mechanism: attention lets any two tokens interact in one step, instead of pushing a signal through a chain of hidden states. This article fills in the mechanism, but only the smallest complete piece of it — the part that takes a handful of token vectors and produces a new vector for each one. By the end you should be able to run the whole computation by hand on three tokens, and say why each step is there rather than just what it looks like.
Every token gets three vectors
A Transformer receives the same thing the RNN did: a sequence of token vectors, each of length . Call them .
For each token, the model computes three vectors by multiplying with three separate weight matrices:
The matrices , , are learned, and — exactly like the RNN's shared weights from last time — the same matrices are applied at every position. That is what makes this self-attention: every token is projected by the same rules, so all tokens end up in one shared query space and one shared key space, where they can be compared with each other.
The names come from database lookup. You make a query, each stored item advertises a key, and when the two match you get back the value. Treat this as a naming convention rather than a description: nothing here is matched discretely or retrieved intact. The match is a number, and the retrieval is a blend.
Why three matrices instead of one? If and were produced by the same projection, the score between tokens and would be symmetric — 's interest in would equal 's interest in . Language is not symmetric. In "the cat sat on the mat", the verb "sat" has reason to care about "cat"; "cat" has much less reason to care about "sat". Separate projections make the relation directed, and let the model learn what "asking a question" and "advertising an answer" mean for its own data.
Two dimensions matter here. Queries and keys must have the same length, call it , because we are about to take their dot product. Values can have their own length , since they only ever get added together. In the original Transformer, and 1.
Scoring a pair: the dot product
The score between query and key is
This is large and positive when the two vectors point in the same direction, near zero when they are unrelated, and negative when they point in opposite directions. It is not a pure cosine — vector length matters too — but the ranking behaviour is what carries the meaning: a high score says "this key is relevant to this query".
Doing this for every query against every key produces an table of scores. That is already visible, and it is where attention's quadratic cost comes from; we will come back to it.
Why the scores are divided by
This is the "scaled" in scaled dot-product attention, and it is the one step that is not obvious. The argument is statistical. Suppose the components of and are roughly independent, with mean 0 and variance 1 — which is what normalization layers in the network are there to encourage. Then each term has mean 0 and variance 1, and is a sum of such independent terms, so
A quantity with variance typically lands around . So as grows, the scores get larger, not merely noisier. With , typical scores sit around instead of .
That matters because the next step is a softmax, which exponentiates. A gap of 8 between two logits means the higher one receives about times the weight of the lower one. So a score table whose entries are spread over roughly produces rows that are nearly one-hot: one token takes essentially all the weight and the rest get almost none. Two things go wrong. First, attention stops being a blend and becomes a hard pick of a single token. Second — and this is what hurts training — when a softmax row is nearly one-hot, its gradient is near zero, so the model barely learns from it. The authors of the original paper put it plainly: for large the dot products "grow large in magnitude, pushing the softmax function into regions where it has extremely small gradients" 1.
Dividing every score by brings the variance back to 1. That is the whole trick, and it explains the specific denominator: pulling a constant out of a variance squares it, so to turn a variance of into 1 you divide the values by , not by . The paper also reports that this is not just theory — for small , scaled and unscaled dot-product attention behave similarly, but for large the unscaled version performs worse, and scaling fixes it 1.
From scores to weights, and from weights to output
Softmax is applied row by row — for each query, across all the keys:
Every is positive, and each row sums to exactly 1. The row is therefore a probability distribution over the tokens: how much of this token's output should come from each position. Softmax never outputs an exact 0 for a finite input — it just makes small scores very small.
The output for token is the weighted sum of the value vectors:
This form is worth dwelling on, because it limits what attention can express. The weights are non-negative and sum to 1, so the output is a convex combination of the value vectors — a point inside the shape they span. If the row is nearly uniform, the output is close to the plain average of all the values. If one weight is near 1, the output is close to that single value. Attention can only re-mix the value vectors it was handed; it cannot produce a direction that is not a blend of them.
The same thing in matrix form
Pack the vectors into rows: is , is , is . Then everything above is one expression:
The shapes trace out the meaning. is , and entry is exactly . The softmax runs along the last axis, so it normalizes each row. Multiplying that weight matrix by the value matrix gives an output — one new vector per input token. In code, skipping masking and dropout, this is about four lines 3:
1scores = Q @ K.transpose(-2, -1) / math.sqrt(d_k) # (n, n)
2weights = scores.softmax(dim=-1) # row-wise, each row sums to 1
3out = weights @ V # (n, d_v)Doing it by hand
Take tokens, , , with simple integer rows:
Step 1 — raw scores . Each entry is the dot product of one query row with one key row:
For instance, entry is , and entry is .
Step 2 — divide by :
Step 3 — softmax each row. Using , , :
- row 1:
- row 2:
- row 3:
Step 4 — multiply the weight matrix by . For row 2:
Doing the same for rows 1 and 3 gives and .
Read the result back against the scores. matched best — score 2, the largest in the table — and token 2's output does lean on . But nothing is sharp: the strongest weight is 0.503, so nearly half of token 2's output still comes from the other two tokens. Rows 1 and 3 happen to produce mirrored weight patterns, purely because and are the mirror images of and in these particular matrices. That is a property of the numbers I chose, not a rule of attention.
Why this is fast, and what it costs
Recall the second problem the previous article raised about RNNs: cannot be computed until exists, so a sequence of length costs dependent steps, and a GPU spends most of that time waiting. Nothing in the recipe above has that shape. Row of the weight matrix depends only on and ; no output feeds into another. The entire score table is a single matrix multiply, with only the row-wise softmax following it. The recurrence is gone.
The price sits in the shape of : it is , so compute and memory grow with the square of the sequence length. Doubling the context length roughly quadruples the cost of the score matrix. Attention traded a long chain of dependent steps for a quadratic table of independent ones — a good trade on parallel hardware, and the reason long context windows are expensive rather than free.
What this piece does not include
Four things are deliberately missing, and it helps to know they are separate pieces so you do not go looking for them inside the formula.
- Position. treats the input as a set, not a sequence. Permute the tokens and the output rows are permuted in exactly the same way, so this computation cannot tell "dog bites man" from "man bites dog". Real Transformers add position information to the token vectors before this step.
- Multiple heads. The original model runs this whole computation times in parallel with eight different, smaller projections (each rather than ) and concatenates the results. Same formula, repeated, then re-projected 1.
- Causal masking. To generate text left to right, you set for every key that lies in the future of query , before the softmax. Those positions then receive weight exactly 0, and each position can only see itself and the past. Nothing else changes 2.
- Self- vs cross-attention. Here , , and all come from one sequence. Bahdanau attention in the previous article was cross-attention: queries from the decoder, keys and values from the encoder. Same arithmetic, different source of the three matrices.