Attention is the thing that made modern language models possible, but it is usually explained as a formula rather than as an answer to a problem. This article is about the problem. By the end you should be able to say precisely what broke in the models that came before, why the fix that appeared in 2014 still wasn't enough, and what "let every word look at every other word directly" actually buys you.

What a sequence model has to do

A language model sees a sentence as a sequence of tokens (roughly, word pieces). Each token is turned into a vector of numbers — an embedding. So the input is not text but a list of nn vectors, x1,x2,…,xnx_1, x_2, \dots, x_n, where nn varies from sentence to sentence. The job is to produce another sequence: a translation, or the next token, or a label.

Two properties of the input matter for everything below: nn is not fixed, and order matters ("dog bites man" is not "man bites dog").

The RNN approach: one vector carried left to right

A recurrent neural network (RNN) processes the sequence one position at a time while carrying a running summary called the hidden state hth_t. At each step it takes the previous summary and the current token and produces a new summary:

ht=f(Wht−1+Uxt)h_t = f(W h_{t-1} + U x_t)

The same weights WW and UU are reused at every position, which is what lets an RNN handle sequences of any length. hth_t is meant to be a compressed memory of everything read so far.

For translation, the standard setup (called encoder–decoder or seq2seq) used two RNNs: an encoder that reads the source sentence until its final state hnh_n, and a decoder that starts from that state and generates the target sentence one word at a time. The whole source sentence was therefore handed over as a single fixed-size vector, usually written cc.

That single vector is where things go wrong.

Two costs of the recurrent design

Everything has to pass through one vector. The encoder produces a vector of fixed width — Cho and colleagues' 2014 model used 1000 hidden units, so 1000 numbers — and that is the only channel from source sentence to decoder, no matter how long the sentence is. The encoder must decide while reading what to keep and what to drop, and the decoder gets whatever survived. The authors of that work conjectured that this fixed-length representation simply doesn't have the capacity to encode a long sentence with complex structure, and that the network may sacrifice some topics in the input in order to remember others. It is a genuine information bottleneck, not a matter of training longer.

The empirical signature is exactly what you'd predict. Cho et al. showed translation quality degrades rapidly as sentence length grows. Bahdanau, Cho and Bengio's 2014 follow-up showed the same: the plain encoder–decoder's BLEU score fell off sharply with sentence length, while a system that used attention stayed flat even past 50-word sentences.

Computation is inherently sequential. h5h_5 cannot be computed until h4h_4 exists, because h4h_4 is an input to it. A 50-token sentence means 50 dependent steps that hardware cannot run in parallel. There's a second consequence buried in the same fact: information from token 1 reaches token 50 only by surviving 50 updates, and gradients during training have to travel back along the same chain. Long-range dependencies are therefore hard to learn, and this difficulty doesn't disappear with better gating — the 2017 Transformer paper notes that the fundamental constraint of sequential computation remains.

The 2014 fix: stop throwing away the intermediate states

Bahdanau et al. attacked the bottleneck, not the sequentiality. The move is small and worth stating plainly: keep every encoder state h1,…,hnh_1, \dots, h_n instead of only the last one, and let the decoder choose which ones to read at each step.

Concretely, when the decoder is about to produce output word ii, it computes a score eije_{ij} for every encoder state hjh_j, turns those scores into weights with a softmax so they're positive and sum to 1,

αij=exp⁡(eij)∑kexp⁡(eik)\alpha_{ij} = \frac{\exp(e_{ij})}{\sum_k \exp(e_{ik})}

and builds the context vector as a weighted sum of all encoder states:

ci=∑jαijhjc_i = \sum_j \alpha_{ij} h_j

The weights are the "attention." αij\alpha_{ij} says how much encoder position jj matters for output word ii. A small feedforward network computes the scores, and it is trained jointly with everything else — the gradient flows through the softmax, so the model learns where to look on its own. Nothing is hard-selected: the model blends, which is why the paper calls it a soft alignment.

A made-up two-dimensional illustration, with made-up weights: suppose h1=(1,0)h_1 = (1,0), h2=(0,2)h_2 = (0,2), h3=(1,1)h_3 = (1,1), and the learned weights are 0.1,0.7,0.20.1, 0.7, 0.2. Then

c=0.1(1,0)+0.7(0,2)+0.2(1,1)=(0.3, 1.6)c = 0.1(1,0) + 0.7(0,2) + 0.2(1,1) = (0.3,\,1.6)

The result is dominated by h2h_2 but still carries a little of the other two. Because the weights sum to 1, this is an averaging operation: the model's context at each step is a blend, and the blend can be different for every output word.

The intuition is a good one: instead of handing the decoder a single summary note and hoping, you hand it the whole notebook each time and let it look up the relevant line. The paper's qualitative check found alignments that matched human expectations of which source words belong to which target words. Quantitatively, the gains were largest for long sentences, and the model reached translation quality comparable to the era's phrase-based statistical systems — a striking result for a purely neural system at the time.

Note carefully what was fixed and what wasn't. The information bottleneck was eased: the source sentence now travels as a sequence of vectors rather than one vector. But the encoder was still a bidirectional RNN reading positions one at a time, and the decoder was still an RNN producing words one at a time. So the sequential step count stayed O(n)O(n), and the path length between distant input positions — how many computational steps a signal must pass through to get from token ii to token jj — stayed O(n)O(n) too. A bidirectional encoder does let each state blend context from both sides, but the signal from x1x_1 still has to walk through dozens of recurrent steps to influence h50h_{50}.

Self-attention: apply the same arithmetic inside one sequence

The 2017 Transformer took the weighted-sum machinery and asked why it needed a recurrence at all. For every position in the sequence, compute a score against every other position (including itself), softmax the scores into weights, and produce a new vector for that position as the weighted sum of all positions' vectors. That is self-attention: the sequence looking at itself.

Nothing here is sequential. Every position's output depends on the inputs directly, so all of them can be computed at once. In the paper's own comparison, self-attention needs O(1)O(1) sequential operations per layer and gives a maximum path length of O(1)O(1) between any two positions, against O(n)O(n) and O(n)O(n) for a recurrent layer. "Every word looks at every other word directly" is a plain-language way of saying that second number: one hop, not fifty.

A minimal sketch for a coder — same recipe as above, without the learned projections yet:

1import torch, torch.nn.functional as F
2
3# x: (n_tokens, dim) — one vector per token
4scores  = x @ x.T                      # (n, n): how much token i cares about token j
5weights = F.softmax(scores, dim=-1)    # each row sums to 1
6out     = weights @ x                  # (n, dim): each token becomes a blend of all tokens

In a real Transformer, x is projected into three different vectors per token — a query, a key, and a value — and the score is the dot product of one token's query with another's key. But the shape of the computation is exactly what you see here. Bahdanau-style attention is the same function with queries coming from the decoder and keys/values coming from the encoder; self-attention is the special case where all three come from the same sequence.

What attention costs, and what to be careful about

Attention is not free, and the 2017 paper is explicit about the trade-offs:

  • Quadratic cost. Per layer, self-attention costs O(n2⋅d)O(n^2 \cdot d) against O(n⋅d2)O(n \cdot d^2) for a recurrent layer. At every position you compare against every other position, so doubling the context length quadruples that part of the work. (The paper notes self-attention is actually cheaper than recurrence when n<dn < d, which is the common case for sentence-length inputs at their dimensions — but not for very long contexts.) The paper suggests restricting attention to a local window of size rr as a mitigation, and indeed much later work on long contexts is about dodging this quadratic term.
  • You have to inject order yourself. An RNN gets word order for free from the direction it reads in. A pile of weighted sums is order-blind — shuffling the input would shuffle the outputs the same way. Transformers add positional encodings, extra vectors that encode each position, to restore that information.
  • Averaging blurs. Combining many positions into one weighted average loses resolution. Multi-head attention — running several attention computations in parallel with different learned projections and concatenating the results — is the paper's countermeasure.
  • Causality needs masking. In a language model, position 5 must not see position 6 when predicting it. That's handled by forcing the future scores to a large negative value before the softmax, so their weights come out as zero.

Where this leaves you

The through-line is a single idea reused twice. First it was used to relieve a memory bottleneck: keep all encoder states, and let the decoder weigh them. Then it was used to eliminate a computational bottleneck: connect every position to every other position in one step, and compute them all at once.

The natural next question is how the scores are produced — what makes one token's query match another token's key, why the dot product is divided by dk\sqrt{d_k}, and how several heads get combined. That's the mechanics of scaled dot-product attention, and it's the right place to pick up from here.