SURYA NARAYAN BARIK
Deep Learning

Why Attention Is All You Need

A short primer on the mechanism reshaping modern ML — why a weighted average of vectors turned out to be one of the most consequential ideas in deep learning.

Recurrent networks read a sequence one token at a time, carrying everything they've seen so far in a single hidden state. That's an elegant idea, and a costly one: the further back a relevant word sits, the more that hidden state has to work to remember it. Attention throws out the queue and lets every position look directly at every other position, in parallel, in one step.

1. The problem with recurrence

An RNN processing "the animal didn't cross the street because it was too tired" has to preserve "animal" across six intervening tokens to resolve what "it" refers to. In practice, gradients shrink over that distance and the signal degrades — the further apart two related words are, the harder the model finds it to connect them. Stacking LSTMs and gates helped, but the sequential dependency itself never went away: token five cannot be processed until token four is done.

2. Self-attention in one picture

Self-attention removes the queue entirely. Every token looks at every other token directly, in one step, and decides how much to borrow from each.

Queries, keys, and values

Each token produces three vectors — a query, a key, and a value — by multiplying its embedding against three learned weight matrices. The query asks "what am I looking for," the key answers "what do I contain," and the value is "what I'll actually hand over if you attend to me."

Turning similarity into weights

The dot product of one token's query and another's key gives a similarity score: of everything else in this sequence, how much should I attend to you? Softmax turns those scores into weights that sum to one, which are then used to blend the value vectors — a weighted average, not a hard lookup.

it the was animal .
Figure 1 — "it" attends most strongly to "animal," weakly to everything else.

3. Scaled dot-product attention

Formally, for query matrix Q, key matrix K, and value matrix V:

# d_k = dimension of the key vectors, used to keep the softmax well-scaled
Attention(Q, K, V) = softmax( Q · K^T / sqrt(d_k) ) · V

Why divide by √d_k

This part is easy to skim past, but it's load-bearing: without it, dot products grow large as dimensionality increases, the softmax saturates, and gradients vanish almost everywhere except the single largest score.

4. Multi-head attention: why one isn't enough

A single attention operation learns one notion of "relevance." Multi-head attention runs several of these in parallel, so the model isn't stuck with just one.

Splitting into heads

Instead of one large attention operation, the model's dimensionality is split across several smaller heads, each with its own learned Q/K/V projections, running scaled dot-product attention independently. One head can track syntactic agreement while another tracks coreference or position.

Recombining the outputs

The outputs of every head are concatenated and projected back down to the original dimension — more perspectives on the same sequence, at roughly the same computational cost as one larger head.

The mechanism doesn't know about grammar or coreference. It only knows how to weight vectors by similarity — the linguistic structure is something the training process discovers, not something we build in.

5. Why it scales

The practical win is parallelism.

Parallelism over sequential computation

An RNN's dependency chain forces token-by-token computation; self-attention computes every pairwise interaction at once, which maps cleanly onto GPU matrix multiplication. That's the trade the Transformer made: quadratic cost in sequence length, in exchange for a computation graph with no sequential bottleneck.

What that trade buys in practice

For the sequence lengths and hardware of the last few years, that trade has paid off enormously — it's the reason a model can look at a token from ten thousand words back exactly as easily as the token right before it.

Closing thought

Recurrence models a sequence the way we read it — one word after another. Attention models it the way we understand it: everything in context at once, weighted by relevance. That shift, more than any single architectural trick, is what "attention is all you need" actually meant.

← Back to all posts