Definition
A Transformer combines attention, position information and local transformations, among other components. A recurrent network updates state over successive inputs. Architecture describes how to compute; the training objective describes what optimisation improves.
Scaled dot-product attention
Q contains queries, K keys and V values. QKᵀ is scaled by √dₖ; row-wise softmax produces weights that combine V.
Compare Q and K
Q [n_q × dₖ]
Kᵀ [dₖ × nₖ]
↓
QKᵀ [n_q × nₖ]Normalise each row
softmax(QKᵀ / √dₖ)
[n_q × nₖ]
Σⱼ wᵢⱼ = 1Combine the values V
W [n_q × nₖ]
V [nₖ × dᵥ]
↓
O [n_q × dᵥ]One head is shown. Q has n_q rows; K and V have nₖ rows. Q and K share dₖ columns. In self-attention, they are projected from the same sequence.
Numerical example: one query, three keysAssumptions and notation
A causal mask excludes future positions in autoregressive prediction. Projections, multiple heads, local layers and residual connections complete a Transformer. Attention weights alone are not causal proof of its behaviour.
Limits of interpretation
Diffusion is a modelling and generation method that can use different networks. DreamerV3 and MuZero organise components into learning and decision systems. These names are not interchangeable categories.
Compare the right levels
- Architecture
- Transformer · Recurrent network
- Learning / generation method
- Supervised · self-supervised · reinforcement · diffusion
- Learning and decision system
- DreamerV3 · MuZero