Transformers: understanding attention
Compare architectures, training methods and complete systems at the right level.
Teaching simulation · fictional values
Numerical example: one query, three keys
K = [[1,0],[0,1],[1,1]], V = [[1,0],[0,1],[1,1]], Q = [a,1] and dₖ = 2. These numbers are chosen for the exercise.
Attention weights
Output = [0.752 ; 0.752]
A causal mask excludes future positions in autoregressive prediction. Projections, multiple heads, local layers and residual connections complete a Transformer. Attention weights alone are not causal proof of its behaviour.
Check my understanding
Before answering, form your own explanation.
Transformer, diffusion and MuZero describe exactly the same design level.
Takeaway
Diffusion is a modelling and generation method that can use different networks. DreamerV3 and MuZero organise components into learning and decision systems. These names are not interchangeable categories.
Go deeperScientific sources
- Vaswani et al. · 2017Attention Is All You NeedarXiv v7 · 2023Original publication
- Hochreiter & Schmidhuber · 1997Long Short-Term MemoryOriginal publication
- Ho, Jain & Abbeel · 2020Denoising Diffusion Probabilistic ModelsOriginal publication
- Jain & Wallace · 2019Attention is not ExplanationOriginal publication
Teaching synthesis of the cited sources. Published results remain tied to their tasks, protocols and budgets.