Skip to contentExperiment
Transformers: understanding attention
Compare architectures, training methods and complete systems at the right level.
Essentials
How it works
A Transformer combines attention, position information and local transformations, among other components. A recurrent network updates state over successive inputs. Architecture describes how to compute; the training objective describes what optimisation improves.
Takeaway
Diffusion is a modelling and generation method that can use different networks. DreamerV3 and MuZero organise components into learning and decision systems. These names are not interchangeable categories.
Compare the right levels
- Architecture
- Transformer · Recurrent network
- Learning / generation method
- Supervised · self-supervised · reinforcement · diffusion
- Learning and decision system
- DreamerV3 · MuZero
From concept to practicePut it into practice
Scientific sources
- Vaswani et al. · 2017Attention Is All You NeedarXiv v7 · 2023Original publication
- Hochreiter & Schmidhuber · 1997Long Short-Term MemoryOriginal publication
- Ho, Jain & Abbeel · 2020Denoising Diffusion Probabilistic ModelsOriginal publication
- Jain & Wallace · 2019Attention is not ExplanationOriginal publication
Teaching synthesis of the cited sources. Published results remain tied to their tasks, protocols and budgets.