Predicting to choose an action
Explore longer horizons and the risk of inaccurate predictions.
Teaching simulation · fictional values
A gain now or a gain later?
Two deterministic paths, without discounting. Compare the model prediction with the toy’s reference outcome.
| Path | Step 1 | Step 2 | Step 3 | Σ |
|---|---|---|---|---|
| A | 4 | 0 | 0 | 4 |
| B | 0 | 0 | 9 | 0 |
| Path | Step 1 | Step 2 | Step 3 | Σ |
|---|---|---|---|---|
| A | 4 | 0 | 0 | 4 |
| B | 0 | 0 | 9 | 0 |
Dimmed steps are outside the horizon.
Model choice : A
Best path at this horizon in the reference : A
DreamerV3 learns a policy on imagined latent trajectories. MuZero predicts reward, value and policy for tree search. Reconstructing every detail of the world is not their common objective.
Check my understanding
Before answering, form your own explanation.
A longer horizon can amplify the effects of a mistaken prediction.
Takeaway
These systems use predictions differently. Model errors can accumulate; looking farther ahead does not guarantee a better decision. Our toy does not run either algorithm.
Go deeperScientific sources
- Hafner et al. · 2023Mastering Diverse Domains through World ModelsDreamerV3 · arXiv v2 · 2024Original publication
- Schrittwieser et al. · 2019Mastering Atari, Go, Chess and Shogi by Planning with a Learned ModelarXiv v2 · 2020Original publication
Teaching synthesis of the cited sources. Published results remain tied to their tasks, protocols and budgets.