Definition
A world model predicts useful aspects of an environment’s evolution. DreamerV3 learns behaviours on imagined trajectories in a latent model. MuZero combines a learned decision-relevant model with tree search.
Equation
Gₕ sums rewards r over h steps with discount factor γ. Our toy sets γ = 1 and compares two deterministic paths.
Assumptions and notation
Gₕ sums rewards r over h steps with discount factor γ. Our toy sets γ = 1 and compares two deterministic paths.
Limits of interpretation
These systems use predictions differently. Model errors can accumulate; looking farther ahead does not guarantee a better decision. Our toy does not run either algorithm.
DreamerV3 learns a policy on imagined latent trajectories. MuZero predicts reward, value and policy for tree search. Reconstructing every detail of the world is not their common objective.