Concept 1 / 5
Tokenization
Text is split into learned units called tokens. A token is not always a word; spelling, language and formatting change token counts and costs.
Section 2 · Mechanisms
An LLM predicts tokens from context. Scale and training make that simple objective surprisingly capable, but the model still generates likely continuations rather than consulting a guaranteed store of facts.
Vue interactive
Choisissez une vue. Le détail n’apparaît que lorsque vous le demandez.
Concept 1 / 5
Text is split into learned units called tokens. A token is not always a word; spelling, language and formatting change token counts and costs.