02-01 Tokenize and Vocabulary
Why this matters
Vocabulary mapping is the numeric interface between text and model computation.
Intuition first (no jargon)
Each distinct token needs a deterministic ID so encoding and decoding stay reversible.
Code walkthrough
jsexport function buildVocab(text) { return { tokenToId: new Map(), idToToken: [] }; } export function encode(text, tokenToId) {} export function decode(ids, idToToken) {}
Your task
Implement buildVocab, encode, and decode.
- Use character-level tokens.
- Ensure
decode(encode(text)) === textfor known tokens. - Throw clear errors for unknown IDs.
Hints
- Keep token ordering deterministic (for stable tests).
Array.from(new Set(text))can help build unique chars.- Use
Mapfor fast lookups.
Check your thinking
- Why do deterministic IDs help debugging?
- What should happen with unknown tokens?
- Why keep both map and array forms?
Stretch (optional)
Add a reserved <UNK> token for unknown characters.
Likely test focus
- Stable vocab creation on repeated runs.
- Correct round-trip encode/decode.
- Errors for invalid IDs.
What should improve
Text is now consistently encoded as token IDs for downstream counting and training.
Bridge to next lesson
Next lesson: count token-to-token transitions.