Lesson 02-01

Tokenize and Vocabulary

12 min
3 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 201 • Runtime includes 4 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

02-01 Tokenize and Vocabulary

Why this matters

Vocabulary mapping is the numeric interface between text and model computation.

Intuition first (no jargon)

Each distinct token needs a deterministic ID so encoding and decoding stay reversible.

Code walkthrough

js
export function buildVocab(text) {
  return { tokenToId: new Map(), idToToken: [] };
}

export function encode(text, tokenToId) {}
export function decode(ids, idToToken) {}

Your task

Implement buildVocab, encode, and decode.

  • Use character-level tokens.
  • Ensure decode(encode(text)) === text for known tokens.
  • Throw clear errors for unknown IDs.

Hints

  • Keep token ordering deterministic (for stable tests).
  • Array.from(new Set(text)) can help build unique chars.
  • Use Map for fast lookups.

Check your thinking

  1. Why do deterministic IDs help debugging?
  2. What should happen with unknown tokens?
  3. Why keep both map and array forms?

Stretch (optional)

Add a reserved <UNK> token for unknown characters.

Likely test focus

  • Stable vocab creation on repeated runs.
  • Correct round-trip encode/decode.
  • Errors for invalid IDs.

What should improve

Text is now consistently encoded as token IDs for downstream counting and training.

Bridge to next lesson

Next lesson: count token-to-token transitions.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.