02-04 Generate with Bigram
Why this matters
Bigram generation is the first model in this curriculum that conditions on immediate context.
Intuition first (no jargon)
The current token selects the next-token distribution at each generation step.
Code walkthrough
jsexport function generateBigram(startId, steps, probs, rng = Math.random) {} export function averageSurprise(ids, probs) {}
Your task
Implement generation and a simple quality metric.
- Generate IDs autoregressively from
startId. - Decode output to text.
- Compute average surprise from predicted probabilities on a sequence.
Hints
- Surprise is
-log(p)averaged over steps. - Guard against
p = 0with a tiny epsilon. - Compare bigram and unigram on the same validation text.
Check your thinking
- Why is lower average surprise better?
- Why can bigram still fail on long dependencies?
- What does autoregressive mean in your own words?
Stretch (optional)
Add a helper that prints side-by-side unigram and bigram samples.
Likely test focus
- Generated length and ID validity.
- Surprise value is finite.
- Bigram beats unigram on toy validation data.
What should improve
You can now generate context-sensitive samples and compare against unigram output.
Bridge to next lesson
Next lesson: move from count tables to trainable neural parameters.