02-02 Build Bigram Counts
Why this matters
Bigram counts capture first-order local order by tracking which token follows which.
Intuition first (no jargon)
A transition matrix stores one-step sequence memory for each current token.
Code walkthrough
jsexport function buildBigramCounts(ids, vocabSize) { // returns vocabSize x vocabSize matrix of counts }
Your task
Implement buildBigramCounts(ids, vocabSize).
- Create a zero-initialized 2D array.
- For each adjacent pair
(ids[i], ids[i + 1]), increment count. - Skip counting when sequence length is less than 2.
Hints
- Row index is current token.
- Column index is next token.
- Validate ID range before incrementing.
Check your thinking
- What does
counts[a][b]mean? - Why do we need both row and column dimensions?
- How is this different from unigram counts?
Stretch (optional)
Track sentence-start and sentence-end markers as special IDs.
Likely test focus
- Correct matrix shape.
- Correct increments for toy ID sequences.
- Handles edge lengths safely.
What should improve
Your model now represents local sequential structure beyond unigram frequency.
Bridge to next lesson
Next: convert count rows into probabilities.