08-02 Top-k Sampling
Why this matters
Top-k sampling restricts decoding to the highest-scoring candidates before drawing a token.
Intuition first (no jargon)
Filter to the top k logits first, then sample within that reduced candidate set.
Paper grounding
- The paper evaluates translation with beam search (beam size 4) and length normalization.
- Top-k here is an explicit curriculum extension used to study decoding trade-offs.
Code walkthrough
jsexport function sampleTopK(logits, k, temperature = 1, rng = Math.random) {}
Your task
Implement top-k filtered sampling.
- Select indices of highest
klogits. - Apply temperature and softmax only to those indices.
- Return sampled ID from that subset.
Hints
- Clamp
kto[1, vocabSize]. - Keep mapping between filtered and original IDs.
- Test with
k = 1as deterministic argmax.
Check your thinking
- How does restricting candidates change output style?
- What is lost when
kis too small? - Why still keep temperature after top-k?
Stretch (optional)
Implement top-p (nucleus) sampling and compare outputs.
Likely test focus
- Output always from top-k set.
- Correct behavior for
k = 1andk >= vocab. - Stable handling of edge cases.
What should improve
You can now enforce candidate filtering before sampling for tighter generation control.
Bridge to next lesson
Next lesson: export attention maps.