Lesson 05-04

Causal Mask (No Peeking)

12 min
2 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 504 • Runtime includes 19 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

05-04 Causal Mask (No Peeking)

Why this matters

Decoder self-attention must block future-token access so position i attends only to positions <= i.

Intuition first (no jargon)

Apply the mask before softmax so illegal future positions receive effectively zero probability.

Paper grounding

  • For decoder self-attention, illegal future positions are masked before softmax by adding a large negative value to those logits.
  • This enforces that each position attends only to itself and earlier positions.

Worked example

  • Inputs: sequence length T = 4
  • Shapes:
    • mask is [4][4]
    • scores are [4][4]
    • masked scores remain [4][4]
  • One computed step:
    • valid causal mask rows are:
      • row 0: [1, 0, 0, 0]
      • row 1: [1, 1, 0, 0]
      • row 2: [1, 1, 1, 0]
      • row 3: [1, 1, 1, 1]
    • row 2 cannot assign attention to position 3.

Code walkthrough

js
export function causalMask(T) {}
export function applyMask(scores, mask) {}
  • causalMask: lower-triangular allow matrix
  • applyMask: large negative value for disallowed score cells

Your task

Implement masking utilities and integrate into attention.

  • causalMask(T) allows positions j <= i only.
  • Replace masked score entries with a very negative number.
  • Confirm masked weights for future positions become near zero.

Common mistakes

  • Flipped mask orientation (j > i allowed)
  • Mask applied after softmax instead of before
  • Non-negative sentinel value used for masked scores

Precision note

This curriculum uses -1e9 as a practical stand-in for negative infinity.

Hints

  • Use -1e9 as a practical negative infinity.
  • Apply mask before softmax.
  • Test with small T = 4 examples.

Check your thinking

  1. Why does masking happen on scores, not weights?
  2. What bug appears if mask is reversed?
  3. Why is this needed in both train and generate modes?

Stretch (optional)

Support an additional padding mask for variable-length batches.

Likely test focus

  • Correct lower-triangular mask structure.
  • Future weights near zero.
  • Causal constraints preserved across rows.

What should improve

Your attention pipeline now enforces autoregressive constraints correctly.

Bridge to next lesson

Next: run multiple attention heads in parallel.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.