Lesson 05-02

Attention Scores and Weights

12 min
2 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 502 • Runtime includes 17 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

05-02 Attention Scores and Weights

Why this matters

Scaled dot-product attention computes QK^T / sqrt(dKey) and applies row-wise softmax to produce attention weights.

Intuition first (no jargon)

Each query row scores all keys, then softmax turns scores into a normalized weighting distribution.

Paper grounding

  • Scaled dot-product attention divides by sqrt(d_k) before softmax.
  • The paper explains this scaling prevents large dot-product magnitudes that would push softmax into very small-gradient regions.

Worked example

  • Inputs:
    • Q = [[1, 0], [0, 1]]
    • K = [[1, 1], [1, -1]]
    • T = 2, dKey = 2
  • Shapes:
    • Q [2][2], K [2][2]
    • S = QK^T is [2][2]
    • weights after row softmax is [2][2]
  • One computed step:
    • S = [[1, 1], [1, -1]]
    • Scale scores by 1 / sqrt(dKey) before softmax.
    • Row-0 softmax after scaling is approximately [0.50, 0.50]; row-1 is approximately [0.804, 0.196].
  • Output: token 1 attends much more to position 0 than position 1.

Code walkthrough

js
export function attentionScores(Q, K, scale = true) {}
export function attentionWeights(scores) {}
  • attentionScores computes QK^T and applies optional 1 / sqrt(dKey) scaling.
  • attentionWeights normalizes each query row independently.

Your task

Implement score and weight computation.

  • Compute score matrix S = Q * K^T.
  • If scale, divide by sqrt(dKey).
  • Softmax each row to get weights.

Common mistakes

  • Softmax over full matrix instead of per row
  • Missing sqrt(dKey) scaling before softmax

Precision note

Decoder self-attention: position i attends only to positions up to and including i.

Hints

  • Reuse stable softmax from Unit 03.
  • Row i corresponds to token i asking for context.
  • Verify each weight row sums near 1.

Check your thinking

  1. Why apply softmax per row?
  2. Why use scaling by sqrt(dKey)?
  3. What does a sharp distribution mean?

Stretch (optional)

Add function to return top-2 attended positions per token.

Likely test focus

  • Score matrix shape [T][T].
  • Row sums near 1 after softmax.
  • Correct scaling behavior.

What should improve

You can now compute stable attention weight matrices that sum to 1 per query row.

Bridge to next lesson

Next lesson: apply weights to value vectors.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.