05-03 Weighted Sum of Values
Why this matters
The paper's attention output multiplies weights by value vectors to produce contextualized representations.
Intuition first (no jargon)
Each output row is a weighted mixture of all value rows for that query position.
Paper grounding
- In Equation (1), softmax-normalized scores are multiplied by
Vto produce contextualized outputs. - This means each output vector is a weighted sum over value vectors.
Worked example
- Inputs:
- attention weights for token 1:
[0.8, 0.2] - value vectors:
V0 = [2, 0],V1 = [0, 4]
- attention weights for token 1:
- Shapes:
weights [T][T],V [T][dValue]- output is
[T][dValue]
- One computed step:
- output for token 1 is
0.8 * V0 + 0.2 * V1 = [1.6, 0.8]
- output for token 1 is
- Output: token 1 becomes a context-mixed vector, not just its original state.
Code walkthrough
jsexport function attentionOutput(weights, V) { // returns [T][dValue] }
Your task
Implement attentionOutput(weights, V).
- For each token row, compute weighted sum across all value vectors.
- Preserve output dimension
[T][dValue]. - Ensure deterministic numeric results.
Common mistakes
- Reused one weight row for every output token
- Wrote outputs with wrong
[T][dValue]shape
Precision note
Deterministic weighted-sum operation with fixed inputs.
Hints
- Triple nested loops are fine for learning.
- Initialize output with zeros.
- Verify tiny hand-calculated examples first.
Check your thinking
- Why can one token attend strongly to itself?
- What happens when weights are almost uniform?
- Why is this more flexible than fixed averaging?
Stretch (optional)
Return both output vectors and entropy of each attention row.
Likely test focus
- Correct weighted sums on toy input.
- Correct output shape.
- No mutation of
weightsorV.
What should improve
You can now turn attention weights into deterministic context vectors.
Bridge to next lesson
Next lesson: enforce autoregressive masking.