Lesson 07-01

Positional Information

12 min
2 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 701 • Runtime includes 24 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

07-01 Positional Information

Why this matters

Section 3.5 adds positional information because attention alone has no recurrence or convolutional order signal.

Intuition first (no jargon)

Token identity answers what the symbol is; positional signal answers where it appears.

Paper grounding

  • Because the model has no recurrence or convolution, the paper adds positional encodings to input embeddings.
  • Positional information is added at the bottom of the encoder and decoder stacks.

Worked example

  • Inputs: token embeddings and matching position embeddings for T = 3, dModel = 3.
  • Shapes:
    • token embeddings [T][dModel]
    • position embeddings [T][dModel]
    • output after add stays [T][dModel]
  • One computed step:
    • output vectors are elementwise sums of token and position vectors.
  • Output: identical token IDs can map to different vectors at different positions.

Code walkthrough

js
export function positionalEmbedding(maxLen, dModel, rng) {}
export function addPosition(tokenEmbeds, posEmbeds) {}
  • positionalEmbedding: position-indexed vectors
  • addPosition: elementwise add for first T positions

Your task

Implement positional embedding creation and addition.

  • Build learnable matrix [maxLen][dModel].
  • Slice first T rows for current sequence length.
  • Add position vectors to token vectors elementwise.

Common mistakes

  • Missing token+position addition before first block

Precision note

This lesson uses learned positional vectors as an implementation choice for easier inspection.

Hints

  • Validate T <= maxLen.
  • Keep token and position dims identical.
  • Use copied arrays to avoid mutation.

Check your thinking

  1. Why can two identical tokens need different vectors?
  2. What fails if no positional signal is added?
  3. Why is elementwise add a reasonable default?

Stretch (optional)

Implement sinusoidal positions and compare with learned positions.

Likely test focus

  • Correct positional matrix shape.
  • Correct addition output shape and values.
  • Distinct position vectors for distinct indices.

What should improve

Your model now injects explicit order information before block computation.

Bridge to next lesson

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.