Lesson 07-03

Batching and Train Step

15 min
2 exports
4 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 703 • Runtime includes 26 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

07-03 Batching and Train Step

Why this matters

Batching makes optimization efficient and keeps tensor contracts consistent across updates.

Intuition first (no jargon)

Sample many windows, compute one loss, and apply one deterministic parameter update step.

Paper grounding

  • The paper trains with Adam using beta1 = 0.9, beta2 = 0.98, and epsilon = 1e-9.
  • It uses a warmup-based learning-rate schedule before inverse-square-root decay.

Worked example

  • Inputs: ids length 20, batchSize = 2, blockSize = 4
  • Shapes:
    • x batch (inputs) is [2][4]
    • y batch (targets) is [2][4]
    • logits from MiniGPT are [2][4][V]
    • loss is scalar []
  • One computed step:
    • sample start indices, build x, and shift by one token for y.

Code walkthrough

js
export function getBatch(ids, batchSize, blockSize, rng) {}
export function trainStep(batch, params, cfg) {}
  • getBatch samples deterministic windows when an RNG is provided.
  • trainStep runs forward pass, computes loss, and applies one parameter update.

Your task

Implement batch sampling and one optimization step.

  • Sample random start positions for contexts.
  • Build input and target tensors.
  • Compute loss and apply parameter update.

Common mistakes

  • Forgot one-token shift between x and y
  • Direct Math.random usage in batching
  • Non-deterministic RNG path in tests

Precision note

This lesson uses simple SGD-style updates to keep mechanics visible. Paper training uses Adam with warmup scheduling.

Hints

  • Keep RNG injectable for deterministic tests.
  • Reuse makeExamples logic conceptually.
  • Log step, loss, and learning rate.

Check your thinking

  1. Why random batches help generalization?
  2. Why is deterministic batching still useful during tests?
  3. Why monitor loss each step?

Stretch (optional)

Add gradient clipping for stability.

Likely test focus

  • Correct batch shapes.
  • Deterministic batches with seeded RNG.
  • Train step lowers loss trend on toy corpus.

What should improve

You can now run reproducible batched training updates.

Bridge to next lesson

Next: add validation and checkpointing.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.