Lesson 07-04

Validation and Checkpoints

12 min
3 exports
5 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 704 • Runtime includes 27 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

07-04 Validation and Checkpoints

Why this matters

Validation estimates generalization, and checkpoints preserve exact model state for reproducible comparison.

Intuition first (no jargon)

Choose the checkpoint by best validation loss, not by latest epoch index.

Paper grounding

  • The paper reports model quality with held-out metrics (for example BLEU on translation tasks).
  • Reliable comparison requires fixed evaluation settings across runs.

Worked example

  • Inputs: one run with trainLoss = [2.4, 2.0, 1.8], valLoss = [2.5, 2.1, 2.3]
  • Shapes:
    • per-eval losses list [N]
    • checkpoint object includes { params, meta }
  • One computed step:
    • checkpoint selection by minimum validation loss.

Code walkthrough

js
export function evaluateLoss(ids, params, cfg) {}
export function saveCheckpoint(params, meta) {}
export function loadCheckpoint(blob) {}
  • evaluateLoss: mean held-out loss
  • saveCheckpoint: persist weights + metadata
  • loadCheckpoint: restore exact state

Your task

Implement evaluation and checkpoint helpers.

  • Compute average loss on validation data.
  • Serialize parameters and metadata.
  • Reload exact state for continued training.

Common mistakes

  • Missing architecture metadata in checkpoint
  • Selection by latest epoch instead of min validation loss
  • Incomplete round-trip equality check

Precision note

Reproducible resume and fair run comparison.

Hints

  • Use plain JSON-compatible structures.
  • Store config fields that affect architecture.
  • Verify round-trip equality on small params.

Check your thinking

  1. Why can lowest train loss be a poor model?
  2. What metadata must be saved besides weights?
  3. Why test checkpoint round-trip?

Stretch (optional)

Keep top-3 checkpoints by validation score.

Likely test focus

  • Evaluation returns finite average loss.
  • Checkpoint serialize/deserialize round-trip.
  • Loaded model reproduces same logits.

What should improve

Your training workflow now supports reliable model selection and exact resume.

Bridge to next lesson

Next: apply decoding controls during generation.

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.