07-04 Validation and Checkpoints
Why this matters
Validation estimates generalization, and checkpoints preserve exact model state for reproducible comparison.
Intuition first (no jargon)
Choose the checkpoint by best validation loss, not by latest epoch index.
Paper grounding
- The paper reports model quality with held-out metrics (for example BLEU on translation tasks).
- Reliable comparison requires fixed evaluation settings across runs.
Worked example
- Inputs: one run with
trainLoss = [2.4, 2.0, 1.8],valLoss = [2.5, 2.1, 2.3] - Shapes:
- per-eval losses list
[N] - checkpoint object includes
{ params, meta }
- per-eval losses list
- One computed step:
- checkpoint selection by minimum validation loss.
Code walkthrough
jsexport function evaluateLoss(ids, params, cfg) {} export function saveCheckpoint(params, meta) {} export function loadCheckpoint(blob) {}
evaluateLoss: mean held-out losssaveCheckpoint: persist weights + metadataloadCheckpoint: restore exact state
Your task
Implement evaluation and checkpoint helpers.
- Compute average loss on validation data.
- Serialize parameters and metadata.
- Reload exact state for continued training.
Common mistakes
- Missing architecture metadata in checkpoint
- Selection by latest epoch instead of min validation loss
- Incomplete round-trip equality check
Precision note
Reproducible resume and fair run comparison.
Hints
- Use plain JSON-compatible structures.
- Store config fields that affect architecture.
- Verify round-trip equality on small params.
Check your thinking
- Why can lowest train loss be a poor model?
- What metadata must be saved besides weights?
- Why test checkpoint round-trip?
Stretch (optional)
Keep top-3 checkpoints by validation score.
Likely test focus
- Evaluation returns finite average loss.
- Checkpoint serialize/deserialize round-trip.
- Loaded model reproduces same logits.
What should improve
Your training workflow now supports reliable model selection and exact resume.
Bridge to next lesson
Next: apply decoding controls during generation.