Lesson 04-04

Train Neural Language Model

15 min
2 exports
4 tests

Lesson blocked by prerequisites

Complete and save a passing attempt for your active lesson before running this one.

Go to active lesson

Lesson workspace sections

Submission + Results
Not run yet

Seed 404 • Runtime includes 15 prerequisite modules
Test Results

Run tests to see case-by-case feedback.

Attempts0 saved

No saved attempts yet.

Lesson README

04-04 Train Neural Language Model

Why this matters

This lesson connects embeddings, context composition, logits, loss, and updates in one reproducible loop.

Intuition first (no jargon)

A full train/eval cycle should produce logits, scalar loss, and trackable metric history.

Worked example

  • Inputs: one batch with B = 2, T = 4, vocab size V = 20
  • Shapes:
    • tokens [B][T] -> embeddings [B][T][dModel]
    • model output logits [B][T][V]
    • loss scalar []
  • One computed step:
    • If validation loss for an epoch is 2.10, then surprise is exp(2.10) = 8.17
  • Output: report loss and optional surprise.

Code walkthrough

js
export function neuralForward(contextIds, params) {}
export function trainNeuralLM(trainSet, valSet, params, cfg) {}
  • neuralForward maps context IDs to next-token logits.
  • trainNeuralLM returns train/validation metric history.

Your task

Implement training and evaluation flow.

  • Build forward pass with embedding lookup and MLP head.
  • Train for several epochs with configurable learning rate.
  • Report train and validation loss each epoch.
  • Optionally report surprise as Math.exp(loss).

Common mistakes

  • Changed seed or split between baseline comparisons
  • Mixed context lengths across runs

Precision note

This lesson uses small dimensions so optimization behavior stays inspectable.

Hints

  • Keep seed fixed when comparing baselines.
  • Start with tiny model dimensions for fast runs.
  • Early stopping can prevent overfitting on tiny data.

Check your thinking

  1. Why can train loss drop while validation loss rises?
  2. Why is baseline comparison important?
  3. What should remain constant for fair comparison?

Stretch (optional)

Save best parameters by validation loss.

Likely test focus

  • Forward pass output dimensions.
  • Training loop returns decreasing trend on toy data.
  • Neural model beats bigram on provided validation set.

What should improve

You can now compare neural and count-based baselines under fixed evaluation controls.

Bridge to next lesson

Monaco Editor

Matches starter

Files

Editor is deferred on smaller screens to keep startup fast.

Autosave is enabled in local storage for this lesson.