04-03 MLP Language Head
Why this matters
An MLP head adds non-linearity so the predictor can model richer token interactions.
Intuition first (no jargon)
Stacked linear layers with activation can represent patterns a single linear map cannot.
Code walkthrough
jsexport function relu(x) {} export function mlpForward(h, params) {}
Your task
Implement relu and mlpForward.
- First affine transform:
h1 = hW1 + b1. - Apply ReLU elementwise.
- Second affine transform to vocab logits.
Hints
- ReLU is
Math.max(0, x). - Keep matrix dimensions documented.
- Reuse your linear helper from earlier lessons.
Check your thinking
- What pattern can ReLU model that linear cannot?
- Why can dead ReLU units happen?
- Why do we still output logits at the end?
Stretch (optional)
Add tanh as an alternate activation and compare loss trend.
Likely test focus
- Correct forward output shape.
- Correct ReLU behavior on negatives.
- Deterministic logits for fixed params.
What should improve
Your head can now model non-linear relationships in local context representations.
Bridge to next lesson
Now combine embeddings, context, and MLP into a full neural language model training loop.