03-02 One-Hot and Linear Logits
Why this matters
A linear logit layer is the simplest trainable scorer from token representation to vocabulary scores.
Intuition first (no jargon)
One-hot selects one token index, and the linear map converts that index to logits.
Code walkthrough
jsexport function oneHot(id, vocabSize) {} export function linearForward(vec, W, b) {}
Your task
Implement one-hot encoding and linear projection.
oneHotreturns a vector of lengthvocabSize.- Exactly one index should be
1. linearForwardcomputesvec * W + b.
Hints
- Validate ID bounds.
- Keep shapes explicit in comments while learning.
- Use nested loops for matrix multiply.
Check your thinking
- Why are logits not probabilities yet?
- What role does bias play?
- Why is one-hot sparse?
Stretch (optional)
Implement batched linear forward for many vectors at once.
Likely test focus
- Correct one-hot vector shape and values.
- Correct deterministic logits for fixed weights.
- Errors for invalid IDs or shape mismatch.
What should improve
You now have a parameterized forward pass suitable for gradient-based learning.
Bridge to next lesson
Next lesson: convert logits to probabilities and compute loss.