06-02 Feed-Forward Layer
Why this matters
The Transformer applies a position-wise FFN: max(0, xW1 + b1)W2 + b2 after attention mixing.
Intuition first (no jargon)
After cross-position mixing, each position gets a deeper non-linear transform independently.
Paper grounding
- The feed-forward sublayer is
FFN(x) = max(0, xW_1 + b_1)W_2 + b_2. - The same FFN parameters are applied to each position independently.
Code walkthrough
jsexport function ffnToken(x, params) {} export function ffn(xSeq, params) {}
Your task
Implement FFN for one token and full sequence.
- Compute
xW1 + b1, apply activation. - Compute second projection back to
dModel. - Process every token independently.
Hints
- Reuse existing linear and activation helpers.
- Keep
xSeq.lengthunchanged. - Check hidden dimension shape carefully.
Check your thinking
- Why is FFN applied per token independently?
- What role does hidden expansion play?
- Why pair attention and FFN together?
Stretch (optional)
Try GELU activation and compare training smoothness later.
Likely test focus
- Correct output shape
[T][dModel]. - Deterministic transform on toy inputs.
- No cross-token mixing inside FFN.
What should improve
You can now implement the position-wise FFN stage used in Transformer blocks.
Bridge to next lesson
Next you add residual connections and normalization for stable deeper stacks.