08-04 Final Model Comparison
Why this matters
A fair capstone comparison keeps prompt, seed, and evaluation protocol fixed across model variants.
Intuition first (no jargon)
Compare architecture changes by pairing metric deltas with concrete output samples.
Paper grounding
- The paper compares model variants with controlled ablations (for example attention heads and dimensional settings).
- Fair comparison requires fixed evaluation protocol so changes in metrics can be attributed to architecture.
Worked example
- Shapes:
- report entries list
[numModels] - each entry includes
{ modelName, metrics, sample, reflection }
- report entries list
- One computed step:
- if bigram validation loss is
2.45and MiniGPT is1.92, improvement is0.53
- if bigram validation loss is
- In the paper, Table 3 studies head count and key/value dimensions while keeping computation roughly fixed.
Code walkthrough
jsexport function compareModels(models, prompt, cfg) { // returns structured report }
Your task
Implement compareModels(models, prompt, cfg).
- Generate text from unigram, bigram, neural LM, and MiniGPT.
- Compute comparable metrics (loss or surprise).
- Return a report object ready for UI display.
- For each model, include reflection in this format:
Change: what architectural capability was added.Evidence: metric delta and one output example.Limitation: what still fails.
Common mistakes
- Different prompts or seeds across models
Precision note
Curriculum comparisons on tiny datasets can be noisy and small-sample biased. Directional evidence under fixed controls.
Hints
- Use fixed seed and same prompt for fairness.
- Keep generation length consistent.
- Include both numeric metrics and short sample strings.
Check your thinking
- Which improvements came from attention specifically?
- Why is fairness setup critical in comparison?
- What remains limited in a tiny model?
Stretch (optional)
Add an "ablation" mode that disables one component and reruns comparison.
Likely test focus
- Report schema correctness.
- Presence of all model variants.
- Reproducible comparison under fixed seed.
What should improve
You can now produce an evidence-backed final comparison report across model versions.
Bridge to next lesson
Curriculum complete: baseline counting through transformer-style generation.