Main takeaways
Traditional ML Metrics:
- Accuracy is not enough: dangerous with class imbalance
- Precision/recall trade-off: choose by false positive vs false negative costs
- Use cross-validation and keep a truly unseen test set
LLM Evaluation:
- Perplexity β truth: fluent text can still be wrong
- LLM-as-a-Judge and G-Eval: scalable evaluation with rubrics
- RAG: ground answers in documents to cut hallucinations
- Red teaming: find vulnerabilities before users do
- Benchmarks have limits: Goodhartβs Law and βbenchmaxxingβ
- Subgroup analysis: always look beyond averages
















