Open Documentation
Data & Evaluation Cookbooks
Before a fine-tune can work, the data has to be right. This section covers deduplication, dataset validation, and the statistics behind trusting a benchmark result, some checks to run before training.
In this collection: Power Analysis for Benchmark Design; Deduplication at Scale: MinHash & LSH in Python; How to Check a Fine-Tuning Dataset Before You Train.
Power Analysis for Benchmark Design
Calculate statistical power and minimum detectable accuracy gaps for LLM benchmarks using Python to avoid reporting sampling noise as signal.
Deduplication at Scale: MinHash & LSH in Python
Implement MinHash and Locality-Sensitive Hashing (LSH) in pure Python to eliminate the quadratic bottleneck of near-duplicate detection for ML dataset curation.
How to Check a Fine-Tuning Dataset Before You Train
Avoid common fine-tuning failures. Learn how to validate chat templates, prevent silent truncation, verify loss masking, and check for data leakage before spending GPU hours.