Open Documentation
Post-Training Techniques Cookbooks
The techniques that actually change a model's behavior: LoRA and QLoRA, DPO, knowledge distillation, and the failure modes that show up along the way, from catastrophic forgetting to the specific bugs that appear when combining these techniques together.
In this collection: DPO Loss Function Explained: The Math of Direct Preference Optimization; Fixing EOS Errors: Why Fine-Tuned Models Talk to Themselves; PEFT Explained: LoRA vs. QLoRA; Train Custom Tokens with LoRA: resize_token_embeddings and trainable_token_indices; Catastrophic Forgetting in Fine-Tuning; LoRA from Scratch in PyTorch: Implementation with Code; Fix "element 0 of tensors does not require grad" (LoRA + Gradient Checkpointing); Knowledge Distillation in PyTorch: Teacher-Student Tutorial with Soft Targets.
DPO Loss Function Explained: The Math of Direct Preference Optimization
Step-by-step derivation of the DPO loss and what the β (KL) term controls. Includes the formula, the gradient, and a minimal NumPy implementation.
Fixing EOS Errors: Why Fine-Tuned Models Talk to Themselves
A technical guide to fixing the infinite generation bug in SFT by properly mapping EOS tokens and chat templates in Hugging Face.
PEFT Explained: LoRA vs. QLoRA
Understand the architectural differences between LoRA and QLoRA, and learn when to use each Parameter-Efficient Fine-Tuning technique based on your VRAM limits.
Train Custom Tokens with LoRA: resize_token_embeddings and trainable_token_indices
Add new special tokens to a model and train them with LoRA. Covers resize_token_embeddings, modules_to_save, and PEFT trainable_token_indices, with code.
Catastrophic Forgetting in Fine-Tuning
What catastrophic forgetting is, why it happens when you fine-tune an LLM, what it costs you, and how to spot it, plus when a specialized small model can safely ignore it.
LoRA from Scratch in PyTorch: Implementation with Code
Implement LoRA (low-rank adaptation) in pure PyTorch: the low-rank A and B matrices, initialization, merging weights, and a training example.
Fix "element 0 of tensors does not require grad" (LoRA + Gradient Checkpointing)
LoRA training crashes with gradient_checkpointing_enable()? The fix: enable_input_require_grads() and use_reentrant=False. Copy-paste code for PEFT and Transformers.
Knowledge Distillation in PyTorch: Teacher-Student Tutorial with Soft Targets
A from-scratch PyTorch knowledge distillation tutorial: teacher and student logits, temperature, soft-target KL loss, and a full training loop.