Open Documentation
Agentic AI
& Post-Training Cookbooks
A growing repository of utility scripts, bug fixes, tutorials, and architectural frameworks to support the community navigating model fine-tuning and agentic deployments. The cookbooks cover LoRA and QLoRA, DPO, knowledge distillation, the KV cache, speculative decoding, tokenizers, MCP servers, self-hosting on AWS, and a multi-part agentic market simulator, each with code you can run.

RAG vs SFT: When to Use Which
A technical breakdown of when to use Retrieval-Augmented Generation (knowledge) versus Supervised Fine-Tuning (behavior) in enterprise AI pipelines.
Agentic Market Simulator, Part 1: The Engine and the Environment
Part 1 of a 10-part series simulating a stock market traded by language-model agents. Build the call-auction clearing engine and benchmark it against zero-intelligence controls before any model is involved.
Agentic Market Simulator, Part 2: Universe, Information and Leakage
Part 2 of a 10-part series simulating a stock market traded by language-model agents. Move from a synthetic fundamental to real S&P 500 data, measure lookahead bias with a fabricated-financials experiment, and build a re-identification probe to test anonymization.
Agentic Market Simulator, Part 3: Actions, Personas and Track Record
Part 3 of a 10-part series simulating a stock market traded by language-model agents. Build a calibrated attention layer that decides who trades and why, grounded in the no-trade theorem, before a single language model call is made.
Using LLMs for Financial Applications: Bias & Considerations
A running reference on bias in LLM-based financial judgment: model memorization of real companies from raw financials, and how framing, naming, and batching change scores even when the underlying data is held fixed.
What Are RL Environments? Rubrics, Verifiable Rewards, and Scaling
An introduction to RL environments: how agents observe, act, and receive rewards, the shift from human grading to verifiable rewards (RLVR) and rubrics, and scaling agent training.
What Is Grokking in Machine Learning? Delayed Generalization Explained
Grokking is when a neural network suddenly generalizes long after memorizing its training data. A clear definition, why it happens, and code to reproduce it.
Self-Hosting an Open-Weight Coding Agent on AWS
Run a private coding agent on your own AWS account: instance choice, setup, costs and security. No per-token API bills, and your code stays in your VPC.
DPO Loss Function Explained: The Math of Direct Preference Optimization
Step-by-step derivation of the DPO loss and what the β (KL) term controls. Includes the formula, the gradient, and a minimal NumPy implementation.
Building Local MCP Servers for VS Code & Cursor with FastMCP
A step-by-step technical guide to building Python MCP servers with FastMCP, configuring VS Code and Cursor mcp.json, and debugging stdio JSON-RPC streams.
Speculative Decoding from Scratch: Acceptance Rule and Why It's Exact
Implement speculative decoding step by step, with the min(1, p/q) acceptance rule and the residual-distribution proof that it preserves the target model's output.
Power Analysis for Benchmark Design
Calculate statistical power and minimum detectable accuracy gaps for LLM benchmarks using Python to avoid reporting sampling noise as signal.
Deduplication at Scale: MinHash & LSH in Python
Implement MinHash and Locality-Sensitive Hashing (LSH) in pure Python to eliminate the quadratic bottleneck of near-duplicate detection for ML dataset curation.
Fixing EOS Errors: Why Fine-Tuned Models Talk to Themselves
A technical guide to fixing the infinite generation bug in SFT by properly mapping EOS tokens and chat templates in Hugging Face.
What is the Model Context Protocol (MCP)?
A technical architecture guide to the Model Context Protocol (MCP). Learn how it solves the M x N integration problem for AI agents and tool calling.
PEFT Explained: LoRA vs. QLoRA
Understand the architectural differences between LoRA and QLoRA, and learn when to use each Parameter-Efficient Fine-Tuning technique based on your VRAM limits.
Train Custom Tokens with LoRA: resize_token_embeddings and trainable_token_indices
Add new special tokens to a model and train them with LoRA. Covers resize_token_embeddings, modules_to_save, and PEFT trainable_token_indices, with code.
Catastrophic Forgetting in Fine-Tuning
What catastrophic forgetting is, why it happens when you fine-tune an LLM, what it costs you, and how to spot it, plus when a specialized small model can safely ignore it.
Inspecting What a Tiny Transformer Actually Learned
A technical guide to probing a character-level PyTorch Transformer. Learn how to measure rule acquisition, test generalization, and ablate attention heads.
How to Check a Fine-Tuning Dataset Before You Train
Avoid common fine-tuning failures. Learn how to validate chat templates, prevent silent truncation, verify loss masking, and check for data leakage before spending GPU hours.
LoRA from Scratch in PyTorch: Implementation with Code
Implement LoRA (low-rank adaptation) in pure PyTorch: the low-rank A and B matrices, initialization, merging weights, and a training example.
Fix "element 0 of tensors does not require grad" (LoRA + Gradient Checkpointing)
LoRA training crashes with gradient_checkpointing_enable()? The fix: enable_input_require_grads() and use_reentrant=False. Copy-paste code for PEFT and Transformers.
Knowledge Distillation in PyTorch: Teacher-Student Tutorial with Soft Targets
A from-scratch PyTorch knowledge distillation tutorial: teacher and student logits, temperature, soft-target KL loss, and a full training loop.
KV Cache from Scratch in PyTorch: Implementation and Speedup
Implement a KV cache for transformer inference in pure PyTorch, step by step. See how caching keys and values cuts generation time, with runnable code.
Tokenizers From Scratch: BPE vs MaxMatch
Stop treating tokenization as a black box. Learn how Byte Pair Encoding (BPE) and MaxMatch split text differently, and what mathematically happens when you add new tokens for fine-tuning.