Open Documentation

Model Architecture & Internals Cookbooks

How transformers actually work, built from scratch rather than taken on faith. Tokenizers, the KV cache, speculative decoding, and what a small model has genuinely learned once training finishes.

In this collection: RAG vs SFT: When to Use Which; What Is Grokking in Machine Learning? Delayed Generalization Explained; Speculative Decoding from Scratch: Acceptance Rule and Why It's Exact; Inspecting What a Tiny Transformer Actually Learned; KV Cache from Scratch in PyTorch: Implementation and Speedup; Tokenizers From Scratch: BPE vs MaxMatch.

Model Architecture & Internals

RAG vs SFT: When to Use Which

A technical breakdown of when to use Retrieval-Augmented Generation (knowledge) versus Supervised Fine-Tuning (behavior) in enterprise AI pipelines.

RAGSFTSystem Design
Model Architecture & Internals·Deep Dive

What Is Grokking in Machine Learning? Delayed Generalization Explained

Grokking is when a neural network suddenly generalizes long after memorizing its training data. A clear definition, why it happens, and code to reproduce it.

Deep LearningGrokkingPyTorchTheoryMechanistic Interpretability
Model Architecture & Internals·Deep Dive

Speculative Decoding from Scratch: Acceptance Rule and Why It's Exact

Implement speculative decoding step by step, with the min(1, p/q) acceptance rule and the residual-distribution proof that it preserves the target model's output.

PythonNumPyLLM InferenceSpeculative DecodingAlgorithms
Model Architecture & Internals·Deep Dive

Inspecting What a Tiny Transformer Actually Learned

A technical guide to probing a character-level PyTorch Transformer. Learn how to measure rule acquisition, test generalization, and ablate attention heads.

TransformersAttentionInterpretabilityPyTorch
Model Architecture & Internals·Deep Dive

KV Cache from Scratch in PyTorch: Implementation and Speedup

Implement a KV cache for transformer inference in pure PyTorch, step by step. See how caching keys and values cuts generation time, with runnable code.

PyTorchInferenceTransformers
Model Architecture & Internals·Deep Dive

Tokenizers From Scratch: BPE vs MaxMatch

Stop treating tokenization as a black box. Learn how Byte Pair Encoding (BPE) and MaxMatch split text differently, and what mathematically happens when you add new tokens for fine-tuning.

PythonTokenizationFine-TuningPreprocessing