Context Distillation for CPU RAG: A Controlled Study on Sub-3B Models
There is a pattern almost every local RAG tutorial follows: chop your documents into 500-token pieces, pull the top 5 to 8 matches from a vector database, and dump all of it — 2,000 to 3,000 tokens — into the prompt. When run on sub-3B models on CPU, prefill latency explodes and hallucination rates surge. We ran a controlled 120-query ablation to prove why context purity beats model size.