Technical Research & Benchmarks

Blogs

Practical empirical studies, CPU optimization techniques, and architecture breakdowns for running lightweight agents on commodity hardware.

Controlled Benchmark 100% CPU RAG Edge AI
August 2026 • 8 min read

Context Distillation for CPU RAG: A Controlled Study on Sub-3B Models

There is a pattern almost every local RAG tutorial follows: chop your documents into 500-token pieces, pull the top 5 to 8 matches from a vector database, and dump all of it — 2,000 to 3,000 tokens — into the prompt. When run on sub-3B models on CPU, prefill latency explodes and hallucination rates surge. We ran a controlled 120-query ablation to prove why context purity beats model size.

95.8%
Grounded Accuracy (±1.8% SE)
1.15s
Total Latency p50 (1.23s p95)
0.10s
TTFT p50 (75% Faster)
3,049 MB
Peak Memory Footprint
SP
Suryaprakash C V
Creator, SLMAgents • Published on Medium