Agentic Memory Index Benchmark: Simple Markdown Wiki Outperforms Commercial AI Agent Memory Solutions
Key Takeaways
- ▸A simple Markdown wiki achieved 98.5% accuracy, outperforming every commercial AI memory product tested, challenging assumptions about vendor-provided solutions
- ▸Significant performance variation by task type: no single tool wins across accuracy, cost, and speed—trade-offs dominate selection decisions
- ▸Cost per successful answer ranges from $341 (Mem0) to $569 (Karpathy Wiki); Anthropic Memory ranks mid-pack at $419.45 with 77.1% accuracy
Summary
Verging Labs released the Agentic Memory Index v0.1, an independent benchmark measuring accuracy, cost, and speed of AI agent memory tools. The striking headline finding: a simple Markdown wiki (Karpathy Wiki) achieved 98.5% accuracy—outperforming all commercial competitors, including Anthropic Memory (77.1%), Mem0 (92.3%), Supermemory (75.1%), and Zep (75.1%). The benchmark evaluated nine memory tools and three search solutions, testing across multiple dimensions including task-specific performance, data retention over extended sessions, and operational cost-efficiency.
The benchmark reveals significant performance trade-offs. Karpathy Wiki leads on accuracy (98.5%) and speed (2.7s response time) but carries the highest cost at $568.93 per 1,000 successful answers. Conversely, Mem0 offers the lowest cost ($341.42 per 1,000 answers) with solid 92.3% accuracy, while Anthropic Memory achieves 77.1% accuracy at $419.45 per 1,000 answers. Claude Code's built-in memory ranked lowest at 67.7% accuracy. Performance also varies significantly by task type, with each tool showing distinct strengths in direct recall versus synthesis-based reasoning.
Verging Labs provides live access to benchmark data via API ($0.035 USDC per query) with a standard subscription API planned. The findings challenge assumptions about commercial AI tool complexity and suggest that simple, well-designed solutions can outperform expensive platforms—though cost and latency considerations remain critical for production deployments.
- Data retention quality varies significantly across tools—early-stage facts don't survive equally well across sessions, a critical factor for long-running agents
- Claude Code's built-in memory ranked lowest at 67.7% accuracy, suggesting external memory solutions may outperform integrated approaches
Editorial Opinion
This benchmark delivers a sobering message to the commercial AI memory market: a simple Markdown wiki outperforms purpose-built systems designed specifically for AI agents. The result suggests many vendors are over-engineered or optimizing for the wrong metrics, prioritizing feature complexity over retrieval accuracy. However, the nuance matters—cost, latency, and task-specific strengths create real trade-offs that justify some commercial solutions depending on use case requirements.



