Last updated: 2026-09-06

Agent memory persistence — Benchmark Sources & Consensus

Agent ability to retain, transfer, and recall context across sessions — measured by task success rates before and after memory handoffs between models or restarts.

Platforms tracked: Hermes · Openclaw · Nemoclaw · Claude Cowork · Chatgpt

Consensus across 3 sources

Across 3 sources, persistent-memory layers lift agent task success: VEKTOR Slipstream scores 0.894 Transfer Continuity (vs PAM 0.880); World Model MCP adds +10.2 pts on SWE-bench repeat-mistakes; a third, self-reported comparison puts a memory graph at 0.831 against 0.801 on LOCOMO. All three margins are narrow and two are reported by the system's own author, so treat the direction as better supported than the ordering.

All Sources

We aggregate published benchmarks; we never run our own tests and never pick winners. Each row links back to the original publication.

SourceDateFindingMethodologyQuality
Medium / Vektor Memory 2026-05-31 VEKTOR Slipstream scores 0.894 Transfer Continuity Score vs Microsoft PAM's 0.880 across 50 engineering scenarios; memory lift ratio 6.61x vs 2.51x. Transfer Continuity Score (task success with vs without memory transfer); 50 engineering scenarios across Q&A, coding, planning; GPT-4 Turbo baseline high
Hacker News / GitHub 2026-06-24 World Model MCP, a harness-neutral memory layer, cut repeat coding-agent mistakes by +10.2 pts paired delta on 49 SWE-bench instances. Pre-registered SWE-bench Verified repeat-mistake test; 49 instances; paired high
GitHub 2026-08-12 A memory-graph implementation scored 0.831 against 0.801 for Memora on the LOCOMO long-conversation memory benchmark LOCOMO benchmark run by the author against one named competitor; scores published with the harness in the repository medium

How we work

OpenClawDatabase aggregates and links to published benchmarks. We don't run our own tests, and we don't pick winners. Our weekly benchmark-aggregator routine scans 7+ live leaderboards (OpenRouter, Aider, SWE-bench, GAIA, LMSYS, BigCodeBench, MMLU-Pro) plus relevant Reddit and Hacker News threads, then writes structured entries into /assets/benchmarks.json. Every row here links back to the original publication.

← Back to all benchmark tasks · See also: Decision guide · Cost calculator

📬 Weekly Digest — In Your Inbox

One email a week: top news, releases, and our deepest new guide. No spam. Same content via RSS if you prefer.