Changelog
Index releases, judge versions, weight changes, task-bank rotations, and our own corrections. This page is the record of the measurement instrument itself.
Index releases
First versioned release: 8 memory tools measured on simulated multi-week company brain operations, with the Claude Code built-in memory reference measured on the same bank. Method detail on the methodology page and the full results on the Agentic Memory Index page.
First versioned release: 9 providers on a 121-task bank, 3 repeats each. Judge gemini-2.5-pro (pinned), 94.1% agreement with human consensus (kappa 0.85, n=34) against labels from at least two humans; the human ceiling is kappa 0.917. Full calibration detail on the methodology page.
Self-correction ledger
Our own errors get logged here in the open: scoring bugs, judge regressions, retracted results, and definition mistakes, each with what changed and why.
Task review: one lookup question's correct answer changed within its source day and reputable sources split, so both answers are now accepted. 14 runs re-graded to passes across 5 providers; all affected numbers updated. Top-5 ordering unchanged.
Onboarding re-measured under a hardened isolation protocol, after we found earlier runs could be contaminated by our own test environment (our side, not the tools').
The time-to-ready column was inconsistent: Zep's 162.7 s was a mean mislabeled as a median, and Hyperspell's 0.454 s came from the wrong run. Every value is now the same statistic: the median wait, on the published run, for a genuinely new fact to become searchable. Zep is now 21.3 s (a quarter of its waits ran 5.8 minutes or longer), Hyperspell 16.0 s, Mem0 2.7 s; Supermemory and Mitosis are unchanged.