Public benchmark
Measured on real memory tasks with two checks
Engram is measured two ways: first, whether it retrieves the right evidence; second, whether an LLM can answer correctly from that retrieved context.
We keep retrieval and answer accuracy separate. Both use public LoCoMo, but they measure different parts of the memory loop.
LoCoMo evidence retrieval
Full public run. Scores show whether the required source evidence appears in the retrieved memories.
| System | MRR | R@1 | R@50 | R@200 |
|---|---|---|---|---|
| Engram | 0.5427 | 40.95% | 93.29% | 97.59% |
| Mem0 OSS | 0.4314 | 31.25% | 85.16% | 94.79% |
| LangMem / LangGraph | 0.4006 | 28.39% | 83.98% | 94.53% |
| AgentMemory | 0.4627 | 35.68% | 83.33% | N/A |
| Session-file BM25 | 0.1858 | 9.11% | 81.90% | 95.05% |
| BM25 | 0.4676 | 36.39% | 78.19% | 85.74% |
LoCoMo answer-level judge
Same full set of 1,536 evidence-bearing tasks and top-50 cutoff for every row. Codex generated the answers and a separate blind judge scored them; this is an Engram-run diagnostic, not an official LoCoMo leaderboard.
| System | Correct / N | Top-50 answer |
|---|---|---|
| Engram | 1,377 / 1,536 | 89.65% |
| Mem0 OSS | 1,294 / 1,536 | 84.24% |
| AgentMemory | 1,280 / 1,536 | 83.33% |
| LangMem / LangGraph | 1,279 / 1,536 | 83.27% |
| Session-file BM25 | 1,235 / 1,536 | 80.40% |
| BM25 | 1,227 / 1,536 | 79.88% |
How to read the two tables
Both published evaluations use all 1,536 evidence-bearing tasks. Evidence retrieval requires source IDs, and answer accuracy uses the same answer and judge models for every included system. Systems that do not satisfy the same evaluation contract are documented separately rather than mixed into the ranking.
How to read the two tables