Public benchmark

Engram benchmark methodology

See the public datasets, retrieval metrics, answer checks, baselines, limitations, and claim rules behind Engram's published results.

LoCoMo evidence retrieval92.19%
LoCoMo answer-level judge66.33%
How It Works
01

LoCoMo evidence retrieval

Full public run. Scores show whether the required source evidence appears in the retrieved memories.

02

LoCoMo answer-level judge

Codex/self-reviewed diagnostic on the same 98 tasks. Scores show whether the generated answer is correct from retrieved context.

03

How to read R@200 and Top-200

Evidence retrieval is cumulative: R@200 can include everything R@50 found plus more, so it should not be lower. Answer accuracy is different. More retrieved context can also add noise or conflicting memories, so Top-50 may answer cleaner while Top-200 tests whether packaging stays useful with more evidence.

1,536

scored evidence-retrieval questions

92.19%

Engram R@50

66.33%

answer accuracy from top-50 context

We keep retrieval and answer accuracy separate. Both use public LoCoMo, but they measure different parts of the memory loop.

LoCoMo ↗

Engram is measured two ways: first, whether it retrieves the right evidence; second, whether an LLM can answer correctly from that retrieved context.

We keep retrieval and answer accuracy separate. Both use public LoCoMo, but they measure different parts of the memory loop.

Support Loop

Help us make Engram sharper

If setup is confusing, a tool behaves strangely, or a pricing limit feels wrong, send it here. The message lands directly with us.

Feedback

Tell us what is missing

Bugs, confusing setup steps, pricing questions, and product ideas go straight to the Engram team.

10 more characters needed0/10