Every number on this page comes from a real, paid run against the exact
harness in
eval/ —
never a proxy, never cherry-picked. Reproduce any of it yourself: see
eval/README.md.Where we are today
70.0% answer accuracy — the mean of three independent 300-question stratified samples of LoCoMo (seeds 42, 1, 2; range 66.0-72.0%), measured 2026-08-22 under the same frozenmem0_calibration protocol as the entry below, now run at Haki’s actual
product default of 2000 context tokens rather than the artificially tighter
900 used for that earlier measurement (see the note under the trajectory
table).
For the first time, that mean is at or above the number Mem0’s original
paper reports on this dataset
(arXiv:2504.19413, 66.9%). We’re
reporting a range rather than a single flattering run because that’s what
three independent seeds actually showed — on the weakest of the three
(66.0%) we’re roughly at parity with that number, not clearly ahead of it.
Full disclosure, because the alternative is someone else pointing it out
first: Mem0 has since published a newer “token-efficient” algorithm (July
2026) reporting 92.5% on LoCoMo and 94.4% on LongMemEval —
mem0.ai/blog.
We have not reproduced those numbers under our own harness, and by that
post’s own description they come from a different single-pass setup and
reflect Mem0’s managed platform rather than its open-source SDK — so we
make no claim against them either way. What we can say and stand behind is
parity with the one Mem0 protocol we’ve actually reproduced end to end,
including its own full-context baseline, below.
What we have is a real, measured trajectory on a frozen protocol, and a
harness honest enough to show its own weak spots.
The trajectory
The 2026-08-22 row is not directly comparable to the one above it on
budget: the context budget changed from 900 tokens (an artificial floor
kept only for comparability with the original Mem0/Zep protocol) to 2000
tokens (what the product actually ships with, since
v0.2.0). Numbers
moved for two reasons at once — nine retrieval fixes and a wider budget
— and this table does not attempt to separate them out. We also switched
from reporting one run to reporting three (seeds 42, 1, 2): a single
300-question sample varied by up to 6 points between seeds on this
protocol, which a single run would have hidden. One more difference: the
2026-08-17 row ran with the cross-encoder reranker on; the 2026-08-22 row
does not (off by default — see “What’s still broken” below, its measured
contribution did not survive a later, more careful re-test).We are deliberately NOT publishing a precise token ratio for Haki’s own
side of that comparison yet. The harness currently records Haki’s context
size from its own internal budgeting estimate, not from an independent
count of the literal prompt sent to the model — while the baseline
number above comes straight from the API’s own billed usage. Those are
not the same kind of measurement, and we caught that mismatch ourselves
reviewing this page rather than measuring it away. A from-source-tokenizer
number for Haki’s side is the next thing we owe this page, not a
headline we’re comfortable shipping on an estimate.
The two earliest runs use different sample sizes (full 1540 vs. a
180-question stratified sample keeping LoCoMo’s real category
proportions). All three are real paid measurements on frozen protocols,
not a proxy — but smaller samples carry wider error bars than 1540.
Treat the trajectory’s direction and magnitude as solid; treat any single
endpoint as carrying real sample-to-sample noise until the next full run.
What it costs
0.235, 0.238) — 0.0035 per ingested conversation. Full breakdown, including the per-run config (frozen dataset checksum, prompts, model prices) that makes this number reproducible:eval/results/
(gitignored in the repo — generated locally when you run the harness).
What’s still broken
Declared as tracked, labeled issues — not hidden, not silently fixed between releases:Temporal accuracy still trails single-hop
49.5% vs. 81.5% on the current protocol — the largest same-row gap of any category, and the clearest place to look next.
The Mem0 gap closed on the mean, not on every seed
70.0% mean vs. Mem0’s published 66.9% — but our own weakest seed (66.0%) lands right on top of it, not above it.
Open-domain: sample still smallish
n=19 per seed — better than the 180-question run’s n=11, still not enough to trust in isolation. (
count and the adversarial categories are out of scope for mem0_calibration by design, matching Mem0’s own published methodology.)Reranker: re-measured, and the gain didn't survive
Re-tested Aug 21 after the ranking rewrite: 85.1% to 86.0% gold evidence served on 308 questions — McNemar p=0.68, not distinguishable from noise. The earlier, larger figure was measured on top of the pre-Aug-21 ranking and didn’t hold once that ranking was fixed. Off by default; the ~1.1s p50 / ~1.4s p95 CPU cost is known and wasn’t the reason it stayed off.
Seed-to-seed variance is real, not noise to wave away
66.0% to 72.0% across three seeds of the same 300-question protocol, mostly driven by multi-hop (49.1% to 72.7%). Read any single Haki-vs-Mem0 number, ours included, as a range until proven otherwise.
known-limitation issues.
Methodology
- Dataset: LoCoMo, pinned by SHA-256 checksum — the same 1540-question set every run draws from, never a moving target.
- Protocol:
mem0_calibrationineval/configs/— reproduces Mem0’s exact question-then-context ordering and judge prompt, so the resulting number is comparable to what Mem0 and Zep publish under their own conditions, not just internally consistent. Context budget is the product’s own default (2000 tokens) as of the 2026-08-22 measurement. - Models:
gpt-4o-minifor both the answering model and the judge, temperature 0. - What’s measured beyond accuracy: abstention rate, contradiction leakage, context tokens per packet, p50/p95 latency, cost per query and per ingestion — metrics the field rarely publishes alongside a headline number.
- Run it yourself: exact commands in
eval/README.md. No account needed — it runs against your own self-hosted instance. This repository’s migration history has gaps in its numbering (jumps like0026to0027) relative to our own internal repo: three migrations for organizations/billing (Cloud-only, never part of this OSS core) are omitted. None of the three touch a table this harness reads or writes — the retrieval and consolidation code paths under test don’t referenceorganizationsor billing tables at all — so the gap changes migration numbers, not reproducibility.

