Skip to main content
Every number on this page comes from a real, paid run against the exact harness in eval/ — never a proxy, never cherry-picked. Reproduce any of it yourself: see eval/README.md.

Where we are today

70.0% answer accuracy — the mean of three independent 300-question stratified samples of LoCoMo (seeds 42, 1, 2; range 66.0-72.0%), measured 2026-08-22 under the same frozen mem0_calibration protocol as the entry below, now run at Haki’s actual product default of 2000 context tokens rather than the artificially tighter 900 used for that earlier measurement (see the note under the trajectory table). For the first time, that mean is at or above the number Mem0’s original paper reports on this dataset (arXiv:2504.19413, 66.9%). We’re reporting a range rather than a single flattering run because that’s what three independent seeds actually showed — on the weakest of the three (66.0%) we’re roughly at parity with that number, not clearly ahead of it. Full disclosure, because the alternative is someone else pointing it out first: Mem0 has since published a newer “token-efficient” algorithm (July 2026) reporting 92.5% on LoCoMo and 94.4% on LongMemEval — mem0.ai/blog. We have not reproduced those numbers under our own harness, and by that post’s own description they come from a different single-pass setup and reflect Mem0’s managed platform rather than its open-source SDK — so we make no claim against them either way. What we can say and stand behind is parity with the one Mem0 protocol we’ve actually reproduced end to end, including its own full-context baseline, below. What we have is a real, measured trajectory on a frozen protocol, and a harness honest enough to show its own weak spots.

The trajectory

The 2026-08-22 row is not directly comparable to the one above it on budget: the context budget changed from 900 tokens (an artificial floor kept only for comparability with the original Mem0/Zep protocol) to 2000 tokens (what the product actually ships with, since v0.2.0). Numbers moved for two reasons at once — nine retrieval fixes and a wider budget — and this table does not attempt to separate them out. We also switched from reporting one run to reporting three (seeds 42, 1, 2): a single 300-question sample varied by up to 6 points between seeds on this protocol, which a single run would have hidden. One more difference: the 2026-08-17 row ran with the cross-encoder reranker on; the 2026-08-22 row does not (off by default — see “What’s still broken” below, its measured contribution did not survive a later, more careful re-test).
For reference, this same protocol’s own full-context baseline — the entire conversation history, no memory system, no compression, the ceiling this whole approach is measured against — scores 73.7% here, within a point of Mem0’s own reported full-context baseline (72.90 +/- 0.19%). That closeness is what confirms the harness reproduces the published protocol rather than measuring something else under the same name. Haki reaches within 3.7 points of that ceiling using meaningfully less context: ~27,300 tokens for the raw transcript, an actual count of what the baseline call billed.
We are deliberately NOT publishing a precise token ratio for Haki’s own side of that comparison yet. The harness currently records Haki’s context size from its own internal budgeting estimate, not from an independent count of the literal prompt sent to the model — while the baseline number above comes straight from the API’s own billed usage. Those are not the same kind of measurement, and we caught that mismatch ourselves reviewing this page rather than measuring it away. A from-source-tokenizer number for Haki’s side is the next thing we owe this page, not a headline we’re comfortable shipping on an estimate.
The two earliest runs use different sample sizes (full 1540 vs. a 180-question stratified sample keeping LoCoMo’s real category proportions). All three are real paid measurements on frozen protocols, not a proxy — but smaller samples carry wider error bars than 1540. Treat the trajectory’s direction and magnitude as solid; treat any single endpoint as carrying real sample-to-sample noise until the next full run.

What it costs

0.238averagetotalfora300−questionstratifiedrun(threeseeds:0.238 average total for a 300-question stratified run (three seeds: 0.235, 0.240,0.240, 0.238) — 0.00067perquery,0.00067 per query, 0.0035 per ingested conversation. Full breakdown, including the per-run config (frozen dataset checksum, prompts, model prices) that makes this number reproducible: eval/results/ (gitignored in the repo — generated locally when you run the harness).

What’s still broken

Declared as tracked, labeled issues — not hidden, not silently fixed between releases:

Temporal accuracy still trails single-hop

49.5% vs. 81.5% on the current protocol — the largest same-row gap of any category, and the clearest place to look next.

The Mem0 gap closed on the mean, not on every seed

70.0% mean vs. Mem0’s published 66.9% — but our own weakest seed (66.0%) lands right on top of it, not above it.

Open-domain: sample still smallish

n=19 per seed — better than the 180-question run’s n=11, still not enough to trust in isolation. (count and the adversarial categories are out of scope for mem0_calibration by design, matching Mem0’s own published methodology.)

Reranker: re-measured, and the gain didn't survive

Re-tested Aug 21 after the ranking rewrite: 85.1% to 86.0% gold evidence served on 308 questions — McNemar p=0.68, not distinguishable from noise. The earlier, larger figure was measured on top of the pre-Aug-21 ranking and didn’t hold once that ranking was fixed. Off by default; the ~1.1s p50 / ~1.4s p95 CPU cost is known and wasn’t the reason it stayed off.

Seed-to-seed variance is real, not noise to wave away

66.0% to 72.0% across three seeds of the same 300-question protocol, mostly driven by multi-hop (49.1% to 72.7%). Read any single Haki-vs-Mem0 number, ours included, as a range until proven otherwise.
See the full list, and comment if you’ve hit something not listed: known-limitation issues.

Methodology

  • Dataset: LoCoMo, pinned by SHA-256 checksum — the same 1540-question set every run draws from, never a moving target.
  • Protocol: mem0_calibration in eval/configs/ — reproduces Mem0’s exact question-then-context ordering and judge prompt, so the resulting number is comparable to what Mem0 and Zep publish under their own conditions, not just internally consistent. Context budget is the product’s own default (2000 tokens) as of the 2026-08-22 measurement.
  • Models: gpt-4o-mini for both the answering model and the judge, temperature 0.
  • What’s measured beyond accuracy: abstention rate, contradiction leakage, context tokens per packet, p50/p95 latency, cost per query and per ingestion — metrics the field rarely publishes alongside a headline number.
  • Run it yourself: exact commands in eval/README.md. No account needed — it runs against your own self-hosted instance. This repository’s migration history has gaps in its numbering (jumps like 0026 to 0027) relative to our own internal repo: three migrations for organizations/billing (Cloud-only, never part of this OSS core) are omitted. None of the three touch a table this harness reads or writes — the retrieval and consolidation code paths under test don’t reference organizations or billing tables at all — so the gap changes migration numbers, not reproducibility.
Report generated 2026-08-22. Next revision alongside the next full-scale run.