Project status: private beta. This documentation is generated from
the real
haki repository code (FastAPI routes, Pydantic schemas, SDKs)
— every request/response example comes from a cited source file, never
a guess.Haki, in one sentence
Picture a brilliant colleague with amnesia: every new conversation starts from zero. That is what an AI agent lives through today. Haki is that agent’s infallible notebook: it writes down what matters, retrieves the right information at the right time, crosses out what is no longer true without ever erasing it, and can show the exact proof behind every memory it uses. That notebook belongs to your agent, whatever its model (OpenAI, Anthropic, a local model…) and whatever its body (Python code, an n8n workflow, Cursor…). You keep your stack; Haki adds the memory.Three ideas
Facts, not a transcript
Haki does not archive your conversations: it extracts structured
facts from them (preferences, constraints, decisions), linked to
their proof — the exact source event that produced them.
Today's truth
Every fact has a validity date and a status. When the user changes
their mind, the old fact is superseded — never silently deleted,
never served as current.
Proof on every answer
Every injected memory packet ships with its sources, its dates, and a
trace explaining what was kept, discarded, or blocked — and why.
How it works
- Capture — your application sends an event (a message, an action, a tool result); Haki records it as immutable proof and answers in milliseconds, with no duplicate on retry.
- Consolidation — in the background, Haki extracts what deserves to become a durable fact, deduplicates, detects changes of mind (supersession) and contradictions (conflict).
- Context — before every answer, your agent asks for memory; Haki only returns facts that are active, valid, in the right scope, ranked by relevance, under a strict token budget.
- Inspect — at any time, the trace shows every memory that was kept, discarded or blocked, with the exact reason.
- Forget — correction or erasure: forgetting propagates to everything derived from it, with a timestamped receipt.
Four ways to use Haki
Coded agent — SDK + CLI
Three lines around your LLM call, in Python or TypeScript.
client.context(...) before the call, capture_turn(...) after.Cursor — MCP server + Hooks
One-click install (
haki mcp) or guaranteed capture via Cursor Hooks
(haki hooks). Four tools, one Project Rule.OpenAI-compatible gateway
Change
base_url, keep your code: memory is injected and captured
automatically on every chat.completions call.Web console
Landing page + app wired to the real API: memories, timeline, traces,
conflicts, keys. Not covered by this site — see
console/README in
the repository.Performance (measured, not promised)
Reproducible benchmark:uv run python scripts/benchmark_context.py
(100 requests per size, local embeddings, Windows dev machine). Re-measured
10 Aug 2026 — the table below replaces an earlier run whose numbers had
drifted (see the script’s own git history if you need the old figures).
This 249 ms figure (p95, 10,000 facts) is the only performance measurement
cited in this documentation outside of its exact context — do not
generalize it to a different token budget, a different data volume, or a
different machine. It holds because embeddings are computed locally
(ONNX CPU, 384-dimension multilingual model): no network call sits in the
critical path of
/v1/context.
Where to start
2-minute quickstart
Docker Compose, migrations, first API key, first
capture → context.
API reference
Every
/v1/* endpoint, exact request/response schemas, typed error
codes.
