The trade-off, in the open.
Give an agent the whole conversation and it answers well — but it re-reads, and re-pays for, everything on every question, and eventually the transcript no longer fits. A memory layer reads a small fraction of that. The real question is how much answer quality you give up to get there. Here is Memomee's answer, measured against a reader that sees everything and against the open-source alternatives — the same test for every system, and every number from a real measurement.
Memomee is a memory governance and retrieval control plane for long-running agents — the kind that write code, run research, or drive workflows across hours and sessions. If yours run long enough to branch, get interrupted, or drift, this page is the evidence. If you only need short-session chat history, you don't need it.
The decision, in one screen.
Five ways to give an agent memory, on one fair test — same benchmarks, same answer model, same grader, same questions; only the memory strategy changes. If you're evaluating alternatives, start here; the evidence follows.
What we're comparing · and why these five
Paste the whole conversation into the model on every question.
Best for: Short chats that still fit the window.
Embed raw turns; retrieve the top-k by similarity.
Best for: Stateless lookup over documents.
An open-source memory system with its own extract-and-store pipeline.
Best for: A drop-in OSS chat-memory layer.
An open-source graph memory system: it builds its own graph over the conversation and runs its own model stack.
Best for: Graph-shaped recall over a corpus, when read cost is not the constraint.
Typed derive → govern → retrieve → assemble — a control plane, not a store.
Best for: Long-running agents that branch, drift, or need governance.
Two bookends — the full-context ceiling and the RAG floor — plus two open-source memory systems, so every comparison is against a meaningful reference, not a strawman. Weighing all six memory approaches, not just these five? Choosing memory for AI agents walks the full landscape.
Which should you reach for?
| If you're building… | Reach for |
|---|---|
| Conversations that fit a context window | Full context |
| One-shot lookup over documents | RAG |
| Agents running for hours or across sessions | Memomee |
| Governance — approval, audit, forgetting, isolation | Memomee |
| Lowest absolute token cost | RAG or mem0 — memomee reads more |
| Long conversations without reading the whole transcript | Memomee |
Memomee at a glance
Four findings
- 01Near-parity recall
Finds the evidence an answer needs about as often as a reader that sees everything — recall@k 0.98–1.00.
- 022.9–10.8× fewer tokens
Reads a compact ~0.2–2.4k-token context, not the whole (and growing) transcript.
- 03Governance nobody else publishes
Approval, audit, supersession, forgetting, zero workspace leaks — re-verified on every pull request and every push to main.
- 04Honest trade-offs
A full-context reader scores 0.041 higher than memomee on LoCoMo and scores above memomee on seven of the nine sweep venues; we show the gap rather than hide it.
The rest of this page is the evidence behind those four — head-to-head, the governance, and where each system wins; the eleven-benchmark sweep and methodology live on /benchmarks. Jump to the recommendation for the call.
Three ideas, then the numbers explain themselves.
The benchmarks below come from the memory-research literature. Three plain-language ideas make every table on this page readable.
The accuracy ceiling: hand the model the entire conversation on every question. It answers well — but it re-reads (and re-pays for) everything, and it breaks the moment the conversation outgrows the context window. It's the bar to measure against, not something you'd ship.
Did retrieval find the right evidence? The share of questions where the fact an answer needs actually reached the model. When recall is high but the answer score is lower, the gap is in writing the answer — not in the memory.
Every system here is answered and graded by the same AI, over the same questions, in a single run. Only the memory strategy changes — so the comparison is fair. That is exactly what public leaderboards usually are not.
Most of the answer, a fraction of the tokens.
What matters in production is answer quality set against how much text you read to get it. The full-context reader sits at the top of the chart because it reads everything. Memomee is level with the full-context reader to three decimals while reading a small fraction of the tokens: most of the answer, for a fraction of the context.
What the chart shows: Memomee is level with the full-context reader to three decimals on LongMemEval and trails the full-context reader by 0.041 on LoCoMo, while reading 2.9–10.8× fewer tokens per question. Reading few tokens isn't the whole story on its own — the RAG baseline on LoCoMo reads almost nothing (194 tokens) and still scores higher than memomee there, while cognee takes the top score on that corpus for 19,927 tokens. What you want is the top-left corner: high answer quality and low token cost.
- LongMemEval — Memomee scores 0.745 reading 2,321 tokens: the full-context reader's answer quality to three decimals for 2.9× fewer tokens, and +0.149 over the next-best system (RAG and cognee tie at 0.596).
- LoCoMo — 0.481 reading 2,419 tokens: 92% of the full-context reader's answer quality for 10.8× fewer tokens. It trails cognee by 0.074 and RAG by 0.011 and leads mem0 by 0.124. cognee scores highest here at 0.555, reading 19,927 tokens to do it.
Where memomee lands
answer score · same test, 5 systemsmemomee is level with the full-context reader to three decimals (Δ +0.000) while reading 2.9× fewer tokens, and leads every other system by 0.149.
cognee leads here (0.555) — the first run in which an external memory product beats memomee on this corpus. It reads 19,927 tokens to do it, against memomee's 2,419. In the same run, RAG also scores higher (0.492) and memomee leads mem0 by 0.124.
Bars scale to a 0.80 score ceiling · full-context reads the whole transcript · mem0 shown lighter to distinguish it from the RAG baseline.
RAG baseline — a deliberately plain vector-retrieval RAG, the "just embeddings" floor. It embeds every raw conversation turn with the same Voyage embedder Memomee's own vector search uses, then for each question retrieves the top 5 turns by cosine similarity and hands them to the same answer model. No extraction, no summarization, no query planning, no graph, no fact store — chunk-and-embed over turns. Because it shares Memomee's embedder and the held-constant answer-writer and grader, the gap to it isolates what Memomee's derive-and-plan pipeline adds on top of plain semantic search.
mem0 (open source) — the mem0 memory system on its own native pipeline (gpt-4o-mini extraction + its own embedder). A full system-vs-system comparison; its memories aren't turn-tagged, so recall@k isn't computed for it.
LongMemEval
1 run · 47 of 50 scored| System | Answer score | recall@k | Read tokens |
|---|---|---|---|
| Memomee | 0.745 | 0.98 | 2,321 |
| RAG baseline | 0.596 | 0.86 | 1,110 |
| mem0 (open source) | 0.553 | — | 432 |
| cognee (open source) | 0.596 | — | 5,262 |
| Full-context · reads everything | 0.745 | 1.00 | 6,629 |
Memomee vs. the strongest other system +0.149 · vs. the full-context reader +0.000 · reads ≈ 2.9× fewer tokens.
LoCoMo
1 run · 44 of 50 scored| System | Answer score | recall@k | Read tokens |
|---|---|---|---|
| Memomee | 0.481 | 1.00 | 2,419 |
| RAG baseline | 0.492 | 0.87 | 194 |
| mem0 (open source) | 0.357 | — | 574 |
| cognee (open source) | 0.555 | — | 19,927 |
| Full-context · reads everything | 0.522 | 1.00 | 26,154 |
Memomee vs. the strongest other system −0.074 · vs. the full-context reader −0.041 · reads ≈ 10.8× fewer tokens.
What the numbers do and don't say.
On LongMemEval memomee is level with the full-context reader to three decimals, and leads RAG by 0.149, cognee by 0.149 and mem0 by 0.191. On LoCoMo it trails cognee by 0.074 and RAG by 0.011 and leads mem0 by 0.124. This is a single run with no variance estimate, so read one corpus against the other rather than either as settled.
A reader handed the entire transcript is level with memomee to three decimals on LongMemEval and scores 0.041 higher than memomee on LoCoMo, and scores above memomee on seven of the nine sweep venues. We show that gap rather than hide it. Most of it comes down to how memory is written — Memomee distills each conversation into typed facts with an LLM before the reader ever sees them, and the quality of that extraction is what narrows the gap.
Whatever the answer score, retrieval finds the right evidence about as reliably as the full-context reader — recall@k of 0.98–1.00 — while reading a compact assembled context: 2.9–10.8× fewer tokens on the head-to-head tests. That token discount is the point.
The same held-constant setup for all five systems: DeepSeek writes and grades every answer, over a single run of 50 sampled questions per benchmark, 47 scored on LongMemEval and 44 scored on LoCoMo. Memomee and the RAG baseline share one semantic-search embedder; mem0 and cognee each run their own native pipeline — a system-vs-system comparison, not a component swap. Compare differences within a test, not raw scores across different tests.
Everyone claims to win LoCoMo.
Almost none of those numbers can be compared. Swap the AI that writes the answers, or the one that grades them, and the score moves more than the memory system ever does. So these are vendor claims, each tagged with the setup it came from — not results we can honestly place next to ours.
| System | Published claim | Benchmark / setup | Why it isn't comparable |
|---|---|---|---|
| Mem0 | 66% · 68% (graph) | LoCoMo (categories 1–4) | gpt-4o-mini answer model, gpt-4o grader; no standard LoCoMo grader |
| Mem0 | 92.5% | LoCoMo · gpt-5 | different answer model; actively contested (Mem0 ↔ Zep) |
| Zep | 63.8% · 71.2% | LongMemEval | gpt-4o-mini / gpt-4o answer model — the answer model differs |
| Letta / MemGPT | 74% | LoCoMo | non-standard grader; session-vs-turn ingestion |
| Supermemory | 81.6% · 84.6% | LongMemEval · gpt-4o / gpt-5 | different answer model (81.6% gpt-4o, 84.6% gpt-5); 84.6% disputed |
| Memobase | 75.8% | LoCoMo | unstandardized grader; disputed |
| LangMem | 51.2% | LoCoMo | third-party run, different setup |
This is why we only compare inside our own setup — the same answer-writer, grader, and search for every system. We never line our score up against a number in this table: a different answer model, a home-grown LoCoMo grader (there is no standard one), or session-by-session vs. turn-by-turn ingestion each move the result more than the memory layer does. Several of these figures are disputed by other vendors, too.
Eleven benchmarks, not two.
The two head-to-heads above aren't the whole story. Across the wider sweep — six long-conversation stress tests, one of them a million-token tier, and three multi-hop question sets — retrieval keeps finding the right evidence (recall@k near 1.0), Memomee scores below the full-context reader on seven of the nine sweep venues, always at a compact ~0.2–2.4k-token read-context. The full tables, the MemoryAgentBench MCC placement, and the methodology live on the benchmarks page — the research paper to this page's decision brief.
From a firehose of events to a small, governed context.
The numbers above come out of one pipeline. It turns everything an agent says and does into compact, typed, governed memory — then reads back only the few facts a question needs. Two things fall out of that design: the token discount you saw in the charts, and governance the store-and-search alternatives can't offer.
Event ledger
Every message, tool call, and result is appended to an immutable log — the raw truth, never edited in place. Corrections arrive as new events, so history stays reconstructable.
Derive
An LLM distills each event into typed memories — facts, decisions, resources, procedures, preferences — instead of keeping raw turns. This is what makes the context small: you read a few distilled facts, not the whole conversation.
Govern
Behaviour-changing memories — procedures, preferences, learned lessons — are born pending and cannot shape an answer until they're explicitly approved and audited. Facts stay freely retrievable; the rules that change how an agent acts are gated. This is the control plane.
Retrieve
For each query a hybrid fan-out searches every substrate at once — lexical, vector, graph, and the typed stores — to gather high-recall candidates. Recall first; precision comes next.
Assemble
Scope, freshness, supersession, and contradictions are resolved before ranking; a reranker and a conflict-collapse step produce one compact, typed, provenance-backed context block. This is where the token discount is made — a tight context, not the transcript.
Answer
The agent answers from that compact block, and a retrieval trace records exactly what was read and what was filtered — so every answer is auditable back to its evidence.
Steps 02 and 05 do the work: the agent reads a compact, roughly constant ~0.2–2.4k-token block of distilled facts instead of a transcript that grows without bound. That's the 2.9–10.8× token discount — and why memomee still answers when the conversation no longer fits a context window at all.
Step 03 is the part nobody else publishes: a memory that would change how the agent behaves is captured but inert until it's approved, every answer carries a provenance trace, superseded facts leave a trail, and one workspace's memory can never surface in another's.
Want the real diagram, stage by stage? How it works walks the full pipeline — every box clickable.
The real difference is governance.
Accurate answers at a token discount are table stakes. What actually sets Memomee apart is what happens after retrieval — the rules that decide which memory is allowed to shape an answer, and the checks that prove those rules hold: behavior-changing memory born pending until approved, facts superseded with lineage, abstention on conflict, zero workspace leaks, governed forgetting. These aren't third-party benchmarks; they're guarantees, re-verified automatically on every pull request and every push to main, with zero tolerance for a regression. Most memory products don't publish them because they don't have them.
What a control plane does that a store doesn't.
Those guarantees exist because Memomee is a control plane, not a store. Storing and searching memory is the easy part; deciding what should survive, change, be forgotten, or shape the next answer is the part the alternatives skip.
| Memomee | Managed memory APIs | Agent frameworks | Graph engines | DIY stacks | |
|---|---|---|---|---|---|
| Resumable working state | yes | no | partial | no | partial |
| Append-only event ledger | yes | no | no | no | partial |
| Fact supersession & lineage | yes | no | no | partial | no |
| Contradiction warnings | yes | no | no | no | no |
| Retrieval trace + filtered counts | yes | no | no | no | no |
| Workflow continuity, not store | yes | no | partial | no | no |
| Framework-agnostic | yes | yes | no | yes | yes |
Examples per category — managed memory APIs: Cloudflare Agent Memory, Mem0 · agent frameworks: Letta, LangGraph + LangMem · graph engines: Zep, Graphiti · DIY: vector DB + summary loop.
Memomee isn't for everyone.
The honest close to a fair comparison: the right choice depends on what you're building. Here is when each approach wins.
…if your conversations comfortably fit the model's window and re-reading everything isn't a cost problem.
…if you need stateless retrieval over documents — not evolving, per-user, governed memory.
…if your agents run long enough to branch, get interrupted, or contradict themselves — and you need to govern what memory is allowed to shape an answer.
If you picked Memomee, it earns its place on long-running agents by reading a fraction of the tokens at near-parity recall and by governing what memory is allowed to shape an answer — the part the alternatives skip. If you picked full context or RAG, that's the right call for your workload, and this page did its job.
What we don't claim
- The best raw recall. Memomee finds the right evidence about as often as a reader that sees everything — it does not out-retrieve a system that reads the whole transcript.
- A broad win on answer quality. Memomee is level with the full-context reader to three decimals on LongMemEval and trails the full-context reader by 0.041 on LoCoMo, and across the eleven-benchmark sweep memomee scores below it on seven of the nine sweep venues (PersonaMem, BEAM are the exceptions — where a full transcript no longer fits), at a large token discount. Nor a clean sweep of the alternatives: cognee leads here (0.555) — the first run in which an external memory product beats memomee on this corpus. It reads 19,927 tokens to do it, against memomee's 2,419. In the same run, RAG also scores higher (0.492) and memomee leads mem0 by 0.124.
- That memory broadly beats full context. Our own stress tests disprove that.
- What we do claim: parity on finding the right evidence, a lead over the strongest other system on LongMemEval (+0.149) in a fair test, a 2.9–10.8× token discount on long conversations, and governance the alternatives don't offer — knowing when to abstain, replacing facts with a trail, and never leaking one workspace's memory into another's.
Methodology: same-setup head-to-head, measured 2026-08-11 and published from eval-results/e2e/competitor/2026-08-11/ — a single run (seed 11), 50 sampled questions (47 scored on LongMemEval, 44 scored on LoCoMo), the same answer-writer and grader across all five systems, turn-by-turn ingestion (Memomee and the RAG baseline share one Voyage embedder; mem0 and cognee each run their own native stack). Single-run means no variance estimate: read these as one measurement, not a mean. Recall here is measured at the evidence-session level (did the right session reach the model); a stricter turn-level measure would be a different quantity, and no committed run reports one — so none is quoted here. The 2.9–10.8× token range is the two head-to-heads (≈ 2.9× LongMemEval, ≈ 10.8× LoCoMo fewer read-tokens than the full-context reader); across the wider benchmark set memomee assembles a compact, roughly constant ~0.2–2.4k-token read-context per question regardless of transcript length. Every "read tokens" figure on this page is the assembled read-context (read_tokens), one metric page-wide. Every delta here is a within-run difference only — it compares arms inside this one run and says nothing about either system across runs. Deltas are computed from unrounded arm scores, so they can differ by 0.001 from subtracting the rounded figures shown. The full sweep, variance, and evidence tiers: /benchmarks. Full documentation is coming to docs.memomee.ai.