Skip to content

Membase· Research

Benchmarking
long-term memory
for agents.

Measured on LoCoMo, LongMemEval and DMR. Powered by episodic extraction and multi-round retrieval that sends the reader a few thousand tokens instead of the whole history.

LoCoMo
93.1
LongMemEval
92.6
DMR
92.2

Mean tokens per retrieval call

  • Membase engine6,562
  • Full-context19,987

3.05× fewer tokens on LoCoMo. On LongMemEval the gap is 11.46×: ~9,000 tokens against a ~103,000-token history.

Benchmark deep-dives

Accuracy and mean context tokens per question on each benchmark. Single pass over the full question set, graded by the benchmark's own judge.

LoCoMo

1,540 questions, 4 categories.
Single-hop, multi-hop, open-domain and temporal recall across multi-session conversations spanning months.

93.1%

Accuracy

6,562

Mean context tokens are 3.05× below the full history.

Gold session reached the reader for 96–99% of questions. Swapping the reader model moves the score by less than 0.1 points.

Accuracy by category

OverallCategory
  • Single-hopn=84194.6%
  • Multi-hopn=28293.6%
  • Temporaln=32191.6%
  • Open-domainn=9683.3%
  • Overalln=1,54093.1%

Methodology

Why the numbers look this way

Each score traces back to a specific part of the architecture, not just asserted.

Recall correctness

Every session is narrated into timestamped episodes that keep who, what and when together, so a fact is retrieved with its context. The retrieval decider reads the first hits and asks follow-up questions before settling, which is why the gold session is in context for 99.95% of LongMemEval questions.

Context footprint

Keyword and vector search are fused by reciprocal rank and only the top twenty episodes enter the prompt. That keeps a LongMemEval call at about 9,000 tokens against a ~103k-token history, and a LoCoMo call at about 6,500 against ~20k.

Response time

Search runs in 1.1–2.5 s at the median, including one to three decider rounds. End-to-end median is 3–15 s depending on the reader model; the memory layer is not the bottleneck.

Performance

Latency and tokens

Search and end-to-end timings, plus how much context each call actually sends to the reader.

LoCoMoLongMemEvalDMR
search latency, p50 / p951.67 s / 7.02 s2.53 s / 6.11 s1.13 s / 1.71 s
end-to-end, p50 / p958.30 s / 18.0 s14.7 s / 30.2 s3.21 s / 6.34 s
context tokens per question6,5628,9701,602
full history per question~20k~103k—
token reduction3.05×11.46×—
reader modelgpt-4.1-minigpt-5.5gpt-4o-mini

ARCHITECTURE

What's inside Membase

Four pieces working together, from how a session is cut up to where the memories live.

  1. 01 · SEGMENT

    Boundary detection

    An LLM pass splits each session into topical cells before extraction, so one episode never straddles two subjects.

  2. 02 · EXTRACT

    Episodic extraction

    Each cell becomes a titled, timestamped narrative from the user's point of view. Optional profile, fact and foresight layers sit beside it.

  3. 03 · QUERY

    Multi-round retrieval

    Hybrid search per sub-query, fused by reciprocal rank; a decider marks core evidence and issues new queries for up to three rounds.

  4. 04 · PERSIST

    Local-first store

    SQLite and FAISS on disk, scoped per user, with any OpenAI-compatible model for extraction, retrieval and answering.

Research Blog

The posts behind the numbers and architecture above.

Membase · benchmark reportSeptember 2026