
Membase: Evidence-Centered Long-Term Memory for AI Agents, 93.1% on LoCoMo
Technical report · Unibase
Summary
Membase is Unibase’s evidence-centered memory system for AI agents. It forms memory from conversations, documents and agent execution traces, revises facts as new evidence arrives without erasing history, and retrieves evidence iteratively before answering. On conversational recall it scores 93.12% on LoCoMo, 92.60% on LongMemEval_S and 92.20% on DMR. Membase is available through Python, CLI, MCP and HTTP.
Abstract
AI applications need continuity across tasks: the ability to retain past experience, maintain useful knowledge, and reuse methods under changing conditions.
We present MEMBASE, Unibase’s evidence-centered memory system for AI agents, spanning episodic, semantic, and procedural functions. Formation extracts contextual narratives and source-supported abstractions; evolution revises factual applicability while retaining event history; utilization selects evidence for the task, with conversational retrieval accumulating a core set across rounds. We explain the design decisions connecting these functions to the implemented system, using one illustrative project throughout.
Historical repository evaluations of conversational recall report 93.12% on LoCoMo, 92.60% on LongMemEval_S, and 92.20% on DMR under dataset-specific configurations. Their diagnostics reveal complementary limits: compression can remove decisive details, mixed representations can obscure time, and available evidence can still be misinterpreted. These findings support evaluating the memory lifecycle as a whole.
The reported experiments cover conversational recall, while broader knowledge use and procedural transfer remain to be evaluated.
1 Why do AI agents need long-term memory?
AI applications encounter information whose value extends beyond the task in which it first appears. An earlier decision may explain a present constraint. A changed requirement may invalidate an old assumption. A successful recovery may suggest how to approach a similar failure. To benefit from that history, an application must preserve more than the text of an exchange: it needs access to what happened, what is known, and what can be learned about how to act.
We use AI memory to describe the persistent information and operations that provide this continuity beyond the current model context. Memory can draw on interactions, documents, and execution experience, and support subsequent reasoning, answers, or task execution. The central design question is how information remains useful as it is compressed, revised, and brought into new situations.
Three tensions shape this problem. Useful abstraction reduces detail, yet an unknown future question may depend on exactly what was omitted. New evidence can change what is currently applicable, yet the history behind a change remains valuable. Relevant memories can improve a task, yet selecting and interpreting them consumes context and computation. Memory quality therefore depends on decisions made throughout its lifecycle, rather than on storage or retrieval alone.
MEMBASE takes an evidence-centered approach: memories retain contextual or source support, factual revisions are tied to new observations, and conversational retrieval accumulates evidence before reading. These choices govern formation, evolution, and utilization. The functional vocabulary is episodic memory for situated experience, semantic memory for knowledge, and procedural memory for reusable methods. These terms describe what memory supports; they do not require a separate subsystem for every concept or bind each source format to one memory type.
This paper explains how source-supported abstraction, factual revision, and iterative evidence selection work together in an implemented memory system. We trace these operations through a shared example and use historical conversational evaluations to examine evidence loss, temporal interpretation, and reader errors. The evaluation scope is narrower than the system design: knowledge use and procedural transfer remain open empirical questions. Benchmark code is available in Membase Bench; Section 6.2 distinguishes the public harnesses from the historical records analyzed here.
2 What is evidence-centered memory?
2.1 Episodic, semantic and procedural memory
The distinction between episodic, semantic, and procedural memory provides an established vocabulary for reasoning about language-agent architectures [1]. Here we use it as an engineering abstraction for the functions that MEMBASE supports.
Episodic memory preserves situated experience: what happened, to whom, when, and under which circumstances. Its value lies in the relationships that explain an event. A past decision includes its motivation; an execution includes the conditions under which it succeeded or failed.
Semantic memory supports knowledge that can be used beyond the original encounter: facts, preferences, constraints, and domain understanding. Some knowledge remains stable; other knowledge has a period of applicability. Its interpretation depends on the distinction between explicit evidence and an inferred conclusion.
Procedural memory supports reusable methods: strategies, approaches, and guidance about their conditions of use. In MEMBASE, this function is supported by explicit guidance derived from execution experience. Its availability does not by itself imply learned model weights, automatic execution, or successful transfer to a new task.
These functions can coexist within the same source. A discussion can describe a failure, establish a requirement, and explain a recovery method. Likewise, an execution can provide both a particular experience and support for a reusable approach. Information sources determine how material enters the system; memory functions determine how that material can help a later task.

Figure 1: Write, update, and read paths in Membase. (1) Formation writes memory; (2) evolution reads existing state and writes revisions; (3) utilization returns retrieved evidence to the application through a caller-selected interface. Source records and derived representations use SQLite and FAISS persistence. These are operational responsibilities, not sequential stages or separate cognitive stores. Figures 2, 3, and 4 expand formation, temporal revision, and conversation retrieval, respectively.
2.2 One Lifecycle Across the Functions
Memory functions describe what a representation supports; lifecycle operations describe how it is formed, maintained, and used. Figure 1 separates the application boundary, memory operations, and persistent state. Source material enters formation through source-specific paths. Evolution reads existing state and writes evidence-supported revisions or consolidated guidance. A task enters a separate retrieval path, which reads memory and returns evidence. These are three responsibilities over persistent representations, rather than a mandatory sequence or three independent memory stores. Formation turns incoming information into interpretable memory. Evolution revises and consolidates that memory as evidence accumulates. Utilization selects and interprets what is needed for the current task. Source attribution, temporal meaning, and persistent storage support the lifecycle throughout.
The application supplies information and later requests relevant memory. It can use returned evidence directly or request an answer from an optional language-model reader. The current implementation provides distinct interfaces for interactions, documents, and agent experience; the caller chooses and combines them. Table 1 relates the implemented representations to these functions without imposing a one-to-one mapping.
2.3 Representations and Persistence
The three functions are realized through source-specific representations over shared persistence. Table 1 provides the implementation correspondence; the remaining sections introduce operators only where they explain a formation, revision, or retrieval decision.

Table 1: Functional interpretation of implemented representations. A representation can support more than one function; these correspondences do not imply complete or uniform capability.
SQLite stores source records, memories, and full-text indexes; FAISS provides a flat inner-product vector index. They do not share a transaction, so record and index consistency require explicit recovery procedures. Conversation operations are exposed through Python, CLI, MCP, and HTTP; knowledge and agent experience have their own engine, CLI, and MCP entry points. The application chooses when to invoke them.

Figure 2: Formation with source support. In panel 1, U1–U3 form one topic; U4 starts another. In panel 2, E1 preserves both dates and the causal relation from U1–U2. A separate factual path extracts the requirement from U3. Below, execution case C1 supports skill S1 under the tested staging conditions. The graph depicts relations within a narrative, not a separate causal-graph store. Illustrative contents.
2.4 A Running Example
Consider a project preparing a software release. On March 10, a rollback test fails. In a March 12 discussion, the team postpones the release because of that failure and explicitly requires rollback verification before launch. A later staging execution restores a snapshot and passes a health check, supplying a supporting case for the same ordered recovery steps. On March 18, an explicit replacement adds a production-like setup to the launch requirement. This is an illustrative scenario, not a measured system trace.
Episodic memory preserves the failed test and the decision in their original circumstances. Semantic memory captures the stated release requirement. Procedural memory makes the recovery approach available for similar tasks, together with its supporting experience. When the application later asks how to prepare another release, it needs the applicable requirement, relevant experience, and a method whose conditions fit the new situation. The following sections use this example to explain how memory acquires, retains, and contributes that meaning.
3 How does Membase form memories?
3.1 The Problem of Unknown Future Use
At formation time, the system does not know every question that will later be asked. Keeping all incoming material preserves detail but makes repeated use expensive. Strong compression reduces that cost while risking the loss of causes, qualifications, or participant relationships. The design objective is therefore to preserve enough meaning for reuse while retaining access to the evidence behind an abstraction.
For the release example, remembering only “release delayed” loses the reason for the decision. Remembering only the rollback requirement loses the episode that explains when and why it was introduced. A useful memory must preserve the relationship between experience and interpretation, while allowing each to be selected for different tasks.
3.2 The Design Choice: Context with Supported Abstraction
MEMBASE forms episodic memory around coherent experiences rather than treating every sentence as independent. The implemented interaction path identifies meaningful boundaries and extracts narratives that retain participants, time, decisions, and outcomes. Execution experience similarly retains a task’s intent and approach. This gives subsequent reasoning a context in which individual statements can be understood.
Semantic formation extracts reusable knowledge from explicit statements and supplied material. The implementation supports durable facts, incremental user summaries, and structured document knowledge. These representations reduce the need to repeatedly interpret an entire source. Their support matters: a requirement explicitly established by the team has a different status from a requirement inferred by a model.
Procedural formation abstracts approaches from execution experience. The implementation considers relevant prior experience and existing guidance when producing reusable methods, preserving links to supporting executions. In the running example, an observed recovery provides evidence for a method under those conditions; it does not establish that the method will succeed for every release.
These operations exist in different ingestion paths. The conceptual relationship between experience, knowledge, and methods does not imply that every source automatically produces all three, or that the system always derives one memory type from another.
Figure 2 traces these abstractions to their source evidence. E1 shows the causal relation preserved within an episode narrative, not a separately stored causal graph. The requirement has its own source support in U3; the figure does not make episode extraction a prerequisite for factual extraction.
Implemented formation. The Observer uses eight-message windows with two-message overlap. The optional episode path performs boundary detection to create MemCells, followed by episode extraction. A MemCell is an extraction input, not a fourth functional memory category. A first pass identifies completed topics; a final pass handles the tail. Extraction covers the whole cell before assignment to selected owners. An owner identifies retrieval placement, not the subject of every sentence.
Document parsing feeds topic-tree extraction and category assignment. Agent trace processing performs boundary detection, case extraction, quality filtering, clustering, and skill extraction. The concept of procedural memory refers to the resulting explicit guidance; no model-weight update is assumed.
3.3 The Trade-off: Compactness and Fidelity
Derived memory is necessarily selective. MEMBASE retains original interaction history and source or supporting-record associations in its other paths, making inspection possible when an abstraction is insufficient. However, retaining a source and including it in a task context are separate operations. The evaluated conversational path uses extracted narratives and does not automatically expand them back to the original messages.
Formation quality should therefore be judged against the details and relationships that later tasks require. If the failed test’s cause is omitted, retrieving the stored narrative more accurately cannot reconstruct it. Source-aware expansion is a candidate improvement, while faithful formation remains necessary for the context actually used.
4 How does Membase update memory over time?
4.1 The Problem of Changing Applicability
New information can supplement a past experience, revise a belief, or change the conditions under which a method is useful. These changes should have different consequences. A later successful rollback does not erase the earlier failure. An explicitly revised release requirement changes what governs the next launch. Additional executions can alter the support for a recovery method.
The design separates the persistence of history from the maintenance of currently applicable knowledge. It also distinguishes revising an abstraction from establishing that the abstraction is correct. This preserves the basis for both present decisions and retrospective questions.
4.2 The Design Choice: Revision with Evidence
The Linker compares recent observations with matching subject and type. Superseded observations retain content and receive an invalidation time and replacement reference. For validity interval [v_o, u_o), a query at t keeps a record when v_o ≤ t < u_o; missing endpoints are unbounded. Undated retrieval hides invalidated observations. Episodes do not use this interval filter.
Historical narratives retain event descriptions and dates. The running example makes the distinction concrete: the rollback failure occurred on March 10 and was discussed on March 12. The date of the report must not become the date of the failure. Current applicability and historical occurrence answer different questions, even when both appear in the same task context.
Profile updates support add, update, delete, and no-op operations, separating explicit information from inferred traits. Clusters organize related experience for consolidation; a Cluster is an organizational device rather than a separate memory function. Optional Foresight extraction represents future-oriented needs or possibilities, not completed events or scheduled actions. The three-function framework does not force every optional output into a new category.
Figure 3 aligns episode history with factual applicability. The March 18 replacement requires rollback verification in a production-like setup before launch. It changes which rule a dated factual query returns without changing the recorded March 10 failure or March 12 decision.
4.3 The Trade-off: Stability and Revision
Memory must remain stable enough to support continuity and responsive enough to accommodate correction. Overwriting history loses explanatory context; retaining every assertion without interpretation leaves contradictions to the application. Likewise, repeated similar experiences may support a method, but a generated confidence or quality assessment is not a measured guarantee of transfer.
A compact abstraction can also hide disagreement or temporal ambiguity. Evolution therefore requires attention to what changed, why it changed, and when the resulting information applies. The recorded temporal comparison in Section 6 illustrates how unclear date semantics can undermine the use of otherwise relevant memory.

Figure 3: Two meanings of time. Panel 1 retains the March 10 failure and the March 12 decision and report in E1. In panel 2, an explicit March 18 replacement closes r1 and opens r2; a March 15 query returns r1, while a March 20 query returns r2. Filled endpoints are inclusive and the open endpoint is exclusive. Both observations remain stored. The factual validity filter does not apply to the episode above. Illustrative example.
5 How does Membase retrieve memory?
5.1 The Problem of Relevance and Sufficiency
A task needs a limited selection of memory. The closest textual match may omit a necessary prerequisite, while a large volume of loosely related material can burden reasoning. Effective utilization must determine both which memories are relevant and whether their combination addresses the task.
For the next release, an application may need the current requirement, earlier failure circumstances, and an applicable recovery method. The application selects memory interfaces and composes their outputs; this example assumes no automatic planner across them.
5.2 The Design Choice: Evidence Before Inference
Exact matching recovers named entities and constraints, while semantic matching finds differently worded descriptions. The implementation first combines these signals, then selects and accumulates useful evidence across queries.
In the evaluated interaction path, reciprocal rank fusion (RRF) [7] combines lexical and dense lists without requiring their raw scores to share a scale. For candidate d and lists L(x) produced for query x,
s(d, x) = Σ over ℓ ∈ L(x) with d ∈ ℓ of 1 / (60 + rank_ℓ(d)) (1)
Ranks start at one. The multi-round path selects episodes when available unless the caller explicitly chooses other units. Lists for multiple requested memory types are interleaved; scores from separate corpora are not treated as calibrated probabilities.
A decider receives the original question, previously selected core evidence, and fresh candidates. It retains useful memories and proposes follow-up queries for unresolved aspects of the question. Selected evidence persists across rounds. Each discovered memory receives its best score across the queries that found it:
ρ(d) = max over x with d ∈ R(x) of s(d, x) (2)
Using a maximum avoids requiring evidence useful to one subquestion to match every query formulation. Final assembly places ranked core evidence first, eligible query reservations second, and remaining candidates third, with identity deduplication before truncation.
5.3 Retrieval Procedure and Controls
Figure 4 illustrates successive states of the same core set: E1 explains the delay, and E2 supplies the March 12 requirement after a follow-up query. Both are episode excerpts in this retrieval example; E2 is not the factual record r1 from the update example. The illustrated reader uses the packed context to produce the answer. Algorithm 1 specifies the procedure, and Table 2 gives its default budgets.
Algorithm 1 describes the default capped-core variant. C and U are identity-keyed containers with insertion order; identity is (source type, record ID). Core entries also retain their selecting query. Candidate block B preserves query labels and can contain the same memory under different queries. ∥ concatenates sequences, mem drops query labels, head_j takes at most j items, and unique retains first occurrences. Score sorting uses stable tie-breaking.
Recall and reservation. M includes the caller’s retrieval-unit, owner, and query-time settings. RECALL combines sparse and dense rankings and caches the result for repeated queries. Multiple requested units are interleaved. RESERVELEAD appends (x, D₁) only when D is nonempty, D₁ is not already core, query x has no prior reservation, and |G| < g. It does not seek a substitute if the first candidate is core. The reservation capacity is the subquery limit times the per-subquery guarantee setting; the original query also participates.

Figure 4: Iterative retrieval: (1) hybrid recall, (2) selection and retention, and (3) core-first packing. The decider retains E1, then selects E2 after a follow-up query while E1 persists. Packing ranks core evidence before eligible reservations and other hits, deduplicates identities, and caps the output. E1 and E2 are episode excerpts; gold highlights requirement content. Solid arrows show data flow, including the reader output; the dashed return issues a query. Ranks and contents are illustrative.

Algorithm 1: Core-first iterative retrieval.
Decision and stopping. DECIDE returns selected candidates S, at most L next queries, and a success flag. Model-call and parsing failures are retried. On persistent failure, the first f candidate occurrences enter fallback selection before deduplication. The no-progress count resets when the core grows; otherwise it increments. With p = 1, two consecutive no-progress rounds terminate the loop.
Assembly and bounds. The algorithm orders core items first, reservations second, and the remaining discovered candidates third. Identity deduplication precedes truncation, giving unique output with |E| ≤ K. Core overflow is disabled in this variant. For positive configured limits, there are at most R decider invocations before retries and 1 + (R − 1)L query-collection requests. These are control-flow bounds, not guarantees of answer quality.

Table 2: Default controls for Algorithm 1. Callers and benchmark harnesses can override them.
From evidence to application. Retrieval remains separate from inference. The application can use evidence directly or ask a reader to synthesize an answer. The evaluated reader receives selected Episode titles and narratives and a query date when the dataset supplies one. A separate legacy context packer can mix facts and source excerpts. Inspecting this actual context distinguishes missing evidence from misinterpretation.
Knowledge search supports lexical, vector, and hybrid modes with category-related ranking. Agent search targets cases, skills, or both and can use supporting-case relationships. These paths have distinct interfaces; they do not all execute Algorithm 1.
5.4 The Trade-off: Coverage, Interpretation, and Cost
Further search may improve coverage, but it consumes model calls and can increase the amount of context to interpret. More evidence can also introduce competing dates or duplicate accounts. Utilization therefore requires a task-appropriate balance between coverage and clarity.
The final check is whether memory changes the task outcome appropriately. A reader may confuse a proposed action with a completed action or miscalculate a time interval. A retrieved method may be relevant but unsuitable under a new constraint. Evidence coverage and downstream correctness must be evaluated separately, together with context size and latency. This connects utilization back to formation and evolution: the consumer depends on what survived and how its meaning was maintained.
6 How accurate is Membase? LoCoMo, LongMemEval and DMR results
6.1 Setup and Scope
The available experiments evaluate conversational recall using episodic narratives. We examine them through formation fidelity, temporal interpretation, and task utilization. They do not measure the complete semantic or procedural functions of the architecture, nor validate the illustrative release scenario.
We analyze the repository’s evaluation records dated September 18–21, 2026 [8]. These runs were not repeated for this manuscript, and their raw outputs and confidence intervals were not independently revalidated. The following results use dataset-specific readers and scoring protocols; comparisons within a recorded experiment are distinguished from comparisons across datasets.
LoCoMo evaluates 1,540 questions in categories 1–4, excluding category 5. Its owner-scoped narrative pipeline uses gpt-4.1-mini for extraction, decision-making, and reading, gpt-4o-mini for judging, and text-embedding-3-small for embeddings. LongMemEval_S uses all 500 questions, one store and owner per question, the supplied query date, and a gpt-5.5 reader. DMR uses 500 MSC-Self-Instruct examples, including four previous dialogues and the current dialogue, with first-person answering and gpt-4o-mini as reader and judge. Extraction and decision-making use gpt-4.1-mini in all three configurations.
Table 3 establishes the recorded operating points. The more informative question for system design is how performance changes when the memory context or reader changes, and where errors remain when retrieval succeeds.

Table 3: Recorded overall results for narrative memory with iterative retrieval. Each row uses its own evaluation protocol; the table does not pool accuracies or compare against matched external baselines.
6.2 Configuration and Reproducibility
The inspected revision is c7766de. The benchmark term “episode-only” means a context of Episode titles and narratives; the LoCoMo mixed-context arm adds Observation text. These are the narrative-only and narrative-plus-facts configurations discussed below, not full comparisons between episodic and semantic memory functions. The product default uses observations and single-pass hybrid retrieval, while the recorded evaluations emphasize optional episodes and iterative retrieval.
At this revision, the engine derives Observer and Linker enablement from not no_llm, rather than directly honoring the corresponding configuration fields. Reproduction must verify effective behavior at the chosen entry point. The ingestion adapter constructs message timestamps from the session date at 30-second intervals, preserving order but not original message times. Failed extraction can skip a cell; complete request handling therefore does not establish complete memory formation.
Membase Bench provides LoCoMo, LongMemEval, and DMR harnesses and judging commands. The public repository is:
👉 https://github.com/unibaseio/membase-bench
The results reported here remain attributed to the historical reproduction record [8], which specifies their datasets, models, and protocols. A new run should pin both the benchmark revision and memory-engine version, and preserve per-question outputs, packed context, extraction failures, retrieval controls, judge exclusions, tokens, and timings. Licensing and algorithm provenance are documented in the respective repositories. Rendering this manuscript does not rerun the evaluations.
6.3 Formation: The Cost of Lost Detail
DMR exposes a different limit. Its stores contain roughly five to seven narratives per question, so the top-20 budget can include the entire derived memory. The record attributes its 39 misses to details omitted during narrative compression, such as retaining a workplace while dropping the exact job role. Once a required detail is absent from all narrative text, retrieving more narratives cannot supply it. A source-aware reading path or a more faithful extraction step would address that mechanism more directly.
6.4 Evolution: Preserving Temporal Meaning
This comparison examines temporal interpretation in retrieved context; it does not directly test the factual validity-interval mechanism in Section 4. The LoCoMo record compares a narrative-only context with one that adds extracted factual statements. Overall accuracy is 93.12% for narratives only and 91.56% for the mixed context. The difference is larger on temporal questions: 91.6% versus 82.9% (Figure 5).
The record attributes part of this difference to date stamps attached to extracted facts. A fact marked with the session date can be read as an event on that date, even when the original dialogue described an earlier event. In the release example from Section 4, the March 12 report date must not replace the March 10 failure date. The example illustrates the mechanism proposed in the run analysis; it is not a benchmark case.

Figure 5: Recorded LoCoMo context comparison. The temporal subset shows a larger difference than overall accuracy. These are historical configuration arms, not a newly controlled ablation.
This comparison supports a concrete design concern: adding another representation can add an alternative interpretation of the same fact. A larger context is useful only if the additional evidence preserves the relationships and time semantics needed by the question. The result does not establish that factual knowledge is generally less useful. Explicit state revision and the interpretation of a past experience are different operations; the comparison concerns how two implemented representations were combined for these questions.
6.5 Utilization: Evidence and Interpretation
For the full LongMemEval_S run, the repository reports 99.95% retrieval recall and all gold sessions present for 99.8% of questions. Its diagnostic attributes the 37 incorrect answers to the reading stage with gold evidence in context. The separate sample analysis highlights cross-session aggregation, distinguishing a user action from an assistant suggestion, and temporal calculations. Finding the relevant sessions does not by itself resolve these operations.
Recorded reader substitutions further illustrate this distinction (Table 4). On the same LoCoMo stores, changing the reader from gpt-4.1-mini to gpt-5.5 changes accuracy from 93.12% to 93.18%: 42 answers improve and 41 regress. On a separate 100-question LongMemEval sample with fixed stores, the corresponding change is 83.0% to 96.0%. The 96.0% sample result is distinct from the 92.60% full-set result.

Table 4: Recorded accuracy (%) after changing the reader on existing stores. Columns abbreviate gpt-4.1-mini and gpt-5.5. The LongMemEval sample takes every fifth question.
The reader change has different effects on the two recorded workloads. This argues for diagnosing answer failures on the intended workload before allocating more computation to either retrieval or generation. It does not establish an architecture-wide preference for either model.
6.6 Utilization: The Cost of Additional Search
Figure 7 reports serial timing samples of 40 LoCoMo, 30 LongMemEval, and 30 DMR questions. Median search latency is 1.67, 2.53, and 1.13 seconds, respectively. Median total latency is 8.30, 14.7, and 3.21 seconds. Search includes model-guided evidence selection as well as vector and lexical retrieval.
Mean rendered context size is 6,562 tokens for LoCoMo, 8,970 for LongMemEval, and 1,602 for DMR. Additional rounds can help acquire missing evidence, but they also add model calls and may expand what the reader must inspect. The present records do not isolate the marginal benefit of each round. Selecting a deployment budget therefore requires measuring accuracy, context size, and latency together; the reported timings are workload observations, not service guarantees.
6.7 Comparison with Published Memory Systems
After the within-system diagnostics, Figure 6 places the recorded results alongside fixed literature configurations. We include the complete GPT-4.1-mini block of Hu et al.’s LoCoMo evaluation [9] and the GPT-4o-mini block of Rasmussen et al.’s DMR evaluation [4]. These are source-reported results, not new Unibase baseline runs or a current product leaderboard.

Figure 6: Historical literature results and separate Membase runs [9, 4, 8]. Both accuracy axes start at zero; dashed lines separate sources. LoCoMo external scores average three judges, while Membase uses one. DMR entries share a reader family but were run separately. Differences in configurations and grading prevent a matched ranking or statistical significance claim.

Figure 7: Total latency in recorded serial timing samples. Different readers and workloads contribute to the observed differences. The samples are smaller than the accuracy sets.
LoCoMo: shared reader, different grading. The external study averages GPT-4o-mini and two auxiliary judges; Membase uses GPT-4o-mini alone. Hu et al. run EverMemOS and MemoryOS end to end, while using official memory APIs for the other systems and standardizing final answer generation. Prompting, retrieval budgets, and memory construction remain system-specific. The numerical proximity of 93.12 and 93.05 does not establish either superiority or equivalence. Membase also includes EverOS-derived components documented in NOTICE; this comparison concerns system configurations rather than independent algorithm attribution.
DMR: retain the reader and context baselines. Zep’s 98.2% and the 98.0% full-conversation result use GPT-4o-mini. The same paper’s 94.8% Zep and 93.4% MemGPT figures use GPT-4-turbo; its MemGPT value is quoted from prior work. Those figures must not be relabeled as GPT-4o-mini results. Membase’s 92.20% is lower than the published GPT-4o-mini Zep and full-conversation scores, consistent with examining the recorded extraction omissions below. Separate runs and incomplete protocol alignment prevent attributing the gap to a single component.
A controlled comparison should fix dataset files, answer and judge revisions, prompts, exclusions, and context budgets, retain per-question outputs, and report paired uncertainty. The full LongMemEval_S Membase result uses a different reader (gpt-5.5), so we do not extend this reader-grouped comparison to that benchmark. Source versions, values, and protocol notes are recorded in docs/data/memory-comparison-sources.json; no comparison experiment was run for this revision.
7 Discussion and Limitations
The lifecycle provides a concise way to diagnose memory failures. Formation asks whether the necessary meaning survived. Evolution asks whether that meaning remains applicable and temporally interpretable. Utilization asks whether the right evidence reached the task and was used correctly. DMR’s detail omissions, LoCoMo’s temporal-context comparison, and LongMemEval’s reader errors illustrate these complementary questions.
The functional framework also clarifies what remains untested. Episodic evaluation should examine contextual fidelity and recall. Semantic evaluation should examine knowledge accuracy, revision, and applicability over time. Procedural evaluation should measure transfer to new tasks under relevant conditions. The reported conversation benchmarks touch parts of this space but do not separately establish all three capabilities. New procedural experiments, in particular, need downstream execution outcomes rather than similarity scores for retrieved guidance.
Our evidence consists of historical repository records with different readers and judges across datasets. The external results in Section 6.7 are literature context; there is no matched external rerun, complete lifecycle ablation, or end-to-end demonstration spanning every input path. The observed comparisons support diagnosis, rather than universal causal conclusions. The system also lacks a uniform temporal model and automatic orchestration across all memory functions. Section 6.2 records the entry-point and timestamp constraints that affect reproduction. Owner filtering is not complete tenant authorization. Model and embedding calls can use external services; a built-in local embedding provider is not implemented, so local persistence and --no-llm do not establish entirely local computation.
8 Related Work
CoALA uses episodic, semantic, and procedural distinctions to describe memory within a broader language-agent architecture [1]. We use this vocabulary to explain the functions of a memory service, with procedural support expressed as reusable guidance grounded in execution experience. This is a functional interpretation, not a claim that the implementation reproduces human cognition or instantiates the full CoALA architecture.
Retrieval-augmented generation couples a language model with retrieved evidence [2]. Our account extends the design discussion to how evidence is formed and maintained before retrieval. MemGPT studies model-directed management of memory tiers [3]; Zep emphasizes temporal knowledge organization [4]. LoCoMo and LongMemEval provide evaluation settings for long-term conversational recall [5, 6]. EverMemOS studies a lifecycle of episodic formation, consolidation, and recollection [9]; Membase integrates EverOS-derived components as documented in NOTICE. Section 6.7 provides source-reported comparisons without claiming matched reruns.
9 Conclusion
MEMBASE connects formation, evolution, and utilization through an evidence-centered design: supported abstraction, factual revision, and evidence accumulation. The architecture preserves the meaning behind information, maintains its applicability as evidence changes, and makes selected memory available to a current task. Historical evaluations expose why these responsibilities must be considered together: lost detail cannot be retrieved, ambiguous time can distort context, and complete evidence can still be misread. Evaluating memory as a lifecycle makes these failures easier to locate and defines the additional evidence needed for broader knowledge use and procedural transfer.
References
[1] Theodore R. Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L. Griffiths. 2024. Cognitive Architectures for Language Agents. Transactions on Machine Learning Research.
[2] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, 33.
[3] Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G. Patil, Ion Stoica, and Joseph E. Gonzalez. 2023. MemGPT: Towards LLMs as Operating Systems. arXiv preprint arXiv:2310.08560.
[4] Preston Rasmussen, Pavlo Paliychuk, Travis Beauvais, Jack Ryan, and Daniel Chalef. 2025. Zep: A Temporal Knowledge Graph Architecture for Agent Memory. arXiv preprint arXiv:2501.13956v1. Table 1 and Sections 4.1–4.2; accessed October 3, 2026.
[5] Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, and Yuwei Fang. 2024. Evaluating Very Long-Term Conversational Memory of LLM Agents. In Proceedings of ACL, pages 13851–13870.
[6] Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, and Dong Yu. 2025. LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory. In International Conference on Learning Representations.
[7] Gordon V. Cormack, Charles L. A. Clarke, and Stefan Büttcher. 2009. Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. In Proceedings of SIGIR, pages 758–759.
[8] Unibase. 2026. Benchmark reproduction records. bench/locomo/REPRODUCE.md, revision c7766de. Evaluation records dated September 18–21, 2026.
[9] Chuanrui Hu, Xingze Gao, Zuyi Zhou, Dannong Xu, Yi Bai, Xintong Li, Hui Zhang, Tong Li, Chong Zhang, Lidong Bing, and Yafeng Deng. 2026. EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning. arXiv preprint arXiv:2601.02163v2. Table 1 and Sections 4.1, A.1; accessed October 3, 2026.