Elephantus: a Memory Layer for AI Apps

By Mohammed Akram Khan Lodi • 11 minutes read •


RAG finds similar text. Memory tracks what is currently true about a user.

Elephantus is a small memory engine for AI apps that runs on your own machine. You send it messages; it pulls out short facts about the user, links each new fact to the ones it already has, and keeps track of which facts are still true. You can use it from a REST API, from Claude Desktop over MCP, or from a terminal-style web UI.

An elephant never forgets, but it does know which of its memories are out of date.

A walkthrough of the web UI: chat, linking decisions, the memory graph, and the RAG-vs-memory comparison.

The problem: similarity isn’t truth

Suppose a user tells an assistant three things over a few weeks:

  1. “I love Adidas sneakers”
  2. “My Adidas broke after a month”
  3. “I’m switching to Puma”

Later they ask, “What sneakers should I buy?”

Most “memory” features in AI apps are really retrieval over a chat log. That works until the user changes their mind, and people change their minds all the time. I built Elephantus to make that difference concrete and measurable.

How it works

The engine

ConceptWhat it does
Documents vs memoriesRaw messages are stored and chunked (that’s the RAG baseline). An LLM extracts short, atomic memories from them.
Container tagsEvery document and memory belongs to a container_tag (for example khan or work). Tags never mix, not even during linking.
RelationshipsEach new fact is compared with the most similar current memories and labelled NEW, UPDATES (the old fact is retired and an edge is stored), EXTENDS (both stay current, with an edge) or DUPLICATE (skipped). Every decision is logged with a reason.
Kindsstatic facts are long-term (where you study, lasting preferences). dynamic facts are recent or ongoing (this week’s project).
ForgettingTime-bound facts (“exam tomorrow”) get an expiry at extraction time and drop out of retrieval once it passes. Facts can also be forgotten explicitly.
Hybrid searchEmbedding similarity plus SQLite FTS5 keyword search, merged with Reciprocal Rank Fusion. Only current, unexpired memories are searched.
ProfileOne call returns the static facts, the recent dynamic facts and, optionally, search results for a query: a ready-made context block for a prompt.

What happens when you add a message

"I'm switching to Puma"
  │ 1. store document, chunk, embed chunks          (RAG baseline data)
  │ 2. LLM extracts facts → "User is switching to Puma sneakers" (dynamic, no expiry)
  │ 3. embed fact, shortlist the 5 most similar CURRENT memories in this container
  │ 4. LLM judges → UPDATES "User loves Adidas sneakers"
  ▼ 5. insert new memory, mark old is_latest=0, add edge new→old, log the decision

Linking is two steps on purpose. Embeddings give a cheap shortlist, and an LLM judge makes the precise call. The judge always sees the top 5 current memories, with no similarity threshold, so it can still catch an update between facts that share almost no words: “switching to Puma” and “loves Adidas” have none in common.

Architecture

The engine (elephantus/engine.py) is a plain Python library that knows nothing about HTTP, MCP or the UI. Three thin entry points sit on top of it:

All three share one engine and one SQLite file, so a fact saved from Claude Desktop shows up in the UI straight away. Behind the engine are:

The demo

The web UI is a dark, terminal-style page with tabs for chat · memories · graph · profile · search · eval · log. A sidebar holds the container tag, a clock for simulating time, live counts, and the current memory list.

Chat: RAG and memory, side by side

Elephantus chat tab: a message is broken into two NEW facts, followed by two answers side by side, one from similarity-only RAG and one from current memory facts, each with the context it used
Each message shows its linking decisions like a coding agent's tool calls. Every question is answered twice with the same model and prompt, once from RAG and once from memory, so the only difference is the context.

The RAG answer’s context is a list of past messages, including stale ones like “I love adidas sneakers”. The memory answer’s context lists only what is currently true, newest first.

Graph: how facts relate

Elephantus graph tab: memory cards connected by blue EXTENDS edges and a red UPDATES edge. 'User loves Adidas sneakers' is struck through and marked outdated
The red UPDATES edge from "switching to Puma" retires "loves Adidas sneakers" (struck through). Blue EXTENDS edges add detail while keeping both facts current.

Profile: what the model actually sees

Elephantus profile tab: static long-term facts on the left, recent dynamic facts on the right, and the exact context prompt below
The profile splits facts into static and dynamic, and shows the exact context prompt an app would inject. Outdated and expired facts are excluded.

Log: every decision, with a reason

Elephantus log tab: a list of linking decisions (NEW, EXTENDS, UPDATES), each with the fact, the judge's one-line reason and a timestamp
Every linking decision is logged with the judge's reasoning, which makes the memory easy to audit and debug.

Forgetting, with a simulated clock

Send “I have an exam tomorrow”, then drag the Clock slider to +72h. The memory turns expired and disappears from answers, the profile and search. The same time_offset_hours parameter is accepted by the API, which is how the evaluation tests expiry without waiting days.

Using it from Claude Desktop (MCP)

The MCP server exposes three tools:

ToolWhat it does
memory(content, action="save" | "forget")Save information (extract + link), or forget the best-matching memory
recall(query, limit?)Search current memories, plus a profile summary
context()The full profile, to inject at the start of a conversation

In one chat say “Remember that I’m switching to Puma sneakers”. In a brand-new chat, ask “What do you know about my shoe preferences?”: Claude calls recall and answers from the same store the web UI is showing.

REST API

The core calls are add, search and profile, with interactive docs at /docs:

curl -s localhost:8000/v1/add -H 'content-type: application/json' \
  -d '{"content": "I am switching to Puma", "container_tag": "khan"}'

curl -s localhost:8000/v1/search -H 'content-type: application/json' \
  -d '{"q": "What sneakers should I buy?", "container_tag": "khan", "mode": "memories"}'

/v1/search takes a mode: memories (hybrid search over current facts), documents (the naive RAG baseline) or hybrid (both). Other endpoints cover chat (answers twice and returns both contexts), forgetting, and listing a container’s memories, graph, log and documents.

Evaluation

elephantus eval runs 25 scripted scenarios in three categories:

CategorynWhat it tests
knowledge_update11A fact changes (city, job, phone, diet…), then a question about the current value
extension7Detail builds up across several messages (job → team → role → language)
expiry7A temporary fact (“exam tomorrow”, “in Tokyo this week”), then a question days later

Each scenario starts a fresh container with 4 unrelated distractor messages, adds its own messages one simulated hour apart, then asks its question. Three retrieval modes return their top 3: RAG (chunk similarity), Memory (hybrid search over current memories) and Hybrid (memories plus chunks).

Results

Run with Azure AI Foundry gpt-4o and BAAI/bge-small-en-v1.5 embeddings, k = 3, about 2 minutes:

CategorynRAG Recall@3Memory Recall@3Hybrid Recall@3RAG staleMemory staleHybrid stale
knowledge_update11100%100%91%100%27%82%
extension793%96%58%n/an/an/a
expiry7100%100%100%100%0%43%
overall2598%99%84%100%17%67%
Elephantus eval tab: the results table above terminal-style bar charts of Recall@3 and stale-fact rate for RAG, memory and hybrid
The same results in the UI's eval tab.

What the numbers show:

An offline run swaps the LLM for a rule-based stand-in and hash embeddings. It reaches only 68% recall and a 33% stale rate, which shows how much of the quality comes from the LLM judge and a real embedding model.

Design decisions

Testing

The pytest suite needs no API key or downloads: the LLM is replaced by a scripted fake. It covers storage and container isolation, chunking, extraction (including malformed output), the sneaker sequence (UPDATES / EXTENDS / DUPLICATE), hybrid search, expiry with simulated time, the profile, forgetting, the REST API, the MCP tools, and the web UI, including a real-browser run of the sneaker flow with Playwright. An optional live test runs end to end against a real provider.

Limitations

Stack

Python 3.11+, FastAPI, SQLite (FTS5), NumPy, fastembed (BAAI/bge-small-en-v1.5), the official MCP Python SDK, and Anthropic / OpenAI / Azure AI Foundry / Ollama as LLM providers. MIT licensed.