<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Akram - agents</title>
    <subtitle>Mohammed Akram Khan Lodi — CS undergraduate researching recursive language models, self-improving agent harnesses, and the energy cost of efficient AI.</subtitle>
    <link rel="self" type="application/atom+xml" href="http://akramlodi.com/tags/agents/atom.xml"/>
    <link rel="alternate" type="text/html" href="http://akramlodi.com/"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2026-10-04T00:00:00+00:00</updated>
    <id>http://akramlodi.com/tags/agents/atom.xml</id>
    <entry xml:lang="en">
        <title>Elephantus: a Memory Layer for AI Apps</title>
        <published>2026-10-04T00:00:00+00:00</published>
        <updated>2026-10-04T00:00:00+00:00</updated>
        
        <author>
          <name>
            Mohammed Akram Khan Lodi
          </name>
        </author>
        
        <link rel="alternate" type="text/html" href="http://akramlodi.com/projects/elephantus/"/>
        <id>http://akramlodi.com/projects/elephantus/</id>
        
        <content type="html" xml:base="http://akramlodi.com/projects/elephantus/">&lt;ul class=&quot;link-row&quot;&gt;
&lt;li&gt;&lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;akramlodi&#x2F;elephantus&quot;&gt;Code on GitHub&lt;&#x2F;a&gt;&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;RAG finds similar text. Memory tracks what is &lt;em&gt;currently true&lt;&#x2F;em&gt; about a user.&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;Elephantus is a small memory engine for AI apps that runs on your own machine. You send it messages; it pulls out short facts about the user, links each new fact to the ones it already has, and keeps track of which facts are still true. You can use it from a REST API, from Claude Desktop over MCP, or from a terminal-style web UI.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;em&gt;An elephant never forgets, but it does know which of its memories are out of date.&lt;&#x2F;em&gt;&lt;&#x2F;p&gt;
&lt;figure&gt;
&lt;video src=&quot;demo.mp4&quot; poster=&quot;demo-poster.jpg&quot; controls muted playsinline preload=&quot;metadata&quot; width=&quot;100%&quot;&gt;&lt;&#x2F;video&gt;
&lt;figcaption&gt;A walkthrough of the web UI: chat, linking decisions, the memory graph, and the RAG-vs-memory comparison.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;h2 id=&quot;the-problem-similarity-isn-t-truth&quot;&gt;The problem: similarity isn’t truth&lt;&#x2F;h2&gt;
&lt;p&gt;Suppose a user tells an assistant three things over a few weeks:&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;“I love Adidas sneakers”&lt;&#x2F;li&gt;
&lt;li&gt;“My Adidas broke after a month”&lt;&#x2F;li&gt;
&lt;li&gt;“I’m switching to Puma”&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;Later they ask, &lt;strong&gt;“What sneakers should I buy?”&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Naive RAG&lt;&#x2F;strong&gt; stores the raw messages and fetches the ones most &lt;em&gt;similar&lt;&#x2F;em&gt; to the question. “I love Adidas sneakers” is the closest match (it even contains the word &lt;em&gt;sneakers&lt;&#x2F;em&gt;), so the assistant recommends &lt;strong&gt;Adidas&lt;&#x2F;strong&gt;. Similarity says nothing about whether a fact is still true.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Memory&lt;&#x2F;strong&gt; turns each message into atomic facts and compares every new fact with what it already knows. “User is switching to Puma” &lt;strong&gt;updates&lt;&#x2F;strong&gt; “User loves Adidas sneakers”. The old fact is marked outdated: it’s kept for history but never retrieved again. The assistant recommends &lt;strong&gt;Puma&lt;&#x2F;strong&gt;.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Most “memory” features in AI apps are really retrieval over a chat log. That works until the user changes their mind, and people change their minds all the time. I built Elephantus to make that difference concrete and measurable.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;how-it-works&quot;&gt;How it works&lt;&#x2F;h2&gt;
&lt;h3 id=&quot;the-engine&quot;&gt;The engine&lt;&#x2F;h3&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Concept&lt;&#x2F;th&gt;&lt;th&gt;What it does&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Documents vs memories&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Raw messages are stored and chunked (that’s the RAG baseline). An LLM extracts short, atomic &lt;strong&gt;memories&lt;&#x2F;strong&gt; from them.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Container tags&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Every document and memory belongs to a &lt;code&gt;container_tag&lt;&#x2F;code&gt; (for example &lt;code&gt;khan&lt;&#x2F;code&gt; or &lt;code&gt;work&lt;&#x2F;code&gt;). Tags never mix, not even during linking.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Relationships&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Each new fact is compared with the most similar &lt;em&gt;current&lt;&#x2F;em&gt; memories and labelled &lt;code&gt;NEW&lt;&#x2F;code&gt;, &lt;code&gt;UPDATES&lt;&#x2F;code&gt; (the old fact is retired and an edge is stored), &lt;code&gt;EXTENDS&lt;&#x2F;code&gt; (both stay current, with an edge) or &lt;code&gt;DUPLICATE&lt;&#x2F;code&gt; (skipped). Every decision is logged with a reason.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Kinds&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;&lt;code&gt;static&lt;&#x2F;code&gt; facts are long-term (where you study, lasting preferences). &lt;code&gt;dynamic&lt;&#x2F;code&gt; facts are recent or ongoing (this week’s project).&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Forgetting&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Time-bound facts (“exam tomorrow”) get an expiry at extraction time and drop out of retrieval once it passes. Facts can also be forgotten explicitly.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Hybrid search&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;Embedding similarity plus SQLite FTS5 keyword search, merged with &lt;strong&gt;Reciprocal Rank Fusion&lt;&#x2F;strong&gt;. Only current, unexpired memories are searched.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;Profile&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;One call returns the static facts, the recent dynamic facts and, optionally, search results for a query: a ready-made context block for a prompt.&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;h3 id=&quot;what-happens-when-you-add-a-message&quot;&gt;What happens when you add a message&lt;&#x2F;h3&gt;
&lt;pre class=&quot;z-code&quot;&gt;&lt;code&gt;&lt;span class=&quot;z-text z-plain&quot;&gt;&amp;quot;I&amp;#39;m switching to Puma&amp;quot;
&lt;&#x2F;span&gt;&lt;span class=&quot;z-text z-plain&quot;&gt;  │ 1. store document, chunk, embed chunks          (RAG baseline data)
&lt;&#x2F;span&gt;&lt;span class=&quot;z-text z-plain&quot;&gt;  │ 2. LLM extracts facts → &amp;quot;User is switching to Puma sneakers&amp;quot; (dynamic, no expiry)
&lt;&#x2F;span&gt;&lt;span class=&quot;z-text z-plain&quot;&gt;  │ 3. embed fact, shortlist the 5 most similar CURRENT memories in this container
&lt;&#x2F;span&gt;&lt;span class=&quot;z-text z-plain&quot;&gt;  │ 4. LLM judges → UPDATES &amp;quot;User loves Adidas sneakers&amp;quot;
&lt;&#x2F;span&gt;&lt;span class=&quot;z-text z-plain&quot;&gt;  ▼ 5. insert new memory, mark old is_latest=0, add edge new→old, log the decision
&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Linking is two steps on purpose. Embeddings give a cheap shortlist, and an LLM judge makes the precise call. The judge always sees the top 5 current memories, with no similarity threshold, so it can still catch an update between facts that share almost no words: “switching to Puma” and “loves Adidas” have none in common.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;architecture&quot;&gt;Architecture&lt;&#x2F;h3&gt;
&lt;p&gt;The engine (&lt;code&gt;elephantus&#x2F;engine.py&lt;&#x2F;code&gt;) is a plain Python library that knows nothing about HTTP, MCP or the UI. Three thin entry points sit on top of it:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;a &lt;strong&gt;REST API&lt;&#x2F;strong&gt; (FastAPI), which also serves the web UI;&lt;&#x2F;li&gt;
&lt;li&gt;an &lt;strong&gt;MCP server&lt;&#x2F;strong&gt; over stdio, which Claude Desktop launches;&lt;&#x2F;li&gt;
&lt;li&gt;the &lt;strong&gt;web UI&lt;&#x2F;strong&gt;, a hand-written static page that only calls the public REST API.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;All three share &lt;strong&gt;one engine and one SQLite file&lt;&#x2F;strong&gt;, so a fact saved from Claude Desktop shows up in the UI straight away. Behind the engine are:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;an LLM provider of your choice:&lt;&#x2F;strong&gt; Anthropic, OpenAI, Azure AI Foundry, or Ollama for fully offline use;&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;local embeddings:&lt;&#x2F;strong&gt; &lt;code&gt;bge-small&lt;&#x2F;code&gt; via fastembed, about 70 MB and downloaded on first use;&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;a single SQLite database&lt;&#x2F;strong&gt;, with FTS5 for keyword search.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;the-demo&quot;&gt;The demo&lt;&#x2F;h2&gt;
&lt;p&gt;The web UI is a dark, terminal-style page with tabs for &lt;strong&gt;chat · memories · graph · profile · search · eval · log&lt;&#x2F;strong&gt;. A sidebar holds the container tag, a clock for simulating time, live counts, and the current memory list.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;chat-rag-and-memory-side-by-side&quot;&gt;Chat: RAG and memory, side by side&lt;&#x2F;h3&gt;
&lt;figure&gt;
&lt;img src=&quot;chat.jpg&quot; alt=&quot;Elephantus chat tab: a message is broken into two NEW facts, followed by two answers side by side, one from similarity-only RAG and one from current memory facts, each with the context it used&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
&lt;figcaption&gt;Each message shows its linking decisions like a coding agent&#x27;s tool calls. Every question is answered twice with the same model and prompt, once from RAG and once from memory, so the only difference is the context.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;The RAG answer’s context is a list of past messages, including stale ones like “I love adidas sneakers”. The memory answer’s context lists only what is currently true, newest first.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;graph-how-facts-relate&quot;&gt;Graph: how facts relate&lt;&#x2F;h3&gt;
&lt;figure&gt;
&lt;img src=&quot;graph.jpg&quot; alt=&quot;Elephantus graph tab: memory cards connected by blue EXTENDS edges and a red UPDATES edge. &#x27;User loves Adidas sneakers&#x27; is struck through and marked outdated&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
&lt;figcaption&gt;The red &lt;code&gt;UPDATES&lt;&#x2F;code&gt; edge from &quot;switching to Puma&quot; retires &quot;loves Adidas sneakers&quot; (struck through). Blue &lt;code&gt;EXTENDS&lt;&#x2F;code&gt; edges add detail while keeping both facts current.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;h3 id=&quot;profile-what-the-model-actually-sees&quot;&gt;Profile: what the model actually sees&lt;&#x2F;h3&gt;
&lt;figure&gt;
&lt;img src=&quot;profile.jpg&quot; alt=&quot;Elephantus profile tab: static long-term facts on the left, recent dynamic facts on the right, and the exact context prompt below&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
&lt;figcaption&gt;The profile splits facts into static and dynamic, and shows the exact context prompt an app would inject. Outdated and expired facts are excluded.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;h3 id=&quot;log-every-decision-with-a-reason&quot;&gt;Log: every decision, with a reason&lt;&#x2F;h3&gt;
&lt;figure&gt;
&lt;img src=&quot;log.jpg&quot; alt=&quot;Elephantus log tab: a list of linking decisions (NEW, EXTENDS, UPDATES), each with the fact, the judge&#x27;s one-line reason and a timestamp&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
&lt;figcaption&gt;Every linking decision is logged with the judge&#x27;s reasoning, which makes the memory easy to audit and debug.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;h3 id=&quot;forgetting-with-a-simulated-clock&quot;&gt;Forgetting, with a simulated clock&lt;&#x2F;h3&gt;
&lt;p&gt;Send “I have an exam tomorrow”, then drag the &lt;strong&gt;Clock&lt;&#x2F;strong&gt; slider to +72h. The memory turns &lt;em&gt;expired&lt;&#x2F;em&gt; and disappears from answers, the profile and search. The same &lt;code&gt;time_offset_hours&lt;&#x2F;code&gt; parameter is accepted by the API, which is how the evaluation tests expiry without waiting days.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;using-it-from-claude-desktop-mcp&quot;&gt;Using it from Claude Desktop (MCP)&lt;&#x2F;h2&gt;
&lt;p&gt;The MCP server exposes three tools:&lt;&#x2F;p&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Tool&lt;&#x2F;th&gt;&lt;th&gt;What it does&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;memory(content, action=&quot;save&quot; | &quot;forget&quot;)&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;Save information (extract + link), or forget the best-matching memory&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;recall(query, limit?)&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;Search current memories, plus a profile summary&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;context()&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;The full profile, to inject at the start of a conversation&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;In one chat say &lt;em&gt;“Remember that I’m switching to Puma sneakers”&lt;&#x2F;em&gt;. In a brand-new chat, ask &lt;em&gt;“What do you know about my shoe preferences?”&lt;&#x2F;em&gt;: Claude calls &lt;code&gt;recall&lt;&#x2F;code&gt; and answers from the same store the web UI is showing.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;rest-api&quot;&gt;REST API&lt;&#x2F;h2&gt;
&lt;p&gt;The core calls are &lt;code&gt;add&lt;&#x2F;code&gt;, &lt;code&gt;search&lt;&#x2F;code&gt; and &lt;code&gt;profile&lt;&#x2F;code&gt;, with interactive docs at &lt;code&gt;&#x2F;docs&lt;&#x2F;code&gt;:&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash z-code&quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;&lt;span class=&quot;z-source z-shell z-bash&quot;&gt;&lt;span class=&quot;z-meta z-function-call z-shell&quot;&gt;&lt;span class=&quot;z-variable z-function z-shell&quot;&gt;curl&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;span class=&quot;z-meta z-function-call z-arguments z-shell&quot;&gt;&lt;span class=&quot;z-variable z-parameter z-option z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-parameter z-shell&quot;&gt; -&lt;&#x2F;span&gt;s&lt;&#x2F;span&gt; localhost:8000&#x2F;v1&#x2F;add&lt;span class=&quot;z-variable z-parameter z-option z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-parameter z-shell&quot;&gt; -&lt;&#x2F;span&gt;H&lt;&#x2F;span&gt; &lt;span class=&quot;z-string z-quoted z-single z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string z-begin z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;content-type: application&#x2F;json&lt;span class=&quot;z-punctuation z-definition z-string z-end z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt; &lt;span class=&quot;z-punctuation z-separator z-continuation z-line z-shell&quot;&gt;\
&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;span class=&quot;z-source z-shell z-bash&quot;&gt;&lt;span class=&quot;z-meta z-function-call z-arguments z-shell&quot;&gt;&lt;span class=&quot;z-variable z-parameter z-option z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-parameter z-shell&quot;&gt;  -&lt;&#x2F;span&gt;d&lt;&#x2F;span&gt; &lt;span class=&quot;z-string z-quoted z-single z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string z-begin z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;{&amp;quot;content&amp;quot;: &amp;quot;I am switching to Puma&amp;quot;, &amp;quot;container_tag&amp;quot;: &amp;quot;khan&amp;quot;}&lt;span class=&quot;z-punctuation z-definition z-string z-end z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;&#x2F;span&gt;&lt;span class=&quot;z-source z-shell z-bash&quot;&gt;
&lt;&#x2F;span&gt;&lt;span class=&quot;z-source z-shell z-bash&quot;&gt;&lt;span class=&quot;z-meta z-function-call z-shell&quot;&gt;&lt;span class=&quot;z-variable z-function z-shell&quot;&gt;curl&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;span class=&quot;z-meta z-function-call z-arguments z-shell&quot;&gt;&lt;span class=&quot;z-variable z-parameter z-option z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-parameter z-shell&quot;&gt; -&lt;&#x2F;span&gt;s&lt;&#x2F;span&gt; localhost:8000&#x2F;v1&#x2F;search&lt;span class=&quot;z-variable z-parameter z-option z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-parameter z-shell&quot;&gt; -&lt;&#x2F;span&gt;H&lt;&#x2F;span&gt; &lt;span class=&quot;z-string z-quoted z-single z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string z-begin z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;content-type: application&#x2F;json&lt;span class=&quot;z-punctuation z-definition z-string z-end z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt; &lt;span class=&quot;z-punctuation z-separator z-continuation z-line z-shell&quot;&gt;\
&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;span class=&quot;z-source z-shell z-bash&quot;&gt;&lt;span class=&quot;z-meta z-function-call z-arguments z-shell&quot;&gt;&lt;span class=&quot;z-variable z-parameter z-option z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-parameter z-shell&quot;&gt;  -&lt;&#x2F;span&gt;d&lt;&#x2F;span&gt; &lt;span class=&quot;z-string z-quoted z-single z-shell&quot;&gt;&lt;span class=&quot;z-punctuation z-definition z-string z-begin z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;{&amp;quot;q&amp;quot;: &amp;quot;What sneakers should I buy?&amp;quot;, &amp;quot;container_tag&amp;quot;: &amp;quot;khan&amp;quot;, &amp;quot;mode&amp;quot;: &amp;quot;memories&amp;quot;}&lt;span class=&quot;z-punctuation z-definition z-string z-end z-shell&quot;&gt;&amp;#39;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;&lt;&#x2F;span&gt;
&lt;&#x2F;span&gt;&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;&lt;code&gt;&#x2F;v1&#x2F;search&lt;&#x2F;code&gt; takes a &lt;code&gt;mode&lt;&#x2F;code&gt;: &lt;code&gt;memories&lt;&#x2F;code&gt; (hybrid search over current facts), &lt;code&gt;documents&lt;&#x2F;code&gt; (the naive RAG baseline) or &lt;code&gt;hybrid&lt;&#x2F;code&gt; (both). Other endpoints cover chat (answers twice and returns both contexts), forgetting, and listing a container’s memories, graph, log and documents.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;evaluation&quot;&gt;Evaluation&lt;&#x2F;h2&gt;
&lt;p&gt;&lt;code&gt;elephantus eval&lt;&#x2F;code&gt; runs &lt;strong&gt;25 scripted scenarios&lt;&#x2F;strong&gt; in three categories:&lt;&#x2F;p&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Category&lt;&#x2F;th&gt;&lt;th&gt;n&lt;&#x2F;th&gt;&lt;th&gt;What it tests&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;knowledge_update&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;11&lt;&#x2F;td&gt;&lt;td&gt;A fact changes (city, job, phone, diet…), then a question about the current value&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;extension&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;7&lt;&#x2F;td&gt;&lt;td&gt;Detail builds up across several messages (job → team → role → language)&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;code&gt;expiry&lt;&#x2F;code&gt;&lt;&#x2F;td&gt;&lt;td&gt;7&lt;&#x2F;td&gt;&lt;td&gt;A temporary fact (“exam tomorrow”, “in Tokyo this week”), then a question days later&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;Each scenario starts a fresh container with 4 unrelated distractor messages, adds its own messages one simulated hour apart, then asks its question. Three retrieval modes return their top 3: &lt;strong&gt;RAG&lt;&#x2F;strong&gt; (chunk similarity), &lt;strong&gt;Memory&lt;&#x2F;strong&gt; (hybrid search over current memories) and &lt;strong&gt;Hybrid&lt;&#x2F;strong&gt; (memories plus chunks).&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Recall@3:&lt;&#x2F;strong&gt; the fraction of expected &lt;em&gt;current&lt;&#x2F;em&gt; facts found in the top 3.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Stale-fact rate:&lt;&#x2F;strong&gt; the fraction of scenarios whose top 3 contains an outdated or expired fact.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h3 id=&quot;results&quot;&gt;Results&lt;&#x2F;h3&gt;
&lt;p&gt;Run with Azure AI Foundry &lt;code&gt;gpt-4o&lt;&#x2F;code&gt; and &lt;code&gt;BAAI&#x2F;bge-small-en-v1.5&lt;&#x2F;code&gt; embeddings, k = 3, about 2 minutes:&lt;&#x2F;p&gt;
&lt;div class=&quot;table-scroll&quot;&gt;
&lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Category&lt;&#x2F;th&gt;&lt;th&gt;n&lt;&#x2F;th&gt;&lt;th&gt;RAG Recall@3&lt;&#x2F;th&gt;&lt;th&gt;Memory Recall@3&lt;&#x2F;th&gt;&lt;th&gt;Hybrid Recall@3&lt;&#x2F;th&gt;&lt;th&gt;RAG stale&lt;&#x2F;th&gt;&lt;th&gt;Memory stale&lt;&#x2F;th&gt;&lt;th&gt;Hybrid stale&lt;&#x2F;th&gt;&lt;&#x2F;tr&gt;&lt;&#x2F;thead&gt;&lt;tbody&gt;
&lt;tr&gt;&lt;td&gt;knowledge_update&lt;&#x2F;td&gt;&lt;td&gt;11&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;91%&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;27%&lt;&#x2F;td&gt;&lt;td&gt;82%&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;extension&lt;&#x2F;td&gt;&lt;td&gt;7&lt;&#x2F;td&gt;&lt;td&gt;93%&lt;&#x2F;td&gt;&lt;td&gt;96%&lt;&#x2F;td&gt;&lt;td&gt;58%&lt;&#x2F;td&gt;&lt;td&gt;n&#x2F;a&lt;&#x2F;td&gt;&lt;td&gt;n&#x2F;a&lt;&#x2F;td&gt;&lt;td&gt;n&#x2F;a&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;expiry&lt;&#x2F;td&gt;&lt;td&gt;7&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;0%&lt;&#x2F;td&gt;&lt;td&gt;43%&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;tr&gt;&lt;td&gt;&lt;strong&gt;overall&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;25&lt;&#x2F;td&gt;&lt;td&gt;98%&lt;&#x2F;td&gt;&lt;td&gt;&lt;strong&gt;99%&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;84%&lt;&#x2F;td&gt;&lt;td&gt;100%&lt;&#x2F;td&gt;&lt;td&gt;&lt;strong&gt;17%&lt;&#x2F;strong&gt;&lt;&#x2F;td&gt;&lt;td&gt;67%&lt;&#x2F;td&gt;&lt;&#x2F;tr&gt;
&lt;&#x2F;tbody&gt;&lt;&#x2F;table&gt;
&lt;&#x2F;div&gt;
&lt;figure&gt;
&lt;img src=&quot;eval.jpg&quot; alt=&quot;Elephantus eval tab: the results table above terminal-style bar charts of Recall@3 and stale-fact rate for RAG, memory and hybrid&quot; loading=&quot;lazy&quot; decoding=&quot;async&quot;&gt;
&lt;figcaption&gt;The same results in the UI&#x27;s eval tab.&lt;&#x2F;figcaption&gt;
&lt;&#x2F;figure&gt;
&lt;p&gt;What the numbers show:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Memory matches RAG on recall and cuts stale facts from 100% to 17%.&lt;&#x2F;strong&gt; RAG finds the right text every time, but it also returns the outdated version every time. That gap is what this project is about.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Expiry: 0% stale.&lt;&#x2F;strong&gt; Facts like “exam tomorrow” are filtered out at query time once they expire. RAG has no notion of time.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;The LLM judge caught all 11 updates&lt;&#x2F;strong&gt;, including “switching to Puma” vs “loves Adidas”, which share no words.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Most of the remaining 27% on updates is a strict metric.&lt;&#x2F;strong&gt; The flagged memories &lt;em&gt;describe&lt;&#x2F;em&gt; the change (“User’s Adidas sneakers broke after a month”) rather than restating the old fact. They mention the old keyword without a new one, so the keyword check counts them as stale. Judged by hand, none of them is outdated.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Hybrid mode is the weak spot at k = 3.&lt;&#x2F;strong&gt; Raw chunks, including stale ones, compete with memories for the same three slots.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;An offline run swaps the LLM for a rule-based stand-in and hash embeddings. It reaches only 68% recall and a 33% stale rate, which shows how much of the quality comes from the LLM judge and a real embedding model.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;design-decisions&quot;&gt;Design decisions&lt;&#x2F;h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;SQLite + FTS5 with brute-force NumPy vectors.&lt;&#x2F;strong&gt; One file and no services. At demo scale, brute-force cosine search is instant and avoids a vector-database dependency. WAL mode lets the API, MCP and UI processes share the file.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;fastembed (ONNX) instead of sentence-transformers.&lt;&#x2F;strong&gt; The same &lt;code&gt;bge-small&lt;&#x2F;code&gt; model without PyTorch, so the install is much smaller and faster.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Safe fallbacks for bad LLM output.&lt;&#x2F;strong&gt; Lenient JSON parsing, one corrective retry, then per-item validation. An unusable relation becomes &lt;code&gt;NEW&lt;&#x2F;code&gt;, so a fact is never lost or wrongly retired.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Short aliases for the judge.&lt;&#x2F;strong&gt; Candidate memories are shown as &lt;code&gt;m1&lt;&#x2F;code&gt;, &lt;code&gt;m2&lt;&#x2F;code&gt;… instead of UUIDs, so the model copies IDs reliably.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Relative expiry.&lt;&#x2F;strong&gt; The LLM outputs &lt;code&gt;expires_in_hours&lt;&#x2F;code&gt; rather than timestamps. Models handle “tomorrow ≈ 48h” more reliably than absolute dates, and the engine converts it using the message’s (possibly simulated) time.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Nothing is ever deleted.&lt;&#x2F;strong&gt; &lt;code&gt;UPDATES&lt;&#x2F;code&gt; flips an &lt;code&gt;is_latest&lt;&#x2F;code&gt; flag, expiry and forgetting are filters, and explicit forget is a soft delete. That keeps the full history for the graph and the decision log.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;A fair side-by-side.&lt;&#x2F;strong&gt; Both chat answers use the same model and prompt, so only the retrieved context differs. The question is stored only &lt;em&gt;after&lt;&#x2F;em&gt; answering, so the RAG baseline can’t retrieve the question itself.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;One client for OpenAI, Azure and Ollama.&lt;&#x2F;strong&gt; All three expose an OpenAI-compatible endpoint, so one small class covers them.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;A static web UI instead of Streamlit.&lt;&#x2F;strong&gt; Hand-written HTML, CSS and JS with no build step, served by the same FastAPI process. It uses only the public API, so it doubles as a working example of it.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;testing&quot;&gt;Testing&lt;&#x2F;h2&gt;
&lt;p&gt;The pytest suite needs no API key or downloads: the LLM is replaced by a scripted fake. It covers storage and container isolation, chunking, extraction (including malformed output), the sneaker sequence (&lt;code&gt;UPDATES&lt;&#x2F;code&gt; &#x2F; &lt;code&gt;EXTENDS&lt;&#x2F;code&gt; &#x2F; &lt;code&gt;DUPLICATE&lt;&#x2F;code&gt;), hybrid search, expiry with simulated time, the profile, forgetting, the REST API, the MCP tools, and the web UI, including a real-browser run of the sneaker flow with Playwright. An optional live test runs end to end against a real provider.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;limitations&quot;&gt;Limitations&lt;&#x2F;h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Small scale by design.&lt;&#x2F;strong&gt; Brute-force vector search and a single SQLite file suit demos and personal use, not millions of memories.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Quality depends on the LLM.&lt;&#x2F;strong&gt; Small local models may mislabel relations or miss expiry hints, as the gap between the offline and real evaluation runs shows.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;One relation per fact.&lt;&#x2F;strong&gt; A fact that both updates one memory and extends another is simplified to a single label.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;Expiry is set once&lt;&#x2F;strong&gt;, at extraction time. “The exam moved to Friday” creates a new fact that updates the old one, rather than rescheduling it.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;The evaluation is keyword-based&lt;&#x2F;strong&gt; and checks retrieval, not answer quality, on a small hand-written dataset.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;stack&quot;&gt;Stack&lt;&#x2F;h2&gt;
&lt;p&gt;Python 3.11+, FastAPI, SQLite (FTS5), NumPy, fastembed (&lt;code&gt;BAAI&#x2F;bge-small-en-v1.5&lt;&#x2F;code&gt;), the official MCP Python SDK, and Anthropic &#x2F; OpenAI &#x2F; Azure AI Foundry &#x2F; Ollama as LLM providers. MIT licensed.&lt;&#x2F;p&gt;
</content>
        
    </entry>
</feed>
