Quick check: does your agent stack have a memory.md in it somewhere? An AGENTS.md? A notes file the agent appends to when something seems worth keeping?
Thought so. Mine did too. It's the pattern everyone converges on, it takes twenty minutes to build, and it genuinely works — right up until the day it hands a customer a fact that stopped being true in March.
This post does three things: shows you exactly why the pattern rots (with a real-shaped sample file we'll dissect), gives you a small script to audit your own file tonight, and walks through the architecture change that actually fixes it. No vendor required for any of it.
The pattern we all built
Strip away the framework and every self-managed memory loop looks like this:
MEMORY = Path("memory.md")
def run_task(task: str) -> str:
context = MEMORY.read_text() # 1. dump everything in
result = llm(SYSTEM + context + task) # 2. do the actual job
note = llm( # 3. agent grades its own homework
"What from this interaction is worth remembering? "
"Reply with one line, or NONE.\n\n" + result
)
if note.strip() != "NONE":
with MEMORY.open("a") as f: # 4. append forever
f.write(f"- {note.strip()}\n")
return result
Enter fullscreen mode Exit fullscreen mode
Be fair to it first: this is human-readable, versionable, greppable, zero-infrastructure. For one agent, one job, small working set — it's honestly hard to beat.
Now look at what it doesn't do. Step 4 is the entire lifecycle. Nothing in this loop ever updates, merges, expires, or questions a line once written. The file has exactly one behavior: it grows.
Dissecting a six-month-old memory file
Here's a condensed, realistic slice of what that loop produces by month six. Read it the way retrieval reads it — every line equally true:
- Customer Acme runs their workload in us-east1
- Acme prefers Slack over email for escalations
- Acme's staging env uses the legacy auth flow
- Acme contact is Priya (prefers email)
- The feature flag `beta_router` must stay ON for Acme
- Acme migrated to europe-west2 for compliance
- Acme staging is on the new auth flow as of the migration
- Reminder: beta_router was removed from the codebase
- Acme runs in us-east1 (confirmed)
Enter fullscreen mode Exit fullscreen mode
Nine lines. Let's count the damage:
-
Two live contradictions.
us-east1vseurope-west2(lines 1, 6, 9 — and note the stale fact got re-confirmed after the migration, because the agent trusted its own notes). Legacy vs new auth flow (lines 3, 7). -
One zombie.
beta_routermust stay ON... for a flag that no longer exists (lines 5, 8). - One soft conflict. Slack vs email preference (lines 2, 4) — maybe both true (channel vs person), maybe not. Nothing will ever decide.
- Zero dates, zero sources, zero supersession. Line 9 looks more authoritative than line 6. It's newer and it says "confirmed." It's also wrong.
This file isn't remembering. It's accumulating — and retrieval over an accumulation returns everything the agent ever believed, in every version, including the dead ones. Which is how an agent politely, fluently tells a customer their data lives in a region they left months ago.
Audit your own file (runnable, ~40 lines)
Run this against your memory file. It flags undated entries, near-duplicates, and candidate contradictions (negation pairs and entries sharing a subject):
import re, sys, itertools
from difflib import SequenceMatcher
from pathlib import Path
lines = [l.strip("- ").strip() for l in Path(sys.argv[1]).read_text().splitlines()
if l.strip().startswith("-")]
print(f"{len(lines)} memory entries\n")
# 1. Undated entries (no ISO date, no month name)
dated = re.compile(r"\d{4}-\d{2}|\bJan|Feb|Mar|Apr|May|Jun|Jul|Aug|Sep|Oct|Nov|Dec", re.I)
undated = [l for l in lines if not dated.search(l)]
print(f"[FRESHNESS] {len(undated)}/{len(lines)} entries carry no date. "
f"None of these can expire.\n")
# 2. Near-duplicates (same fact, drifted wording)
for a, b in itertools.combinations(lines, 2):
if SequenceMatcher(None, a.lower(), b.lower()).ratio() > 0.7:
print(f"[DUPLICATE?]\n {a}\n {b}\n")
# 3. Contradiction candidates (shared subject, different claims)
def subject(l):
words = re.findall(r"[A-Z][a-z]+|\b\w+_\w+\b", l)
return words[0].lower() if words else None
by_subject = {}
for l in lines:
by_subject.setdefault(subject(l), []).append(l)
for subj, group in by_subject.items():
if subj and len(group) > 2:
print(f"[REVIEW '{subj}'] {len(group)} entries share this subject "
f"— read them side by side:")
for g in group:
print(f" - {g}")
print()
Enter fullscreen mode Exit fullscreen mode
Run it on the sample above and it flags the us-east1/europe-west2 cluster, the duplicate region claims, and reports 9/9 entries undated.
But here's the part that matters more than the script: notice what it can't do. It can surface that lines 1 and 6 share a subject. It cannot tell you which one is true — that requires knowing that a migration supersedes a location, that "confirmed" from a self-referencing agent means nothing, that line 8 kills line 5. Detecting contradiction candidates is string matching. Resolving them is judgment. Judgment is not a regex. Judgment is a job.
Which is the whole point.
The job nobody in your stack has
Go back to the loop at the top. Steps 3–4 quietly assign the working agent a second occupation: memory manager. Every "is this worth keeping?" decision is paid in tokens and attention out of the task budget — and the curator is grading its own homework, keeping whatever felt important mid-task. Multiply by a fleet of agents, each with its own file, and you have N diaries, zero reconciliation, and no component anywhere whose job is to notice when they disagree.
Write out what the job actually requires and it stops looking like a file API:
| Verb | Engineering requirement | Your memory.md
|
|---|---|---|
| Curate | Decide keep vs. noise, with criteria, outside the task loop | The busy agent, mid-task |
| Reconcile | Detect + resolve contradictions across all agents' knowledge | Nobody |
| Consolidate | Merge duplicates, supersede stale facts, keep the estate compact | Nobody |
| Brief | Deliver per-task relevant context — not the whole file |
read_text() (the whole file) |
| Provision | Bootstrap a new agent with what the fleet already knows | Copy-paste, if you remember |
Five verbs, mostly unstaffed. That's not a storage gap. It's a staffing gap.
The architecture change
The fix is division of labor: pull the management loop out of the working agents and give it to a specialist — a Memory Agent whose entire function is managing the memories of the other agents. The working agents' interface collapses to two calls:
# working agent — memory is no longer its problem
def run_task(task: str) -> str:
context = memory_agent.brief(agent_id=ME, task=task) # relevant slice only
result = llm(SYSTEM + context + task)
memory_agent.report(agent_id=ME, interaction=result) # raw material, not decisions
return result
Enter fullscreen mode Exit fullscreen mode
# memory agent — runs its own loop, on its own budget, fleet-wide
while True:
curate(inbox) # keep vs. noise — with criteria, not vibes
reconcile(estate) # cross-agent contradictions surfaced & resolved
consolidate(estate) # merge, supersede, expire
# brief() / provision() served on demand
Enter fullscreen mode Exit fullscreen mode
Note what changed. The working agent spends zero tokens on memory decisions. Curation criteria live in one place instead of N prompts. Contradictions have an owner. And the memory estate becomes a first-class system you can scope per-agent, audit, and export — instead of a text file with commit history.
Try it honestly
The honest carve-out first: one agent, small working set, no compliance surface → keep your markdown file. Sincerely. This architecture earns its keep at fleet scale — several agents, shared knowledge, contradictions that reach users.
If that's you, two paths:
- Build the loop yourself. The pseudocode above is genuinely enough to start — the hard parts you'll hit (contradiction resolution policy, decay, briefing relevance) are exactly the parts that make this a discipline. Instructive either way.
-
Run one that exists. This is the problem we work on: Memanto is an open-source (MIT, free) Memory Agent that runs alongside whatever agents and framework you already have — grab it at
github.com/moorcheh-ai/memantoand point it at your fleet. Its estate is stored in an open format (OKF) with a shippedmemanto migrateCLI, so trying it doesn't marry you to it.
Either way, run the audit script first. Then ask the one question your stack should be able to answer and probably can't:
What does your fleet believe today?
If the answer is a 3,000-line file — memory is a job, not a dump. Staff it.
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.