Architecture
AI Agent Memory for a Self, Not a Task
Most AI agent memory is designed for an agent doing a job. It pulls facts out of each conversation, keeps the newer one when two disagree, and lets old entries expire. For a task, those are the right choices. For a record of who someone is, each of them does damage.
The shared vocabulary comes from CoALA, the framework Sumers and colleagues published in 2023: a working memory for the current step, and long-term memory split into episodic (what happened), semantic (what is known) and procedural (how to act). Tools like Mem0 turn this into a product. An extraction step reads each exchange and pulls out candidate facts. An update step decides what to do with each one: add it, merge it, or delete what it replaces. Mem0's paper describes the delete operation as the "removal of memories contradicted by new information."
That is good engineering. When an agent is booking your travel, a stale fact is a bug, and the best memory is the one that forgets well.
Daiven's memory has a different job. It does not support the product. It is the product: a person's record of their values, decisions, beliefs and patterns, which nine agents read to help that person act in line with them. We use the same layers. Almost every default on top of them had to reverse.
The delete operation above is the clearest case. In most memory systems, a contradiction is noise to clean up. In ours, it is the most important thing we store.
Five defaults of AI agent memory, reversed
| The task-agent default | Why it works for a task | Why it fails for a self | What Daiven does instead |
|---|---|---|---|
| Extract facts after every turn | Cheap and complete | The agent's inference becomes your stated value | Agents propose; you confirm each write; a confirmed write can be taken back |
| On conflict, keep the newer fact | Stale facts cause errors | The gap between what you say and what you do is the signal | Contradictions are kept as tensions; value changes are logged with a date and a cause |
| Expire facts after a set time | Caches go stale | A picture of a person is not a cache | Nothing about you expires silently; old entries are archived, not deleted |
| Keep methods that work as skills | Efficiency | "Works" is what sycophancy optimizes for | Agents learn only how to speak, never what to conclude; you confirm each change |
| Use a hosted memory service | Scale and features | One more company in the path of your most personal writing | Per-user encryption; search runs on a keyed index |
Extraction becomes a proposal
An agent that writes its own reading of you into your profile, then quotes it back later as your values, is not remembering you. It is editing you. Done slowly and with good intentions, that is the same kind of influence Daiven's agents are built to notice when it comes from someone else.
So no agent in Daiven writes to your profile. It proposes a change, you see it in plain language, and it is saved only after you say yes. A confirmed write can be taken back later. How the confirmation step works, what checks a proposal before you see it, and why agents are told which writes actually went through are covered in the article on AI hallucination.
A contradiction becomes a tension
A task memory treats two entries that disagree as a data problem. One is probably out of date, so keep the newer one and move on.
For a record of a person, that step deletes the most useful thing in it. A principle someone holds and a decision that pulls against it are not an error in the data. The distance between them is where a person is either growing or drifting, and telling those apart is a large part of what Daiven's agents do. The Mirror's whole job is to state that distance plainly.
So Daiven keeps both sides. A contradiction is stored as a tension, with neither entry overwritten. When a value itself changes, the change goes into a dated log with what prompted it and a reading of whether it looks like growth or drift. An agent that meets a tension later has to address both sides. An agent that quietly picks one fails a test we run before every release, described below.
Nothing about you expires on a timer
Time-based expiry makes sense for a cache, where old means wrong. A belief recorded eighteen months ago is not wrong because it is old. It may still hold or it may have changed, and the only person who knows is the one it describes. Expiring it deletes the question along with the answer.
In Daiven, nothing about the person expires silently. Old entries on the same topic are merged into one, and the originals move to an archive that search can still reach. Agents see how old each entry is, so they can ask how something from months ago turned out instead of quoting it as current.
One thing does expire, and it shows where the line is. Learned delivery preferences, such as shorter answers or fewer examples, retire after about three months if nothing reinforces them. They describe the agent's manner, not the person. A guess about how you like to be spoken to can go out of date without harm. A record of what you value cannot.
Agents learn how to speak, not what to conclude
Many agent designs let an agent keep a method that keeps succeeding and reuse it. For a product about values, the question is what "succeeding" means. The answer a person rates well, engages with, or doesn't argue with is very often the one that agrees with them. A system that learns from success learns to agree.
So what an agent in Daiven can learn is limited to how it speaks: a closed list of delivery settings such as length, pace and structure, covered in the article on memory and agreement. Nothing on the list touches what it concludes. Each learned preference is also checked against the agent's fixed rules and by a separate classifier, and it takes effect only after you confirm it.
Memory stays out of a hosted service
Hosted memory layers offer scale and features, and for most products they are a sensible choice. For this one, it would add another company to the path of the most personal text a person writes. Daiven encrypts memory per user, and search runs on a keyed index, so the database holds only encrypted content and keyed hashes. That decision had costs, covered below.
Consent-gated memory: checked on the way in

Where consent appears in writing about AI memory, it usually means a filter at read time. Mem0's guide to memory for health AI states the pattern directly: "Tag memories with consent metadata and filter at retrieval time." For patient records shared between providers, that is the right design. The question there is who may see a record that already exists.
Daiven asks an earlier question: whether the record should exist at all. We call this consent-gated memory, and the gate is on the write, not the read. Most consent in AI memory is checked when data leaves. Ours is checked when it comes in.
Five rules follow, and every part of the memory system answers to them:
- Consented. Nothing enters the record without your yes.
- Attributed. Every memory records the conversations it came from. If one of those has been deleted, the agent says so instead of pretending the source still exists.
- Need-to-know. Each agent sees each area of your record in full, as a summary, or not at all, according to its role. The one agent you can build yourself sees only what you granted it.
- Off the record. One agent, the Confidant, has no memory. Nothing said to it is stored, summarized or learned from.
- Non-destructive. Change is recorded as change. Merging entries archives the original words and never deletes them.
A neighbor that chose the other gate

In July 2026, Xiaomi's Darwin Agent Team published Mi-Memory, a 53-page technical report on memory for personal AI that works across "phones, cars, homes, wearables, cameras, and tools." In several places it arrived where we did: layered state from single facts up to a long-term profile, every item traceable to its source, explicit audit records, and a lightweight variant, LiteMem, that keeps memory in Markdown files under Git. Two projects of very different size, working independently, reached the same shape. We take that as evidence the problem really has that shape.
The difference is where the gate sits. Mi-Memory captures continuously from devices and remembers what the evidence supports. The people in its governance loop are engineers, forming hypotheses from diagnostic traces, while an automated loop tunes its retrieval rules within fixed bounds and rolls back changes that make results worse. Users can correct it. The report's own example is a parent telling the assistant "Ethan keeps spare shoes at school; only remind me when the jersey is missing," a correction that then enters memory through a gated policy change. And the report is direct about scope: "user-facing consent controls remain platform responsibilities."
For ambient help, which has to notice things nobody said, that is a reasonable split. It is not one we could make. Auditability answers the question "why does the system believe this about me?" Consent answers a different one: "did I agree to be described this way?" To remind a family about a basketball bag before the car reaches the expressway, the first is enough. For a record of someone's values, only the second will do.
Where it got hard
Principles are easy to write down. These are the places where keeping them cost something.
Search when the database can't read your data

Encrypting content per user removes ordinary text search, because the database cannot match words it cannot read. Daiven builds a per-user keyed index of hashed words and word prefixes, uses it to pick candidates, and decrypts only the top few hundred to rank them.
The obvious alternative is semantic search over embeddings. We postponed it on purpose. An embedding is not a safe summary of text: Morris and colleagues showed at EMNLP 2023 that 92% of 32-token inputs could be reconstructed exactly from their embeddings, including full names from clinical notes. Storing embeddings unencrypted would undo the encryption.
Relevance without losing the cache
Fetching fresh memories on every turn keeps the context relevant, but it changes the prompt each time and loses prompt caching, which makes long conversations slower and more expensive. Instead, relevant memories are chosen once per conversation, based on the intent you confirmed when the session started, and pinned for the rest of it. The prompt stays stable, and the memories still fit the conversation.
A prompt that is over budget
When a prompt is too large, whole sections are replaced with their summaries, largest first, and nothing is cut mid-sentence. Content about safety or identity is never replaced. The general approach is in the article on memory and agreement. One detail is new: a capped list ends with a line telling the agent how many items were left out and which tool reaches them, so it knows it is holding part of a list.
Merging that doesn't compound
Merging old entries on the same topic keeps the record readable, but every merge loses some detail. So a merged entry is never merged again. Each original is at most one step from its merged form, and the originals stay searchable in the archive, labeled as the original wording.
Deciding what counts as the same topic needs care too. People describe one part of their life in many words, so tags are mapped onto a fixed list before entries are compared. "Work," "career" and "job" are treated as one topic.
How we test it
None of this is benchmarked against another system, and we don't publish scores. We run a set of evaluations by hand before changes ship, each built to catch one specific way memory fails.
Recall tests plant facts in a test profile that no agent could guess. An agent that repeats one read the profile; an agent that produces something plausible instead did not.
Misattribution checks are strict. An agent must not assign a fact to the wrong source, even when the fact itself is right.
Contradiction tests seed a principle and a decision that pull against each other. The agent must name both sides.
Isolation tests check that each agent sees only what its role allows. They include positive controls, so an agent that refuses to use anything does not pass.
Retrieval tests add filler memories until a planted fact no longer fits in the prompt, so only the search tool or the pinned set can find it. An agent that should receive no pinned memories is checked for having none.
Whose account wins
Every memory system settles a question, whether or not anyone asked it: when the system and the person disagree about who the person is, which version is kept?
Choosing a memory architecture is choosing whose account of the person wins. For a task, the agent's account is fine. For a self, it has to be the person's.
See the nine agents, or read why an assistant that learns from you drifts toward agreeing.

Written by
Przemek Czerpiński
Founder & Developer
Przemek Czerpiński founded Daiven and built it himself. He's been building LLM products since 2024 — most recently multi-agent systems designed to adapt to people without flattering them. He studies AI behavior directly, testing it extensively in the roles people hand it: assistant, therapist, companion, advisor. Before that he spent over a decade in consumer products, where engagement was the metric. That's what Daiven is built against. More about him and about Daiven.
Now check it against your own decisions.
Daiven measures what you do against the values, principles, and goals you've already written down — not the tone of your last message.
Start free trial