AI
AI Hallucination: Why Models Make Things Up About You
On this page
- 1.Why a model invents anything
- 2.Two kinds of hallucination, and only one gets corrected
- 3.Don't make it guess who you are
- 4.Don't let it guess what you meant
- 5.A rule against guessing needs somewhere else to go
- 6.Stopping an error from having descendants
- 7.Old is not the same as wrong
- 8.You can read the whole thing
You ask an assistant to help you think through a job offer. It opens with this:
Given how much you've prioritized stability over the last few years, the lower base salary is the real problem here.
You never said that. And you are not certain it is wrong, which is the part worth stopping on. It took two things you had told it and produced a third that arrives sounding exactly like the other two.
That third thing is an AI hallucination. Not the famous kind — no fake citation, no invented statistic, nothing anyone could check. It made something up about you.
Why a model invents anything
No part of that is a failure of the model's reasoning. It was asked a question it could not settle from what it was holding, and it answered anyway.
And it answered anyway is the interesting half, and it has now been measured. In April 2026, Nature published work from OpenAI and Georgia Tech arguing that hallucination is less a defect in the weights than a product of how models are trained and graded. Their finding: dominant metrics like accuracy systematically reward guessing over admitting uncertainty. A model that says I don't know scores the same as one that answers wrong, and worse than one that guesses and gets lucky. These systems were optimized to be good test-takers, and a good test-taker never leaves a question blank.
So two conditions have to hold before a model invents. It has to be short of what it needs. And it has to be better off guessing than saying so.
The second condition is not ours to change. It is set upstream, in training and evaluation, an industry away from any product built on top of a model. The first condition is a product decision, start to finish.
That is why Daiven treats hallucination as an information problem rather than a model problem. Not because the model isn't involved. Because the model is the part we cannot reach.
Two kinds of hallucination, and only one gets corrected
Before any of the mechanisms, the boundary. It decides which of them are even relevant.
AI hallucination gets discussed as though it were one failure. It is two, and they behave nothing alike.
The first is about the world. An invented citation, a date that is off by a year, a court ruling that does not exist. These are checkable. Someone opens the link and it goes nowhere. The error is embarrassing and it is also self-limiting, because reality is sitting right there holding the answer.
The second is about you. Your values, your history, why you chose what you chose, what you said in March, how you behave under pressure. Nothing external checks that. You are the only instrument available, and you are being handed the claim at the exact moment you were hoping to be understood — which is the worst possible moment to be evaluating a sentence for accuracy.
| About the world | About you | |
|---|---|---|
| What it invents | A citation, a date, a statistic, a quote | A value, a motive, a pattern, something you said |
| What checks it | The record. Anyone can look | You, in the moment, hoping to be understood |
| How it reads when it's wrong | Confidently incorrect | Almost right |
| What it costs to miss it | One bad answer | A premise every later answer inherits |
That last row is the whole difference. A wrong fact about the world stays wrong in one place. A wrong fact about you, in a system that stores what it learns, stops being an answer and becomes a starting condition.
Which sets the limit on what any of this can do. Daiven does not make the model underneath it more accurate about the left column. It will not stop an invented citation about monetary policy, and nothing here should be read as claiming it will. What it can do about the left column is make agents label it: mark which part of a response comes from your own record and which comes from general knowledge, so the two do not arrive looking alike. Everything else in this article is built against the right column.
That is also the column nobody instruments. Hallucination benchmarks measure facts, because facts can be scored. There is no benchmark for whether a system was right about a person, and there cannot be one, since the only ground truth is sitting on the other side of the screen.
Don't make it guess who you are
The standard fix for a model that guesses about you is to tell it more about you. That is half right in a way that makes it dangerous.
The half that's right: structure. Daiven keeps what it knows about you as separate kinds of record — who you are, what you value, how you behave, what you believe, the principles you hold, the habits you keep, what you're working toward, what you've decided, and what happened worth remembering. Each has its own shape, its own dates, its own connections to the others, and each is rendered into a request as clean labeled text.
Compare that to the common pattern: memory as one running document. A document hands the model an undifferentiated wall of words and asks it to work out what matters. That decision is itself a guess, remade from scratch on every turn, and it degrades as the document lengthens. Daiven made those decisions once, at schema level. The storage side of it, including what gets merged and what gets retired, has an article of its own.
The half that's wrong is volume. More context is not more care.
Each Daiven agent gets a defined slice of your profile and nothing beyond it: some sections in full, some as a summary, some not at all. The Mirror reads your values, your patterns and your decisions in full, because holding stated values against actual choices is the entire job. The Witness knows your name and your situation and close to nothing else, because receiving what you say without an agenda does not require your decision history. The intake step runs on summaries across the board.
This is not rationing. Irrelevant context measurably degrades a model's reasoning: Shi and colleagues at ICML 2023 took a grade-school math benchmark, added a single irrelevant sentence to each problem, and watched accuracy fall sharply. One sentence, on problems the models could otherwise solve. Give a model material with no bearing on the question and it will find a pattern in it anyway. That is what pattern-matchers do.
So the agent reading a message someone sent you does not also hold your goals for the year. Precision beats volume, and the two are frequently in opposition.
Don't let it guess what you meant
The other guess happens earlier, and you never see it.
Every Daiven session starts with the same agent, and it never answers the question. Its stated function is to refuse to pass a request downstream while it is still ambiguous. It produces a specification — the subject, where you stand on it, what you are actually asking for, the context that bears on it, the emotional weight it carries, and which agent should take it. You confirm that or correct it. Nothing runs until you do.
The line in its own instructions is the one the rest of the product is built around: ambiguity is debt. An unresolved ambiguity does not evaporate when a request moves downstream. It gets resolved anyway, by a model, silently, and from that point on the guess is load-bearing. The second rule follows from the first: never infer when you can ask. Two plausible readings means you pick, not the model.
This has a cost and the product treats it as one. Intake is turn-budgeted: a soft nudge to converge, and a hard ceiling that does not consult the model at all. Asking instead of assuming is the mechanism. It is not a license to interview you.
The same instinct runs through the agents themselves. Ask a thin question and you get an answer aimed at the median of everyone who has ever asked something like it. For one specific person that lands as a near-miss dressed up as insight, which is worse than being asked a question, because it closes the matter instead of opening it. So the Mirror asks rather than supplying a motive. The Judge establishes facts rather than presuming them. The Sentinel asks about the relationship and what is at stake before it will read the message.
The version of this that operates below meaning
When you paste something in for analysis — a message, an email, the paragraph you have now read eleven times — the specification carries it verbatim, and the platform checks it byte for byte against what you actually typed.
That check exists because models tidy. Ask one to reproduce a long quote and it will normalize the spacing, smooth an awkward clause, quietly fix what looks like a typo. In most contexts, harmless. For manipulation analysis it is fatal, because the agent would then be grading language the sender never wrote. And in that kind of reading, the signal is often in exactly the thing a model would smooth. The odd repetition. The missing word. The sentence that changes direction halfway through.
Daiven restores your bytes where it can and flags the drift where it can't. It never repairs by guessing.
A rule against guessing needs somewhere else to go
Every agent carries an explicit instruction against filling a gap it cannot fill. These are lines from the shipping agent definitions, not summaries of them, for instance:
- The Mirror: Infer from evidence; never assume the unstated.
- The Witness: Never fill in what they haven't said.
- The Sentinel: Read what's there; never assume what isn't.
- The Devil's Advocate: Never invent a study to win an argument.
- The Truth Teller: Accuracy over completeness. I will leave gaps rather than fill them with plausible guesses.
Read as a set they are one instruction: prefer a visible gap to an invisible fabrication. That is the trade a general assistant makes in the other direction, and it is not a matter of tone. A gap you can see is information. A gap filled well is indistinguishable from knowledge.
But an instruction not to guess is close to worthless on its own. Tell a model it may not invent, leave it with no other move, and it will invent anyway, politely, with a caveat attached. The rule only works if there is somewhere else for it to go.
There are two. Agents hold scoped tools against your own record: search what you have said, list your decisions, open a value and read its full history, pull up a goal. When the context loaded up front is not enough, the agent goes and looks instead of estimating. And where a section was loaded only as a summary, the agent can formally request the full version and receive it on the next turn.
That is the structural answer to the deepest cause of hallucination. A model with no way to say let me check will produce something rather than nothing. It will do it every time, and it will sound fine doing it. The stakes rise once an agent can act rather than only answer.
Stopping an error from having descendants
Everything so far lowers the odds of a wrong inference on a single turn. None of it touches what happens to the wrong ones that get through. That is the failure people actually recognize, because almost anyone who has used a memory-keeping assistant has been remembered wrong about something and found out weeks later, inside an answer that had quietly built on it.
So, in Daiven: when an agent decides something about you is worth keeping, it does not write it. It proposes it. You see the entry in plain language, and you confirm it, and only then does it become part of your profile.
The reason that gate matters more than any single-turn safeguard is arithmetic. In an ungated system, one wrong inference becomes a stored fact. Every agent in every later session then reads that fact and reasons from it correctly. Correct reasoning from a false premise is exactly the problem, because it manufactures further false premises, each one properly supported by the last. Error does not sit still. It has descendants.
The confirmation step is where that line ends. It does not catch every error. It makes sure an uncaught one has to get past you first.
Before a proposal reaches you, it passes a second, cheaper model whose job is to repair a malformed write. That model may add something the first agent missed, under one condition: it has to quote you. Evidence that does not appear word for word in what you actually wrote is discarded. And it is never shown your existing profile — deliberately, because a verifier that can see what is already there will start reasoning about what is missing, and a verifier that fills gaps has stopped verifying and started originating.
There is a smaller version of the same discipline that is easy to miss. Agents are handed the list of writes that actually went through, and told plainly that absence from that list proves nothing. So the system does not hallucinate about its own behavior either. An assistant that tells you it has remembered something when it hasn't has made the one error you have no way at all to catch.
Old is not the same as wrong
A record of a person is not a set of facts. It is a set of facts with dates attached, and the dates carry as much information as the facts do.
Three things follow. Every request is anchored with today's date and the model is told to treat it as now. Every dated entry arrives with its age attached, so a decision from March shows up as six months old rather than as an undated fact. And change is tracked on purpose rather than incidentally: values carry an evolution log with an assessment of whether the movement looks like growth or drift, and decisions come back for review at 30, 90, and 180 days.
The consequence is a distinction most systems cannot draw. An agent that knows a belief was recorded eighteen months ago can ask whether it still holds. One that doesn't will quote it back to you as current. Old information gets treated as old — not as false, and not as fresh.
One more surface belongs here, and it is one most products do not instrument at all: where a memory came from. Every entry records the conversations that produced it. So where did you get that? is a question with an answer, rather than an invitation to improvise one. Attribution is its own hallucination surface, and a fluent one: a system that cannot trace a claim will happily describe a conversation that never happened. Consolidation doesn't break the chain either; when related entries are merged, the originals are archived rather than deleted and the provenance carries across.
You can read the whole thing
Every mechanism above lowers a rate. Not one of them takes the rate to zero, and an article about hallucination that claimed otherwise would be performing the exact failure it describes.
Which is why the last one is different in kind. Your profile is not an opaque store the system consults about you. It is a page. You can open it, read what Daiven believes about you in the same plain language it was proposed to you in, change it, and delete from it. What each agent is for, and what each will not do, is on the agents page.
Hallucination you can see is hallucination you can correct. That is the entire reason the page exists, and it is the only safeguard that still works once all the others have failed.
Everything in this article is an attempt to make that page boring to read. It will not always be. When it isn't, it is one entry, and it is yours to delete.
See the nine agents, or read how Daiven compares to a general assistant.

Written by
Przemek Czerpiński
Founder & Developer
Przemek Czerpiński founded Daiven and built it himself. He's been building LLM products since 2024 — most recently multi-agent systems designed to adapt to people without flattering them. He studies AI behavior directly, testing it extensively in the roles people hand it: assistant, therapist, companion, advisor. Before that he spent over a decade in consumer products, where engagement was the metric. That's what Daiven is built against. More about him and about Daiven.
Now check it against your own decisions.
Daiven measures what you do against the values, principles, and goals you've already written down — not the tone of your last message.
Start free trial