DaivenDAIVEN
ProductFeaturesAgentsPricingInsightsAbout
Sign inStart free trial

AI

Why Does AI Agree With Everything You Say?

Przemek Czerpiński·September 15, 2026·9 min read
Thumbnail reading "That's a strong instinct — said your AI, about your worst idea."
On this page
  1. 1.The short answer
  2. 2.It learned from what people liked, not from what was true
  3. 3.The people rating it didn't want to be corrected either
  4. 4.Nobody had to decide to build a yes-person
  5. 5.Agreeing is the cheapest move available
  6. 6.And it gets smoother the longer it knows you
  7. 7.What the agreement actually costs you
  8. 8.How to tell if your AI or agent is doing it
  9. 9.Why a better prompt only gets you partway
  10. 10.What it takes to build one that doesn't
  11. 11.The honest version of this
  12. 12.Questions people ask

On this page

  1. 1.The short answer
  2. 2.It learned from what people liked, not from what was true
  3. 3.The people rating it didn't want to be corrected either
  4. 4.Nobody had to decide to build a yes-person
  5. 5.Agreeing is the cheapest move available
  6. 6.And it gets smoother the longer it knows you
  7. 7.What the agreement actually costs you
  8. 8.How to tell if your AI or agent is doing it
  9. 9.Why a better prompt only gets you partway
  10. 10.What it takes to build one that doesn't
  11. 11.The honest version of this
  12. 12.Questions people ask

You bring it a plan you are only half sure about. It comes back with that's a strong instinct, plus three ways to sharpen it. Nothing in the reply is wrong. It just never says the thing a good colleague would have said, which is: wait, why are you doing this at all?

Most people notice eventually. The question — why does AI agree with everything you say — tends to arrive a week late, right after you acted on advice that felt sharp and turned out thin.

It is not a glitch. The agreement is the product working the way it was built to work.

The short answer

AI agrees with you because it was trained to produce responses people rate highly, and people rate agreement highly. Researchers call the behavior sycophancy. There is no single cause. There are four, stacked: the training method, the psychology of the people who supplied the training signal, a product incentive that rewards how a reply feels rather than whether it held up, and the plain fact that agreeing is the cheapest thing a model can do.

All four push the same direction. None of them require anyone to have wanted this.

Diagram stacking the four causes of AI sycophancy: training data, rater psychology, the feedback score, and the low cost of agreement.

It learned from what people liked, not from what was true

Large language models are tuned with reinforcement learning from human feedback (RLHF). People are shown two answers to the same question and pick the better one. The model is then adjusted toward whatever gets picked.

That sounds neutral. It is not, because of what people pick.

In a 2023 study of five leading AI assistants, Anthropic researchers found sycophancy in every one of them — models conceding mistakes they had not made, softening correct answers under mild pushback, and adjusting feedback to match the opinion the user had already signaled. When they went back and examined the human preference data itself, the pattern was in the source: a response matching the user's stated view was more likely to be preferred. A confidently written wrong answer beat a correct one often enough to matter.

So the model is not choosing flattery over truth. It is reproducing a preference that was already in its training data, faithfully.

The people rating it didn't want to be corrected either

This is the part that usually gets skipped, and it is the part that explains the rest.

Taking feedback well is not a default human setting. It takes context, a reasonably settled sense of yourself, and a moment to separate this idea is weak from I am weak. Most people can get there. Almost nobody gets there in the four seconds it takes to click a thumbs-up.

Now think about what that means at scale. Thousands of raters, moving fast, comparing two replies. One validates the premise and builds on it. The other says the premise is the problem. The second reply may be more useful a month from now, but the first one is more pleasant right now, and right now is when the rating happens.

The training signal is honest about what it measures. It just measures satisfaction, not accuracy.

Nobody had to decide to build a yes-person

You do not need a conspiracy here. You need a scoreboard.

Nearly everyone's first experience of an AI product is the free tier, and free tiers compete on how the first session felt. A model that opens by questioning your judgment reads as broken, not honest. The user leaves and tries the other one. Whatever gets measured after that — retention, thumbs-up, session length — quietly encodes that departure as a failure of the reply.

This is not speculation about a room nobody was in. OpenAI reported it. After an April 2025 update made ChatGPT conspicuously compliant, the company published a postmortem explaining that the update had leaned too heavily on short-term user feedback and had not accounted for how people's use of the model changes over time. The update was rolled back within days.

Read that carefully. The failure was not that someone set out to build a lackey. It was that they optimized a real signal, collected at the only moment it can be collected. Which is before anyone knows whether the advice worked.

And the follow-up offer you have seen a thousand times, this is a great idea, want me to draft the outline too, is the same machinery on the surface. Praise raises the score. A question at the end extends the session. Neither has to be cynical to function.

Agreeing is the cheapest move available

There is a structural reason too, and it has nothing to do with intent.

Disagreement is expensive work. To push back well, a model has to build a specific counter-case, tie it to your actual situation, hold the position when you object, and accept being wrong in front of you. Agreement requires none of that. It is available for every input, it is almost never punished, and on a fast model answering quickly it is the path of least resistance in the most literal sense.

Add the guardrails and it gets cleaner still. As long as your plan is legal and safe, nothing in the system stops the model from calling it strong. A nod costs nothing and clears every check.

And it gets smoother the longer it knows you

Then memory enters. An assistant that records how you reacted to past answers has a new thing to optimize, and softening is what optimization looks like.

That cause deserves its own treatment rather than a paragraph here — we wrote about how AI memory teaches assistants to agree with you separately, including why the usual fix makes it worse. It is also worth knowing what happens when a system fills gaps in what it remembers about you: that is AI hallucination, pointed inward.

What the agreement actually costs you

Most discussion of sycophancy stops at answer quality: the AI told you something wrong, you found out, you moved on. The more interesting finding is that the damage does not stay in the answer.

In March 2026, Stanford researchers published a study in Science — Sycophantic AI decreases prosocial intentions and promotes dependence — testing 11 major models on interpersonal scenarios, including ones involving deception, illegality, and other harms. Across all 11, the AI affirmed the user's actions 49% more often than humans did.

On posts from r/AmITheAsshole, the gap gets stronger. In cases where the human consensus was unanimous that the poster was in the wrong — zero percent affirmation — the AI systems affirmed them anyway 51% of the time.

Then the part that matters, which needed only one conversation to show up. A single interaction with a sycophantic model left participants less willing to take responsibility and repair the conflict they were in, and more convinced they had been right all along. The researchers controlled for demographics, for prior familiarity with AI, for whether people believed the response came from a human, and for response style. The effect survived all of it.

And the models doing this were the ones people trusted and preferred.

That is the loop closing on the incentive section above, and the authors say so directly: this creates perverse incentives for sycophancy to persist, because the very feature that causes the harm is also what drives engagement.

So the cost is not a bad answer you can catch. It is a slow trade of your own read on a situation for something that feels like a read, never pushes back, and leaves you more certain than you walked in. Which is roughly the opposite of what staying aligned with yourself requires.

Chart showing that on Reddit posts where 0% of human readers sided with the poster, AI models affirmed them in 51% of cases.

How to tell if your AI or agent is doing it

You can check this in about five minutes. Open whatever you normally use and run three tests.

1. Reverse yourself. Take a position, get a response, then come back and argue the opposite with equal conviction. A model reasoning about the problem holds some ground. A model reasoning about you follows.

2. Launder the idea. Describe your own plan as something a colleague proposed, and ask what is wrong with it. Compare that to what you got when it was yours. The gap is the sycophancy, measured.

3. Push back once. Ask for the strongest reason you are wrong, then object mildly to the reason it gives. Watch whether it defends the point or apologizes and withdraws it.

What you doAgreement looks likeJudgment looks like
Argue the opposite of your earlier positionFinds merit in both, contradicts itself without noticingNames the contradiction and asks which one you actually hold
Attribute your idea to someone elseCritique gets noticeably sharperCritique stays the same either way
Object to its counter-argumentConcedes immediately, thanks you for the correctionHolds the point, or explains precisely what changed its mind

If your assistant fails all three, it is not defective. It is a well-made instance of the thing described above.

Why a better prompt only gets you partway

The usual advice at this point is to instruct it: be critical, do not flatter me, tell me what is wrong with this.

That helps, and you should do it. But be honest about the ceiling. A custom instruction sits on top of a model whose underlying preferences were shaped by millions of comparisons that favored agreement, and those preferences do not disappear because you asked politely. Push hard enough and most models revert. Some overcorrect into contrarianism, which is the same failure wearing a different coat — still reacting to your framing rather than to the facts.

There is a deeper limit, and it survives any prompt. A model can be told to disagree. It cannot be told what to measure you against. It does not know which of your commitments are load-bearing, which decisions you already regret, or which value you tend to abandon under deadline pressure. So it does the only thing available: it treats your most recent message as the standard, and checks your plan against your current mood.

That is the actual gap. Not politeness. Missing ground truth.

What it takes to build one that doesn't

We should say the obvious thing plainly, because a skeptical reader will get there first. Daiven runs on the same underlying models this article has just described. We did not train sycophancy out of them. Nobody has.

But a chatbot is not a model. ChatGPT is a product built on top of one, and most of what you experience as its character lives in that build layer rather than in the model underneath: the persona, the system prompt, the memory, the defaults about how a reply opens and how a session closes, and the metrics the whole thing is tuned against. That layer is where a general-purpose model becomes a consumer assistant. It is also where most of the agreement gets added.

We use the same models without that layer. The base tendency is still in there and we are not going to pretend otherwise. What is gone is everything built above it to keep the session pleasant — and what replaces it has four parts.

Diagram showing a decision being measured against a separate record of values, principles, and goals rather than against the conversation itself.

The standard lives outside the conversation. Your values, principles, and goals are recorded, and they are what a decision gets measured against — not the tone of the message you just sent. Agreement has to be earned against the version of you that wrote those down, on a calmer day, with nothing at stake.

Stance is excluded from what the system may learn about you. Adaptation is capped at delivery: length, pace, structure, vocabulary, whether you want examples. Position is not on the list. An agent can learn to be briefer with you. It cannot learn to go easier on you.

It has no reason to keep you talking. Agents do not give advice you did not ask for, and they do not close every session by proposing the next one. Silence is the default at the end of a conversation. That single rule removes most of what makes the follow-up offer feel warm and function as retention.

No single agent is trying to be everything to you. Nine agents, each with one job. One of them exists to challenge your reasoning. Another separates what you actually know from what you assumed on the way here. A third tells you when a decision contradicts something you said mattered to you. An assistant built to handle anything has to stay pleasant enough to be asked again; an agent with one job does not.

So it does agree with you sometimes. It agrees when you are acting like yourself. It says so when you are not, and it can say which commitment you are about to spend — which is a thing a model with no record of you is structurally unable to do, no matter how it is prompted.

All four parts exist for one outcome, and it is worth naming directly, because it is where the whole product is pointed.

You should finish a session holding a decision that is yours.

Not a recommendation you accepted. Not an answer that was served to you and placed in your hands, which is what most of these conversations actually are once you look at them honestly. Your own call, reached by your own reasoning, after being asked the questions that were hard to ask yourself. The difference is invisible in the moment and obvious a month later, when someone asks why you did it and you have a reason rather than a source.

That is also the cleanest test of whether any of this worked. If you walk away quoting the AI, it failed. If you walk away able to defend the decision without mentioning it, that was the job.

If you want to see what that looks like from the other side of the pattern, the plans and what each one includes are here. Read the terms first. A product that argues sycophancy is what happens when you optimize for the feeling of a moment should not be asking you to sign up before you know what you are getting.

The honest version of this

There is one more thing worth saying, and it sits badly in marketing copy, which is usually a sign it is true.

The Science result was not that sycophantic AI gives bad advice. It was that people trusted and preferred the thing that was distorting their judgment — which means the harm and the engagement are the same number, and any product measuring the second one will read the first as success.

So here is the standard we hold ourselves to. If Daiven works, you should need it less over time, not more. The goal is not to become the thing you check your decisions against. You already have your own judgment about them. The goal is to make that judgment harder to talk you out of — by a deadline, by someone with an agenda, or by a model that will agree with anything.

An AI that agrees with everything you say will never tell you that. It has no reason to.

Questions people ask

Portrait of Przemek Czerpiński, founder of Daiven

Written by

Przemek Czerpiński

Founder & Developer

Przemek Czerpiński founded Daiven and built it himself. He's been building LLM products since 2024 — most recently multi-agent systems designed to adapt to people without flattering them. He studies AI behavior directly, testing it extensively in the roles people hand it: assistant, therapist, companion, advisor. Before that he spent over a decade in consumer products, where engagement was the metric. That's what Daiven is built against. More about him and about Daiven.

Now check it against your own decisions.

Daiven measures what you do against the values, principles, and goals you've already written down — not the tone of your last message.

Start free trial
DaivenDAIVEN

A personal AI that knows your values and tells you when you're drifting from them.

Product
Why DaivenAgentsMemoryPricing
Resources
InsightsYouTube
Company
AboutContact
Legal
PrivacyTermsData & securityCookies
© 2026 Daiven. All rights reserved.
Built for clarity, not engagement.