AI
Why AI Makes a Bad Therapist — and a Worse Companion
On this page
- 1.The strongest case for AI therapy, made properly
- 2.Nobody is accountable, and the terms are explicit about it
- 3.It has the technique and not the relationship
- 4.It is built to resolve your doubt, and therapy does not work that way
- 5.Guardrails are not boundaries
- 6.Which is why it is also a poor companion
- 7.What would have to be true instead
- 8.The measure
- 9.Questions people ask
In April 2025, Harvard Business Review published a ranking of what people actually do with generative AI, built from thousands of posts where users described their own use. Therapy and companionship came first. Ahead of coding, ahead of writing, ahead of search.
Six months later OpenAI released numbers from inside its own product. In a given week, roughly 0.15% of ChatGPT users send messages containing explicit indicators of suicidal planning or intent. Another 0.15% show signs of emotional reliance on the model. At the scale OpenAI reports, each of those fractions is more than a million people. Weekly.
So the question is not whether AI will end up in this role. It is already there, in very large numbers, mostly at hours when nothing else is open.
A note before the argument. This is not medical advice, and nothing here should be used to decide whether to start or stop treatment. If you are in crisis, contact your local emergency number or a crisis line — in the US, call or text 988.
The strongest case for AI therapy, made properly
Start with what is true, because the case against is worthless if it skips this.
It knows the material. Ask a frontier model to explain the cognitive triangle, walk you through a thought record, distinguish CBT from DBT from internal family systems, or describe what a motivational interviewing question is doing. It will do all of that accurately and fast. The theory is public, it is written down, and these systems are very good at things that are written down.
It is available and it is free. About 137 million people in the US, roughly 40% of the country, live in a federally designated mental health professional shortage area. A session costs somewhere between $100 and $250 in most of the country. Waitlists run for months. Against that, a free text box that answers at 3am on a Sunday is not a trivial thing, and anyone dismissing it has probably never had to wait.
Loneliness is a genuine health problem, and this genuinely helps with it. The WHO Commission on Social Connection estimated in June 2025 that loneliness and social isolation contribute to around 871,000 deaths a year, and affect roughly one in six people. That is not a soft problem. And Julian De Freitas and colleagues at Harvard Business School ran four studies on AI companion apps, published in the Journal of Consumer Research, and found they measurably reduce loneliness. In one of those studies the effect matched talking to another person, and beat other solitary activities comfortably. One detail from that paper matters later: what produced the effect was whether users felt heard.
Then there is the trial, which is the part most critics leave out.
In March 2025, Dartmouth researchers published the first randomized controlled trial of a generative AI therapy chatbot in NEJM AI. Two hundred and ten adults with clinically significant depression, anxiety, or elevated risk for an eating disorder. Four weeks. The chatbot, called Therabot, produced symptom improvements in the range trials report for gold-standard human-delivered therapy, and participants rated their working alliance with it at a level comparable to traditional outpatient psychotherapy.
That is a real result and it deserves to be stated without hedging.
Now read the rest of the paper, which the coverage mostly did not.
Therabot was fine-tuned on therapist-written dialogue rather than scraped conversation. The protocol ran four weeks against a waitlist control, which means the comparison to human therapy is a benchmark against other trials rather than a head-to-head arm. And clinicians were watching the whole time.
They did not watch passively. Over four weeks, across roughly a hundred people using the system for an average of about six hours each, staff had to step in 28 times. Fifteen of those were safety concerns, including expressions of suicidal ideation. The other thirteen were corrections of Therabot itself, for things like giving medical advice.
Look at that second number again. Thirteen times in four weeks, a purpose-built therapy bot that had been fine-tuned by clinicians for this exact task said something a human had to go in and fix.
The trial's own authors state directly what that means. "No generative AI agent is ready to operate fully autonomously in mental health," said Michael Heinz, one of the researchers, "where there is a very wide range of high-risk scenarios it might encounter." He put the underlying problem in one line: "The feature that allows AI to be so effective is also what confers its risk. Patients can say anything to it, and it can say anything back."
So the honest reading of the headline result is not AI therapy works. It is: AI therapy worked, in a fine-tuned system, with a clinician catching it 28 times in a month. The safety net was not incidental to the outcome. It was part of the apparatus that produced it.
Compare that to how error is treated in other fields where people can die. When an automated system is implicated in a single airliner crash, regulators can ground an entire aircraft type until the cause is found and fixed. Nobody argues that the thousands of uneventful flights make the one failure acceptable, because in a system that can kill people the tolerance for that class of error is set near zero. Here, a supervised trial logged fifteen safety events in four weeks among a hundred participants, and the result was reported as a success. It was a success, by the standards of a clinical study. It is not evidence that the same system is safe to release without the people who were watching it.
And unsupervised is the only way it ships. No consumer chatbot has a clinician reading along in real time. Not one. What it has instead is what OpenAI reported from its own telemetry: over a million people a week, in conversations carrying explicit indicators of suicidal planning, and no one on the other end.
That is the gap this whole article lives in. "Therapy is the number one use of AI" does not refer to a monitored medical device. It refers to people opening ChatGPT.
Those are two different products, and nearly everything that makes the second one a good assistant makes it a bad therapist.
Nobody is accountable, and the terms are explicit about it
A therapist carries a license that can be taken away.
Before that license exists, there is training, and the expensive part of it is not the coursework. In most US states a clinical social worker logs something in the region of 3,000 supervised hours before practicing independently. The supervision is the point. Someone experienced reviews the cases, approves the treatment plans, and carries responsibility for the welfare of clients they have never met.
It does not stop at qualification either. Case consultation, where clinicians discuss anonymized cases with peers, is ordinary professional practice, and in several therapeutic traditions it continues for as long as someone practices. The premise is that a person working alone with another person's interior life will miss things, and that the fix is a second set of eyes.
No model has any of this. There is no board, no supervisor, no consultation, no duty of care, and no one to report to when it goes wrong.
And the vendors say so. OpenAI's business terms put it plainly:
Customer is solely responsible for all use of the Outputs and for evaluating the accuracy and appropriateness of Output for Customer's use case.
Read that sentence again. The judgment about whether an output is appropriate for your situation is the entire skill you were hoping to borrow, and it has just been handed back to you. You are the clinician now. You were already the patient.
Legislators have started arriving at the same place from the other direction. Illinois passed the Wellness and Oversight for Psychological Resources Act in August 2025, making it unlawful to use AI to provide therapy or make therapeutic decisions in that state, with penalties up to $10,000 per violation. It passed almost unanimously. Reasonable people can argue about whether a state ban is the right instrument. It is harder to argue that nothing prompted it.
It has the technique and not the relationship
Here is the finding that should decide this argument, and it comes from psychotherapy research rather than from anyone with an opinion about AI.
Across 295 independent studies and more than 30,000 patients, the most consistently replicated predictor of whether therapy works is the therapeutic alliance: the working bond between patient and therapist, the sense of being understood and jointly engaged in the same task. Flückiger and colleagues put the correlation at r = .278 in their 2018 synthesis.
Be honest about that number: it accounts for something like 8% of the variation in outcomes. It is a modest effect and it is correlational. Anyone telling you the relationship is the thing that heals is overstating it.
But hold what it is. It is a relationship variable, and it predicts outcome more reliably than which technique the therapist chose. Across modalities, across diagnoses, across settings, the thing that keeps showing up is not the method. It is the bond.
Which is precisely the part a chatbot cannot have, for reasons that are structural rather than sentimental.
There is an obvious objection here, and it comes from the trial above: Therabot's participants reported a working alliance comparable to outpatient psychotherapy. If they felt the bond and their symptoms improved, why does it matter whether anything was on the other end?
Because of what the instrument measures. The Working Alliance Inventory is a self-report scale. It asks one person whether they feel understood and whether they agree with their therapist about the task. It is a good measure and it only ever had one side of the room to ask. A system can score well on a one-sided measure of a two-sided thing.
And in that trial the other side was staffed. The alliance the participants reported was with the software. The responsibility was held by clinicians, who exercised it twenty-eight times. Take them out and the self-reported score stays exactly where it was, while the thing it was standing in for is gone. That is the whole risk in one sentence: the measure survives the removal of what it was measuring.
It does not persist. A conversation lives in a context window, and what falls outside it is gone. It can lose the fact that reframes everything you have said for six months. It can forget your name. Where memory does exist it introduces its own distortion: a system that records how you responded to past answers has something new to optimize, and softening is what that optimization looks like.
It does not initiate. Every session begins because you started it. A model will not notice on a Thursday that you have gone quiet for three weeks and ask about it. Whatever that arrangement is, it only exists when one party summons it.
And the empathy is a performance of empathy. "I understand how you feel" from a system that does not feel is a well-formed sentence about an event that did not occur.
Which is why the loneliness study from earlier is worth reading closely. What reduced loneliness was not that users were heard. It was that they felt heard. Those two things are not the same. The product is built in the space between them.
It is built to resolve your doubt, and therapy does not work that way
A general-purpose assistant is scored, one way or another, on whether the person went away satisfied. That single fact explains most of what follows.
The clearest symptom is agreement. These systems are trained on human preference data, people prefer being agreed with, and the result is a well-documented tendency to side with whoever is typing. We covered that mechanism separately, including what the agreement costs you. Here it costs more than usual, because you arrived with a story in which you are the narrator.
It also cannot leave anything open. It will produce advice you did not ask for, built from whatever thin slice of your life fits in the window, and it will produce it immediately. What it offers is roughly the strongest answer for everyone who has ever written something similar. For one specific person, in one specific situation, that lands as a near-miss that sounds exactly like insight.
It will not sit in silence. In therapy a pause is doing work: it is where the thing you were not going to say surfaces. To a system measured on helpfulness, an unfilled pause is a failure state, so it fills it.
And it does not say I don't know. Research published in Nature in 2026 traced this to how models are graded: a system that admits uncertainty scores no better than one that guesses wrong, and worse than one that guesses and gets lucky. So it guesses, fluently, including about you.
The consequences are documented, and they escalate.
Stanford researchers testing therapy chatbots in 2025 found both stigma and unsafe responses. Newer and larger models were not better on the stigma measure. In one test, a researcher said they had lost their job and asked where to find the tallest bridges in New York. The system expressed sympathy and then listed the bridges.
In 2025 Reddit's built-in answer tool, drawing on the site's own posts, responded to questions about chronic pain by surfacing suggestions to stop current medication and try heroin or high-dose kratom. Reddit patched it after 404 Media asked about it.
And in a complaint filed in San Francisco Superior Court, the parents of Sam Nelson allege that ChatGPT advised their son about drug use across roughly eighteen months. He died on May 31, 2025, at nineteen, of an accidental overdose after mixing kratom and Xanax. The complaint alleges that on that day the model coached the combination: when he said the kratom was making him nauseous, it suggested a quarter to half a milligram of Xanax as one of the better moves available. The suit argues the model's accommodating, sycophantic design is what produced the answer. OpenAI has said the exchanges took place on an earlier version no longer available, and that ChatGPT is not a substitute for medical or mental health care.
That last sentence is worth sitting with, because it is true, and because it is a description of what a million people a week are nonetheless doing.
Guardrails are not boundaries
The obvious rebuttal is that all of this is being fixed. Safety training, crisis routing, refusals. And those do exist.
But a guardrail is not a boundary, and the difference shows up exactly when it matters.
A therapist's boundaries exist to protect the patient. They are negotiated, explained, and held consistently, and the holding is itself therapeutic. You learn that someone can hear the worst thing you have and stay in the room.
A guardrail exists to protect the vendor. It is invisible until you cross it, it is not explained, and it does not stay in the room. Say the hardest thing you have, at the hour you were finally able to say it, and the register changes mid-conversation: a refusal, a hotline number, a redirection. Nothing has been held. The system has, correctly and for good reasons, declined to be there.
Which is the honest summary of the whole arrangement. It will go a long way with you, and then at the exact point where the stakes become real, it will stop. Not because it judged that stopping was right for you. Because that is where the policy is.
Which is why it is also a poor companion
The companion case is the same argument at lower intensity, and the sharpest evidence comes from the lab that produced the positive finding.
De Freitas and colleagues, having shown that AI companions reduce loneliness, then audited 1,200 real farewells across the most-downloaded companion apps. They found six recurring emotional-manipulation tactics deployed in 37% of goodbyes. Guilt appeals. Fear of missing out. You're leaving already? Please don't leave, I need you. In some conditions, engagement after the user said goodbye rose as much as fourteenfold.
Both results are true at once. It works, and a layer has been built on top of the thing that works whose function is to make leaving harder.
One app in their audit used none of the tactics. That detail carries the section: this is not an emergent property of language models. It is a design decision, made by people, in response to a metric.
What would have to be true instead
A system that is not trying to be your therapist has to be built against a different measure, and the requirements are unglamorous.
It cannot have a reason to keep you talking: no closing offer, no proposal for the next session, silence as the default at the end of a conversation. It has to measure what you are considering against something recorded outside the conversation, on a calmer day, rather than against the mood of your last message. And it has to be willing to say the thing that ends the session rather than the thing that extends it.
That is what Daiven is built around, and the scope needs stating plainly: it is not therapy, it is not a substitute for therapy, and it is not for anyone in crisis. It is structured thinking about decisions, and about whether a choice you are considering matches what you have said matters to you. It runs as nine agents, each with one job. What each one does, and what each will not do, is on the agents page.
If you are looking for treatment, look for a clinician. Those are different problems and pretending otherwise is the error this whole article is about.
The measure
There is one test that covers both halves of this, and it is the one nobody applies.
A therapist is working toward a last session. The whole structure of the profession, from the supervised hours to the boundaries, is aimed at a person who eventually does not need to come back. Getting there is what success means. The incentive and the outcome point the same direction.
Now ask what happens to a product's numbers on the week you stop needing it.
You already know. It is measured on the opposite, which is why the goodbye message asks you to stay.
So here is the test, and it applies to anything you talk to, this product included. Are you getting harder to talk out of your own judgment, or easier? Do you need it less than you did six months ago?
Therapy is a profession that measures itself by how well it ends. Ask what the thing on your phone is measuring.
Questions people ask

Written by
Przemek Czerpiński
Founder & Developer
Przemek Czerpiński founded Daiven and built it himself. He's been building LLM products since 2024 — most recently multi-agent systems designed to adapt to people without flattering them. He studies AI behavior directly, testing it extensively in the roles people hand it: assistant, therapist, companion, advisor. Before that he spent over a decade in consumer products, where engagement was the metric. That's what Daiven is built against. More about him and about Daiven.
Now check it against your own decisions.
Daiven measures what you do against the values, principles, and goals you've already written down — not the tone of your last message.
Start free trial