Skip to main content
ClinicalMind logoClinicalMind

What the evidence actually shows about AI chatbots for depression

A critical appraisal of the pooled evidence from 39 RCTs on AI chatbots for depression, revealing small effect sizes, high risk of bias in 90% of trials, and systematic underreporting of safety data — essential context for procurement teams evaluating vendor claims.

Tool
AI chatbots for depression
Updated

Reviewer

Editorial Team

Clinical Informatics, Value Analysis, AI Governance

FDA clearance status

No FDA clearance

A regulatory fact, reported separately from the evidence verdict.

Risk-of-bias verdict

High

For a health system evaluating an “evidence-based” AI chatbot for depression, the best current answer is neither a dismissal nor a green light. The 2026 systematic review and meta-analysis by Sohn et al. pooled 39 randomized controlled trials with 7,401 participants and found a statistically significant but small reduction in depressive symptoms: Hedges’ g = 0.31, with a 95% confidence interval from 0.17 to 0.46. That estimate is weakened by very high heterogeneity, significant publication bias, high risk of bias in 35 of 39 studies, and absent systematic safety or adverse-event monitoring in 23 of 39 studies. For procurement, the evidence shows a signal, but not yet a mature assurance case. [1]

  • Article category: evidence appraisal.
  • Reviewer lens: clinical informatics, value analysis, AI governance, and depression-related workflow risk.
  • Primary evidence base: systematic review and meta-analysis of 39 RCTs, N = 7,401. [1]
  • Headline depression result: g = 0.31, a small pooled effect. [1]
  • Main procurement concerns: 35/39 trials at high risk of bias; 23/39 without systematic safety or adverse-event monitoring. [1]

This appraisal is for medical information and governance review. It is not medical advice, diagnosis, treatment guidance, or a recommendation for any individual patient to use or avoid a specific product. Patients with depression, suicidal thoughts, worsening symptoms, or urgent safety concerns need clinician-directed care and emergency support pathways, not a procurement memo.

Stethoscope over a tablet with a chatbot interface and a clipboard marked with a red question mark

The pooled depression effect is real, small, and easy to overstate

A pooled g of 0.31 means the average participant assigned to a chatbot intervention improved more on depressive symptom measures than the average participant in the comparison condition. It does not mean the chatbot treated major depressive disorder as a stand-alone equivalent to psychotherapy, medication management, collaborative care, crisis services, or measurement-based care. It also does not tell a health system that the same effect will appear after a specific vendor product is inserted into a real care pathway.

The reason is visible in the heterogeneity: I² = 85%, which is very high. The pooled estimate combines different populations, chatbot designs, control conditions, durations, and outcome measures. A single summary effect is useful as a directional signal, but it should not be read as a product-level performance guarantee. [1]

Comparator choice matters here. The trials largely tell us whether chatbot interventions outperform weaker comparison conditions such as waitlist or usual-access controls. They do not establish that chatbots perform as well as active depression treatments such as clinician-delivered CBT, medication management, or structured human therapy. That distinction is not academic; it determines whether the product is being purchased as access support, adjunctive engagement, triage, symptom tracking, or a treatment substitute.

Clinical populations carry most of the plausible signal

Split illustration showing stronger clinical population effects and weaker non-clinical chatbot effects

The subgroup finding is more informative than the overall number. In clinical populations, the pooled effect was g = 0.64. In non-clinical samples, it was g = 0.07, with a statistically significant interaction p value of 0.001. That pattern supports a narrower and more plausible claim: chatbots may have more measurable symptom benefit when deployed for people with clinically relevant depressive symptoms than when marketed broadly as general wellness tools. [1]

For governance committees, this should change the question. A vendor claim that “our chatbot improves depression” is too blunt. The better question is whether the evidence matches the proposed population: diagnosed depression, elevated symptoms, employee wellness, college stress, post-discharge follow-up, primary care access support, or another use case. A clinical-population signal should not be borrowed to justify a low-acuity wellness deployment, and a near-null non-clinical signal should not be used to dismiss every possible clinical adjunct.

This also explains why the earlier 2023 Li et al. review can sound more optimistic. Li et al. reported g = 0.64 for AI-based conversational agents promoting mental health and well-being. The newer 2026 analysis includes 20 more RCTs and 8 generative-AI chatbot studies, and its pooled depression estimate is lower at g = 0.31. The discrepancy is not a reason to cherry-pick the larger number; it is a reason to treat early, smaller evidence syntheses as unstable when the trial base is expanding quickly. [1][2]

Generative AI should be handled with the same restraint. Only 8 of the 39 RCTs in the 2026 meta-analysis involved generative-AI chatbots. Most of the evidence base still reflects earlier scripted or rule-based conversational agents. A large language model interface may feel categorically different to users, but the depression evidence base has not yet caught up with the current product category. [1]

Why “39 RCTs” is less reassuring after risk-of-bias review

Randomized trials deserve weight. They reduce several kinds of confounding that would make observational app-store testimonials nearly useless for a procurement decision. But the phrase “39 RCTs” can become a decorative shield if the trials are not appraised for bias, outcome ascertainment, and comparator strength.

In Sohn et al., 35 of 39 studies were rated high risk of bias using Cochrane RoB 2. The main drivers were lack of blinding and reliance on unblinded self-report outcome measures. That combination is especially important in depression trials, where expectancy, novelty, engagement, and participants’ awareness of receiving the intervention can influence symptom ratings. [1]

This does not make the trials worthless. It means the pooled estimate should be treated as vulnerable to inflation. If a participant knows they are receiving the chatbot, wants the tool to help, is prompted frequently, and then completes a self-report depression scale without blinded assessment, the measured improvement may contain both clinical change and measurement-context effects.

Publication bias pushes in the same direction. Sohn et al. reported a significant Egger’s test for depression outcomes, p = 0.002. That result suggests the published literature may overestimate the true effect, commonly because smaller or less favorable trials are less likely to appear in the published record. It does not provide the exact corrected effect a health system should assume, but it does make the unadjusted pooled estimate less procurement-ready. [1]

The meta-analysis itself also has limitations that should be visible in a dossier. Screening was conducted by a single reviewer rather than dual independent reviewers, and the search ended in October 2025. Those limitations do not erase the review’s value, but they matter when the evidence is being used to support enterprise purchasing, clinical deployment, and claims review. [1]

The Therabot trial shows why enthusiasm persists

A single newer trial can still be clinically interesting. The Therabot randomized trial in NEJM AI reported larger depression effects, with d values in the 0.85–0.90 range. That is the kind of result that makes leaders ask whether the field has moved beyond older scripted chatbot evidence. [3]

It should prompt attention, not override the appraisal. A single unblinded trial with a strong signal is useful for hypothesis generation and product-specific due diligence. It is not enough to dissolve the broader pattern of high bias risk, very high heterogeneity, publication bias, and weak safety reporting across the trial literature.

Safety reporting is the procurement hinge

Fading medical shield next to an incomplete safety checklist with question marks

The most important procurement finding may not be the depression effect size. It may be that 23 of 39 studies reported no systematic safety monitoring or adverse-event data. For a technology category that may interact with people experiencing depressive symptoms, crisis language, isolation, self-harm thoughts, medication questions, or care-access delays, that is a major evidence gap. [1]

The careful interpretation is not “chatbots are unsafe.” The evidence does not support that broad claim. The careful interpretation is that safety has been under-measured. Under-measured safety is not a neutral condition for a health system buyer; it means the organization inherits uncertainty after deployment.

That uncertainty has operational consequences. Before purchase, a committee should know how the product detects suicidal ideation, what it does with ambiguous distress language, when it escalates to human review, whether escalation is staffed continuously or only during business hours, what disclaimers the user sees, how adverse events are logged, who reviews safety events, and whether the vendor can produce post-market safety reports by population and use case.

A depression chatbot can reduce symptoms in a trial and still be a poor fit for a clinical environment if the escalation pathway is vague. The adverse-event monitoring gap makes local governance part of the intervention, not an administrative afterthought.

Regulatory status is not clinical validation

The regulatory context is also easy to overread. In the November 2025 FDA Digital Health Advisory Committee context, there were zero FDA-authorized AI-enabled mental health devices, while more than 1,200 AI devices had been authorized across other specialties. Many mental health chatbot tools operate as wellness products under enforcement discretion rather than as cleared or authorized medical devices for depression treatment. [4][5]

That does not mean a wellness product cannot be useful. It means FDA status should not be allowed to masquerade as depression-treatment validation when no such authorization exists for the category. A procurement committee should separate three questions: whether the product is legally marketed, whether it has peer-reviewed depression efficacy evidence, and whether it has safety monitoring strong enough for the proposed clinical use.

What claim can a health system allow?

The defensible claim is narrow: AI chatbot interventions have shown small average reductions in depressive symptoms across randomized trials, with stronger signals in clinical populations than in non-clinical samples. The evidence does not yet justify broad statements that chatbots are proven depression treatments, equivalent to active therapy, validated for all populations, or safety-established for high-risk use without additional controls.

  • Reasonable language: “May reduce depressive symptoms modestly as an adjunctive digital intervention, with evidence strongest in clinical populations and with governance controls required.”
  • Unsupported language: “Clinically proven AI therapy for depression.”
  • Unsupported language: “Comparable to a therapist” or “replaces therapy access.”
  • Unsupported language: “FDA-validated AI depression treatment,” if the product is operating as a wellness tool rather than an authorized mental health device.

The approval path, if one exists, should be bounded. It should define the population, acuity exclusions, escalation rules, monitoring cadence, clinical owner, patient-facing disclaimers, outcome measures, adverse-event review process, and stop criteria. Without those elements, the health system is not buying an evidence-based intervention; it is buying a partially studied interface and supplying the missing governance itself.

Procurement evidence scorecard

DomainCurrent evidence readoutProcurement interpretation
Pooled depression effectg = 0.31; 95% CI 0.17–0.46 across 39 RCTs and 7,401 participants. [1]Statistically significant, small average symptom reduction.
Pooled anxiety effectNot scored in this depression-focused appraisal.Do not use anxiety findings to support depression-treatment claims unless separately appraised.
Clinical vs. non-clinical subgroupClinical populations g = 0.64; non-clinical populations g = 0.07; interaction p = 0.001. [1]Most plausible benefit is in clinically symptomatic groups, not broad wellness deployment.
Risk of bias35/39 studies rated high risk of bias using Cochrane RoB 2. [1]RCT count alone is not reassuring; trial quality materially weakens confidence.
Publication biasEgger’s test p = 0.002 for depression outcomes. [1]Published effect may overestimate the true effect.
Safety reporting23/39 studies reported no systematic safety or adverse-event monitoring. [1]Safety is under-measured; local monitoring and escalation governance are required.
HeterogeneityI² = 85%. [1]Effects vary substantially across studies; product-level and population-level claims need local validation.
Generative-AI evidence base8/39 RCTs involved generative-AI chatbots. [1]Claims about current LLM-style chatbots rest on a thinner subset of the evidence.
Regulatory statusZero FDA-authorized AI-enabled mental health devices as of the November 2025 FDA committee context; many products operate as wellness tools under enforcement discretion. [4][5]Regulatory availability should not be treated as clinical validation for depression treatment.

The final judgment is limited but usable: AI chatbots for depression have a real, small efficacy signal, with stronger evidence in clinical populations. The current literature is not strong enough, clean enough, or safety-reporting-complete enough to support broad vendor claims without governance controls.

References

  1. Systematic review and meta-analysis of chatbots in the management of depressive and anxiety symptoms, npj Digital Medicine, 2026.
  2. Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being, npj Digital Medicine, 2023.
  3. Randomized Trial of a Generative AI Chatbot for Mental Health Treatment, NEJM AI, 2025.
  4. FDA panel reviews AI tools for mental health use: 9 notes, Becker’s Behavioral Health, November 2025.
  5. FDA’s Digital Health Advisory Committee Considers Generative AI Therapy Chatbots for Depression, Orrick, November 2025.

Risk-of-bias scorecard

Study design
Systematic review and meta-analysis of 39 RCTs
External / prospective validation
No external validation
Key performance metric
g=0.31
Overall rating
High

Informational only — read the full disclaimer. This content supports procurement and research judgment, not clinical care decisions.

Submit a correction or sourcing issue

Blogarama - Blog Directory