The procurement answer is narrower than most product language suggests: current NLP and machine-learning models can detect linguistic correlates of therapeutic alliance, but the best reported alliance prediction in the peer-reviewed studies is only around r = 0.20, with no external validation and no prospective evidence that acting on the model output improves patient outcomes. In the largest study, therapist speech features predicted client-rated alliance at Spearman's rho = 0.15, and the automatic speech recognition pipeline had a 36.4% word-error rate before the model ever saw the language it was supposed to analyze.[1]
That is enough to make transcript analytics scientifically interesting. It is not enough to treat an AI score as a reliable monitor of the psychiatrist-patient relationship, especially if the score would affect supervision, quality review, payment, or how a clinician behaves in the next session.

| Question for procurement | What the peer-reviewed evidence shows |
|---|---|
| Study design | Observational or naturalistic studies; no randomized trials |
| Sessions analyzed | From 28 dyads in the smallest study to 1,235 sessions in the largest study [1][3] |
| Best alliance prediction | About r = 0.20 in Lalk et al.; rho = 0.15 in Goldberg et al. [1][2] |
| Transcript source | Session transcripts or automatic speech recognition outputs, depending on study |
| External validation | None reported across these studies |
| Prospective outcome data | None; no study shows that ML-based alliance monitoring improves clinical outcomes |
| Speech-to-text reliability | 36.4% word-error rate in the largest study's ASR pipeline [1] |
What would count as measuring the relationship?
Therapeutic alliance is not a decorative psychotherapy concept. A major meta-analysis found that alliance predicts treatment outcome at about r = 0.28, which is large enough to justify clinical attention and small enough to remind everyone that alliance is not the whole treatment.[4] The construct matters because it captures something patients and clinicians can often recognize before a chart review does: whether the work feels collaborative, credible, and tolerable enough to continue.
But the clinical importance of alliance does not automatically validate a transcript-derived proxy. A monitoring tool has to answer a different question: whose alliance rating is being approximated, how much of that rating is captured, and what action follows when the score changes? A model that weakly correlates with client-rated alliance may help researchers search for linguistic signals. It should not be silently promoted into a substitute for instruments such as the Working Alliance Inventory or Session Rating Scale.
This is where the phrase "AI monitoring of psychiatrist-patient relationships" becomes risky. Monitoring implies repeated measurement, interpretable change, and some threshold for response. The evidence base described here does not yet provide those pieces. It provides retrospective associations between language features and alliance ratings.
The largest study found a real but weak signal
Goldberg et al. is the study that should receive the most weight in a committee discussion because it is the largest of the three: 1,235 sessions, 386 clients, and 40 therapists. The investigators used machine learning and natural language processing features from psychotherapy sessions to predict client-rated alliance. Using therapist speech features alone, the model predicted alliance at Spearman's rho = 0.15, with p < .001.[1]
The p value says the signal is unlikely to be random in that dataset. It does not say the signal is strong enough for clinical monitoring. A correlation of 0.15 leaves most of the variation in patient-rated alliance unexplained. In practical terms, many sessions with acceptable alliance would be expected to look linguistically similar to sessions with weaker alliance, and many sessions flagged by a model could reflect something other than a relationship problem.
The 36.4% word-error rate matters just as much as the model statistic.[1] Transcript-based monitoring depends on the words being available in the first place. If more than one in three words is wrong in the upstream automatic speech recognition output, the downstream model is not analyzing the session as it occurred. It is analyzing a damaged representation of the session.

That problem is not merely technical housekeeping. Alliance-relevant language is often subtle: hedging, repair, shared formulation, hesitation, disagreement, and the small conversational moves by which patients test whether the clinician has understood them. If the transcript is noisy, the model may preserve enough signal for a statistically significant paper and still be too unreliable for a supervision dashboard.
Goldberg et al. is therefore best read as evidence that therapist language contains measurable alliance-related information, not as evidence that alliance can be continuously monitored with acceptable accuracy. The study does not provide external validation across other sites or populations, and it does not test whether clinicians who receive model feedback make better decisions or achieve better outcomes.[1]
Lalk et al. points to a different, stronger use case
Lalk et al. analyzed 552 German psychotherapy transcripts from 124 patients using BERTopic neural topic modeling. Alliance prediction remained modest: r = 0.20 from therapist topics and r = 0.12 from patient topics.[2] The therapist-topic result is the strongest alliance prediction in the studies summarized here, but it is still a weak measurement signal if the intended use is monitoring an individual therapeutic relationship.
The more striking finding is not the alliance result. It is that symptom severity prediction was stronger, reaching r = 0.45 from patient topics.[2] That contrast matters because it suggests transcript NLP may be better suited to detecting what patients talk about and how symptom burden appears in language than to estimating the relational quality of the session.

That distinction is easy to lose in vendor-facing language. Adoption of transcript analytics for one measurement target does not validate it for another. A model may find symptom-relevant topics because patients explicitly discuss sleep, anxiety, mood, substance use, conflict, or impairment. Alliance is more relational and more evaluative. It may be expressed in what is said, but also in timing, rupture, repair, expectation, culture, power, and what the patient decides not to say.
Lalk et al. is useful because it does not simply add a second modest alliance estimate. It helps narrow the claim. NLP may have a more plausible near-term role in adjunctive symptom tracking or research phenotyping than in replacing patient- or clinician-rated alliance instruments. The alliance signal is present, but it is not yet a dependable measurement layer for the relationship itself.
Ryu et al. is mechanistic, not validation evidence
Ryu et al. studied 28 dyads and identified specific language markers associated with therapeutic alliance. Therapists' use of "we" was negatively correlated with alliance, while patient non-fluency, including hesitations and filler words, was positively correlated with alliance.[3] Those findings are intellectually interesting because they move beyond a black-box score and ask which features of session language might carry relational information.
They also show why interpretation has to be cautious. A therapist saying "we" could reflect collaboration in one clinical context and premature joining in another. Patient hesitations could reflect anxiety, careful disclosure, cognitive load, shame, trust, or ordinary conversational style. The computational marker does not become a clinical meaning without context.
The sample size also limits what the study can carry. Twenty-eight dyads can support exploratory signal finding; it cannot validate a monitoring system for broad psychiatric care. Ryu et al. belongs in the evidence base as an early mechanistic study, not as proof that automated relationship monitoring is ready for deployment.[3]
Why r = 0.15 to 0.20 is not a monitoring-grade result
The central measurement issue is not whether the correlations are statistically significant. It is whether they are accurate enough for the implied decision. A monitoring system changes the clinical environment. Someone has to decide whether a low score prompts a supervisor review, a clinician nudge, a documentation burden, a patient outreach message, or no action at all.
At r = 0.15 to 0.20, the score is too noisy to be treated as a trustworthy estimate of alliance for an individual session. It may rank groups weakly or help researchers identify features associated with alliance ratings. It does not provide the kind of patient-level precision usually expected before a measure is used in clinical quality monitoring.
Alliance ratings also tend to suffer from ceiling effects: many patients rate alliance highly. That compresses the variance a model can learn from and makes clinically meaningful deterioration harder to detect. A weak statistical signal in a compressed rating distribution may still be useful for research, but it is a poor foundation for automated alerts.
There is also a direction-of-measurement problem. A transcript model trained against client-rated alliance is not the same thing as clinician-rated alliance, observer-rated alliance, or rupture detection. If the target is the patient's experience, then the patient rating remains the nearer measure. If the target is supervisory review, the system needs evidence that its output improves supervisory judgment, not just that it correlates weakly with a questionnaire.
The missing studies are the ones deployment would require
All three studies are retrospective, cross-sectional, observational, or naturalistic in the relevant sense: they analyze existing session language and compare model outputs with ratings or clinical variables. None is a prospective trial of AI alliance monitoring. None shows that giving clinicians or supervisors an automated alliance score improves retention, symptom outcomes, rupture repair, safety, satisfaction, or any other patient-centered endpoint.
External validation is also absent in the evidence summarized here. That matters because psychotherapy language is not a single stable measurement environment. Population, diagnosis, acuity, modality, therapist style, culture, language, recording quality, and institutional setting can all change the distribution of words and the meaning of conversational features. A model that learns alliance-associated language in one setting may not preserve calibration in another.
For a governance committee, the relevant gap is not simply that the models could be improved. The gap is that the clinical action pathway has not been validated. A score that makes a therapist self-conscious, changes documentation behavior, or causes supervisors to focus on false-positive relationship concerns can create burden without benefit. A score that misses actual ruptures may create false reassurance.
What is evidence-supported now?
The evidence supports transcript NLP as an adjunctive research signal. It can help investigators search for linguistic markers, compare language patterns across sessions, and generate hypotheses about how alliance is expressed in clinical conversation. Goldberg et al. shows that a weak alliance signal can be extracted at scale; Lalk et al. shows that symptom severity may be more detectable than alliance; Ryu et al. shows that specific features such as pronoun use and non-fluency can be studied mechanistically.[1][2][3]
The evidence does not support replacing validated alliance measures. It also does not support representing current transcript models as prospective monitors of psychiatrist-patient relationships. If a health system wants to pilot these tools, the defensible framing is quality-improvement or research with safeguards, not routine clinical measurement.
- Use model outputs as exploratory signals, not as stand-alone alliance scores.
- Keep patient- or clinician-rated instruments as the measurement reference when alliance is the clinical concern.
- Require site-specific validation before any score enters supervision, quality review, or operational reporting.
- Treat speech-to-text performance as part of the clinical measurement problem, not as an IT implementation detail.
- Do not infer outcome improvement from retrospective correlation.
A procurement-ready verdict is therefore simple: candidate adjunctive research signal, not replacement for patient- or clinician-rated alliance measures, and not an evidence-supported prospective monitoring tool until externally validated models and outcome-impact studies exist.
References
- Machine learning and natural language processing in psychotherapy research: Alliance as example use case, Journal of Counseling Psychology, 2020.
- Measuring Alliance and Symptom Severity in Psychotherapy Transcripts Using Bert Topic Modeling, Administration and Policy in Mental Health, 2024.
- A natural language processing approach reveals first-person pronoun usage and non-fluency as markers of therapeutic alliance in psychotherapy, iScience, 2023.
- The alliance in adult psychotherapy: A meta-analytic synthesis, Psychotherapy, 2018.