Ask most clinicians which part of a diagnostic assessment is the least reliable, and structured interviews are rarely the first answer. We were trained to treat standardized diagnostic interviews (SDIs) — instruments like the SCID, DIVA, MINI, CIDI, DIS, and the various substance-use interview schedules — as the gold standard against which every self-report questionnaire gets measured. A new systematic review and meta-analysis in JAMA Network Open should make us less confident in that assumption. The finding, in short: even our best structured interviews are only moderately reliable, and nobody can yet fully explain why the reliability swings so much from study to study. That has real implications for how we think about the reliability of diagnosis. In everyday busy clinical practice, clinicians are typically not doing so-called gold standard, standardized diagnostic interviews. Clinicians are typically using their intuition and experience — but these methods are known to be even less reliable than standardized protocols.
Reliability of Standardized Diagnostic Interviews
Xie, Nordgaard, Sheldrick and colleagues (McMaster University, University of Copenhagen, and UMass Chan Medical School) looked at the results of 57 studies. They estimated pooled test-retest reliability for common adult mental disorders and substance use disorders (SUDs), then tested which study characteristics explained the variation between studies. It’s methodologically serious work.
How Eeliable are Standardized Psychometric Interviews?
Overall pooled reliability was kappa 0.69 (95 percent confidence interval, 0.66 to 0.72), “substantial” by the conventional benchmark of 0.61 to 0.80, but built on enormous heterogeneity between studies.
Substance use disorders were significantly more reliable than mental disorders overall. Within mental disorders, reliability ranged from kappa 0.55 for nonaffective psychoses up to kappa 0.74 for bipolar disorder, with anxiety, depressive, and personality disorders clustered around kappa 0.63 to 0.64.
Put plainly: even our best structured interviews, administered by trained interviewers days apart on the same person, land somewhere between “fair” and “excellent” agreement. And for some of the most common presentations in outpatient practice, we sit closer to the fair end.
We Can’t Explain the Variance
This is the part I think matters most for practice. For mental disorders, none of the tested moderators — which specific interview was used, diagnostic criteria version, sample size, retest interval, clinical versus nonclinical setting, retest-aware versus independent administration, or reported response rate — reached statistical significance. Heterogeneity remained at 94 percent after accounting for all of them. For substance use disorders, only diagnostic criteria version explained meaningful variance, a 16 percent reduction, with later diagnostic manuals outperforming earlier ones. Structural standardization alone may not be sufficient to ensure consistent psychiatric diagnosis.
That’s a notable admission from inside the diagnostic-reliability research tradition — a checklist executed identically by two different interviewers on two different days does not guarantee the same diagnostic conclusion, because so much of what determines the answer lives in things a reliability statistic doesn’t capture: how a question was asked, how much space was given for a hesitant answer, whether the interviewer picked up on something the patient mentioned in passing three questions earlier.
What This Means for a Real Patient
A reliability level in that range isn’t just a statistical curiosity — it translates into a meaningful chance that two independent, competent assessments of the same patient a week apart would disagree on the diagnosis. In a single assessment that disagreement is invisible. Across a series of assessments — intake, review, medico-legal report, insurance or NDIS determination — it surfaces as a real clinical and administrative problem: different diagnoses, different treatment plans, different funding decisions, for the same person, produced by the same “gold standard” method.
Where Psychometric Measurement Fits
None of this makes validated psychometric scales infallible, they carry their own measurement error. But they have two structural advantages a categorical diagnostic interview doesn’t. First, their test-retest reliability is published, scale-specific, and calculated the same way every time, so a clinician knows exactly how much confidence a given score deserves. Second, because most produce a continuous severity score rather than a binary present or absent judgement, a small amount of measurement error moves a result up or down a severity band rather than flipping it across a diagnostic threshold entirely.
Also worth building into a routine assessment battery: AUDIT for alcohol use, MDQ or the General Behaviour Inventory for bipolar screening, and the MSI-BPD or BSL-23 for personality presentations.
The full library is at novopsych.com/assessments, built on the same underlying philosophy as NovoPsych Psychometrics: quantify what can be quantified, and be explicit about the measurement error in what can’t.
The Skill Clinicians Need to Cultivate
None of this is an argument against the clinical interview — quite the opposite. A skilled interviewer elicits information a checklist cannot: they notice the flat affect that contradicts a denied low mood, they follow up on a detail the patient mentioned and then moved past, they adjust pacing and phrasing when someone is guarded or ashamed. That is exactly the “contextual and phenomenological information” this paper’s authors say diagnostic assessment needs more of. It’s a real clinical skill, and it isn’t one a rigid interview schedule protects — if anything, a rigid schedule can crowd it out.
But we should be honest about the limits of human judgement here too. We’ve known for a long time that clinical judgement is vulnerable to bias, and inter-rater reliability on clinician-rated scales has never been particularly strong either. The unreliability this meta-analysis documents in structured interviews shows up, in different forms, wherever a human is doing the rating. That’s the uncomfortable symmetry: neither the checklist nor the clinician’s impression, on its own, is as reliable as either of us would like.
This is where I think artificial intelligence is genuinely useful — and where it isn’t. AI can standardise the parts of the process that don’t need a human: consistent coverage of relevant questions, consistent documentation, consistent scoring against a scale’s published cut-offs. What it can’t do, at least not yet, is replace the moment-to-moment judgement of deciding what to ask next, when to slow down, when to gently push, when to let a silence sit. That remains squarely a human skill — and it arguably becomes more important as everything downstream of the interview gets more standardised.
So here is where I think this is heading: AI, layered with validated psychometric instruments, will increasingly do more of the analysis and the deciding — flagging likely diagnoses, tracking severity over time, checking a clinical impression against a normed cut-off. The clinician’s highest-value skill shifts toward conducting a warm, attentive interview that actually elicits the information all of that downstream analysis depends on. Garbage elicitation still means garbage data, no matter how sophisticated the layer sitting on top of it.
Ambient AI and the Next Step
This is also why I think ambient AI, AI that listens to and helps analyse the clinical interview itself, is one of the more important developments in our field. Done well, it takes the coding and documentation burden off the clinician entirely, freeing them to be fully present in the conversation rather than dividing attention between listening and writing.
It also means the analysis step — matching what was said against diagnostic criteria, flagging inconsistencies, prompting relevant follow-up questions grounded in validated psychometric scales — can happen consistently every time, rather than depending on which interviewer happens to be in the room that day. Ambient AI’s potential here goes beyond transcription: large language models applied to interview audio and transcripts can already extract clinically meaningful signal in real time.
Chen and colleagues, publishing in npj Mental Health Research in 2025, trained large language models on audio recordings from 1,160 outpatients with depression and anxiety, using them to identify clinical symptoms directly from psychiatrist to patient dialogue. The models detected the presence of clinical annotations with 86.9 percent accuracy, and correctly identified specific anxiety and depression symptoms roughly three quarters of the time — even picking up disorder-specific markers, like anhedonia and reduced volition in depression versus tension and an inability to relax in anxiety, from the language used.
Separately, speech itself carries measurable personality signal: Lukac, publishing in Scientific Reports in 2024, used acoustic and linguistic embeddings from free-form speech samples of over 2,000 participants to predict Big Five personality traits, finding correlations with self-report ranging from 0.26 for extraversion to 0.39 for neuroticism — modest on their own, but rising as high as 0.60 once corrected for the unreliability of the self-report measure itself.
The future points to something really useful: an ambient AI scribe that gives a standardised, quantified second read on the same conversation — a consistency check the clinician can weigh against their own judgement, rather than a replacement for it. Given that this meta-analysis found no explanation for why structured-interview reliability swings so much from rater to rater, adding that kind of quantified second signal is one of the more promising ways to shore up the weak link the study exposes.
This is something I’m genuinely proud NovoNote is part of building. We’re not trying to replace the interview or the interviewer — this meta-analysis is a good reminder of why that would be a mistake. We’re trying to standardise the layer underneath it: reliable capture, reliable scoring against validated psychometric instruments, and consistent documentation, so the clinician’s time and skill go into the part of the job a checklist, an algorithm, or an AI model still can’t do as well as a good clinician can — eliciting the truth from another human being.
This is the same principle I wrote about when reviewing the evidence on AI-generated psychological and psychiatric report accuracy: AI is most trustworthy, and most useful, when it’s grounded in objective, quantified data rather than left to freely interpret a transcript. The lesson from this reliability meta-analysis is the mirror image of that argument, applied one step earlier in the pipeline — at the interview itself, rather than the report that follows it.
Practical Takeaways
- Don’t treat a single structured-interview diagnosis, however standardised, as infallible — especially for anxiety, depressive, personality, and psychotic presentations, where reliability was lowest and least explained by study quality.
- Pair the interview with a validated, disorder-matched psychometric measure and re-administer it over time. A continuous severity score degrades more gracefully than a categorical diagnosis when measurement error is present.
- If you’re using an AI scribe or ambient AI to support interviews, use it to standardise capture and scoring against objective psychometric data — not to replace your own probing and follow-up questions. The evidence says that’s still the part of the process that needs a skilled human.
- Re-assess. A reliability level in the 0.55 to 0.70 range means a diagnosis made at one point in time deserves revisiting, not permanent status.
Where I Land
This meta-analysis is a useful corrective to the idea that structure alone equals reliability. Standardized diagnostic interviews earned their “gold standard” reputation because they’re more reliable than an unstructured clinical impression — but “more reliable than nothing standardized” and “reliable enough to hang a permanent diagnosis on” are different claims, and this paper is evidence for the first, not the second.
My reading of where the field is heading: psychometric instruments, and increasingly AI — eventually ambient AI — will carry more of the analytic and scoring load, and do it more consistently than a human rater managing that task alongside forty other things in a session. That leaves the clinician’s core skill as the thing it probably always should have been: conducting an interview that actually gets the patient to tell you what’s going on. Standardising the interview format was always meant to serve that goal, not replace it — and the evidence increasingly suggests we should be standardising the analysis, not the humanity, of the conversation.
References
- Xie, W., Nordgaard, J., Sheldrick, R. C., Ahmad, J. F., Gomes, F. A., & Duncan, L. (2026). Test-retest reliability of standardized diagnostic interviews for common adult psychiatric disorders: A systematic review and meta-analysis. JAMA Network Open, 9(5), e2615039. doi.org/10.1001/jamanetworkopen.2026.15039
- Chen, J., et al. (2025). Identifying psychiatric manifestations in outpatients with depression and anxiety: a large language model-based approach. npj Mental Health Research, 4, 63. doi.org/10.1038/s44184-025-00175-1
- Lukac, M. (2024). Speech-based personality prediction using deep learning with acoustic and linguistic embeddings. Scientific Reports, 14, 30149. doi.org/10.1038/s41598-024-81047-0
Related reading: Are AI Psychological and Psychiatric Reports Accurate? What the Research Actually Says (NovoPsych, 2026)