All writing
17 min readAI in medicine, Psychiatry, Diagnosis

The Wrong Object

Psychiatric evidence isn't found in the patient. It's made in the room, which may explain why AI has reached every diagnostic frontier except this one.

In 1914, Archduke Franz Ferdinand was assassinated. Historians agree this triggered World War I. They disagree on what caused it, because the cause wasn't an event. It was a structure: alliances, industrial capacity, competing empires, a system already primed for violence. The assassination mattered only because it detonated something already there. Treating the trigger as the cause is a category error that collapses mechanism into moment.

Psychiatry makes a version of the same error, with a twist. The diagnostic interview gets treated as a trigger, a probe that sets off evidence already sitting inside the patient, the way the shot in Sarajevo set off a continent. But the interview doesn't reveal a pre-existing psychiatric fact. It generates one. And unlike the alliances of 1914, the structure that gives a patient's answer its meaning does not precede the event. It is assembled in the room, turn by turn. The patient's response to a specific clinician's question, posed in a specific moment, with a specific history between them, constitutes the evidence. Change the interviewer, the phrasing, the timing, the therapeutic stance - the evidence shifts. Not because the diagnosis was vague or the patient inconsistent, but because psychiatric evidence has no existence independent of the relational process that produces it.

This, I have come to think, is why AI diagnostic tools have found no purchase in psychiatry. They treat the interview as data extraction, input variables generating a fixed output. They don't account for the fact that the output is the relationship. The evidence isn't hidden in the interview. It's produced by the interview.

I read a paper by a Lancaster University group studying how mental health professionals actually use technology in practice. The headline finding - uneven adoption of digital tools across the stages of clinical practice - was not surprising. Clinicians' reluctance to adopt new tools is well documented. What caught me was the shape of the unevenness. The specialized AI tools clinicians had adopted clustered almost entirely around treatment: symptom trackers, CBT modules, relapse-prevention tools, meditation apps. For diagnosis, the picture was thin. Practitioners reached for general-purpose software to document what they had found, and reached for almost nothing specialized to help them find it in the first place.

This seemed backwards. If you had to bet, sight unseen, on which half of psychiatric practice would fall to AI first, diagnosis would look like the safer wager. Diagnosis is pattern recognition over data, and pattern recognition over data is the one thing AI does uncontroversially well. It is the basis of the technology's most legible medical wins: systems that read mammograms, retinal scans, and ECGs at or above specialist level. Treatment, by contrast, is supposed to be the stubbornly human part of medicine: judgment, rapport, the slow accumulation of trust with someone in crisis. So why would the easier problem, on paper, be the one nobody had touched?

The obvious answers don't hold up well. One is that the tools simply aren't good enough yet. The paper - by Abeer Alotaibi and Corina Sas at Lancaster - gestures at something like this: limited evidence of clinical effectiveness, a lack of training. But it doesn't say the tools fail to perform. It says clinicians don't trust them enough to use them, and use them too rarely for the evidence base to grow, which keeps the distrust intact. That's a real dynamic, but it isn't a psychiatric peculiarity. Radiologists spent years worried about deskilling and liability before AI-assisted reads became routine. Cardiologists ask identical questions about black-box risk scores. Distrust shows up everywhere a model meets a license to practice. If generic distrust were doing all the explanatory work here, you'd expect roughly even resistance to diagnostic AI across specialties - not a sharp split where psychiatry holds the door shut on diagnosis while waving treatment tools straight through.

One detail in the practitioner interviews pointed somewhere more interesting. A recurring concern was that diagnostic tools risked interrupting the emotional contact between clinician and patient. It is an obvious point - so obvious that one might mentally shrug and move on. It gave me pause instead, and I chose to take it seriously as a claim about evidence. What if breaking the texture of the conversation is not merely less than ideal? What if it removes the mechanism that produces the diagnostic evidence in the first place? (I don't mean face-to-face contact, or anything uniquely human. I mean the shape of the interaction: the psychiatric interview, a form cultivated over millennia of inquiry, centuries of medicine, and decades of clinical craft.)

Put a CT scan in front of a radiologist and the lesion, if it's there, is already there - present in the tissue before anyone looks, indifferent to who's looking or how carefully. That independence is exactly what makes the task tractable for a pattern-matching system: train on enough labeled images and a model can learn to find the same fact, because the fact doesn't move. Cardiology, dermatology, most of oncology - these specialties run on evidence that precedes the clinician. It sits in the tissue, the rhythm strip, the blood panel, waiting.

Psychiatry doesn't always have that. Rose McCabe, who has spent two decades filming and transcribing psychiatric consultations to study them with the tools of conversation analysis, has made the point bluntly: unlike almost every other field of medicine, psychiatry rarely has biomarkers or lab values to check a diagnostic judgment against. The evidence is the conversation. And a conversation, unlike a lesion, doesn't sit still waiting to be observed - it gets built, turn by turn, out of what the clinician asks and how the patient answers.

This isn't unique to psychiatry in some absolute sense. Every clinical history shapes what a doctor finds, in cardiology as much as anywhere else. The difference is the ability to course correct. In cardiology, an off-target interview is a recoverable error - order the echo, run the troponin, let the test correct for what the conversation missed. In psychiatry there is, in the overwhelming majority of cases, no independent test to correct against. The interview isn't one input among several. Most of the time, it's nearly all there is.

That claim sounds like philosophy until you look at what conversation analysts have actually found by studying real psychiatric interviews on tape. McCabe's research on "repair" - the back-and-forth by which doctor and patient catch and fix misunderstandings mid-conversation - found that patients who clarified or corrected a psychiatrist's talk during a consultation went on to show better treatment adherence, and that training psychiatrists to build a shared understanding of a patient's experience of psychosis measurably increased how often the psychiatrists repaired their own talk in response. None of this is speculative. It's the kind of finding you only get by treating the interaction itself as the object of study, rather than as a wrapper around some more real signal sitting underneath it.

There's a sharper version of the same point in how question wording moves disclosure. Ask someone whether they've ever tried to kill themselves and you may get a flat no. Ask the same person, a few minutes later, whether they've ever hurt themselves on purpose, and you may get a different answer. This isn't a trick of phrasing; it's a known hazard of clinical interviewing generally, and researchers using conversation analysis have specifically studied how clinicians and patients negotiate exactly this kind of disclosure around suicide risk. In most of medicine, a wording problem like that costs you a slightly noisier estimate. In psychiatric assessment, where the disclosure often is the evidence, a wording problem can be the difference between a diagnosis and its absence.

None of this means psychiatric evidence is entirely manufactured by the interview. Sleep loss, psychomotor slowing, a frankly disorganized speech pattern, the visible signs of self-neglect - these can be observed rather than elicited, and a sufficiently attentive system, human or otherwise, could in principle register them without asking anything clever. What seems to depend on the interaction specifically is the more decisive evidence: the texture of guardedness, the inconsistency that reveals more than a consistent answer would, the difference between a patient minimizing out of shame and one who genuinely doesn't experience the symptom you're asking about. The evidence that shows up only when somebody asks in a way that lets it.

This is also an old argument inside psychiatry itself, and one with real data behind it. In 2012, the Danish phenomenological psychiatrist Julie Nordgaard and colleagues at the University of Copenhagen tested how well the Structured Clinical Interview for DSM-IV - the most widely used standardized diagnostic instrument in the field - agreed with diagnosis reached the older way: a consensus judgment by two experienced clinicians, drawing on every source available to them, including a long, videotaped, semi-structured conversation. On a sample of a hundred first-admitted patients, the two methods agreed at a kappa of 0.18 - a figure the statistical literature files under "questionable agreement," barely distinguishable from chance. A 2023 replication on a fresh sample of first-admission psychosis patients landed at 0.21. Nordgaard traced the mismatch to the structured instrument's dependence on what patients said when asked a fixed question in a fixed way: over-reliance on self-report, vulnerability to patients who minimized or dissembled, a tendency to chase comorbidity checklists rather than the shape of the underlying disorder. Her broader argument, developed with Josef Parnas and the philosopher Louis Sass, is that a fully structured interview narrows a patient's possible answers into categories - yes, no, sometimes - that don't fit how a lot of psychopathology actually presents, especially the subtler disturbances of self-experience that resist reduction to a checklist item. Their conclusion was unambiguous: full structure isn't a more rigorous version of the psychiatric interview. It's an instrument poorly matched to what it's trying to measure.

"The nature of this information is conceived on analogy of a substantial, temporally enduring thing, almost like a table or a chair." — Nordgaard, Sass & Parnas, 2013

A skeptic should refuse to let that finding pass unexamined, and the strongest version of the objection goes like this. Low agreement between two methods tells you that they disagree; it does not tell you which one is wrong. Nordgaard's gold standard - senior clinicians conferring until consensus - is exactly the kind of expert judgment that a half-century of research, from Paul Meehl onward, has shown can be beaten by dull actuarial rules. And in the 2012 study the SCID was administered by a trained non-clinician, which means the trial arguably tested who was asking as much as it tested structure itself. All of this is fair. But follow the objection to its end and notice where it lands. If there is no external criterion - no scan, no assay - that could settle which method got those patients right, then the disagreement itself becomes the finding. Two procedures, applied to the same hundred people, produced two different sets of psychiatric facts, and nothing outside the procedures can adjudicate between them. That is not a rebuttal of the interaction-dependence argument. It is the argument, stated in kappa.

You can see the same problem from the engineering side, in how the field has actually built systems to detect mental illness from interview data. The DAIC-WOZ depression-detection corpus, built around interviews conducted by Ellie, a virtual interviewer controlled by a hidden human (à la Wizard of Oz), became a canonical benchmark for automatic mental-health inference. The standard move was to treat the interviewer's turns as context, nuisance, scaffolding. The patient's words, face, and voice were presumed to contain the true signal.

When framed like this, depression behaves like a hidden internal state leaking out through speech, and Ellie's questions are just the funnel that gets it talking.

A 2024 paper out of the Idiap Research Institute complicated that picture in a useful way. The researchers found that models incorporating Ellie's prompts performed better, then showed why: those models weren't learning anything about how the clinician's questions shaped what the patient disclosed. They were learning to spot the specific moments in the transcript where Ellie asked about a history of mental illness, and using the mere presence of that question as a shortcut to the answer. Exploiting nothing but that one bias on purpose, the researchers got an F1 score of 0.90 - the best published result on the dataset using text alone.

This is an elegant failure. It proves the interviewer matters while missing the reason the interviewer matters. The model found that some questions predict labels. It did not learn that the meaning of an answer may depend on the clinical move that made the answer possible.

A three-second silence after "Have you ever wanted to die?" is not the same clinical object as a three-second silence after "When you say you are tired, do you mean sleepy, empty, or done?" The duration matches. The evidence does not. One is a warning; the other is a person deciding between three words. Extraction converts both into the same timestamped gap.

This is the deeper version of the Nordgaard problem, ported into machine learning. It isn't that researchers in this field have failed to notice the interviewer matters - the Idiap paper is proof they have. It's that the underlying architecture, built to find a fixed feature somewhere in a span of input the same way a model finds a tumor in a scan, has no native way to represent the differences between evidential responses that emerge from different prompts: that a three-second pause following a gentle, well-timed question means something different from a three-second pause following an abrupt one. It can find the pause. It can't yet find the relationship that produced it. That's a fair description of this specific kind of system - a classifier. It is not a description of what a modern, agentic dialogue system does, and it would be a mistake to treat it as one.

The more interesting question is what happens once you swap that classifier for such an agent - a system that asks its own questions and adapts them turn by turn.

Google's AMIE is the fairest test of the optimistic case. It is not a crude classifier searching a fixed transcript. It is an adaptive conversational diagnostic system trained through self-play, designed to ask questions, update uncertainty, and move through clinical dialogue. In a Nature evaluation, AMIE was compared with primary-care physicians in simulated text consultations and performed strongly across diagnostic and communication axes. The architecture matters because it shows that the old objection - AI cannot conduct a conversation - is no longer sufficient.

And yet the gap remains. AMIE's evaluation excluded psychiatry. Later clinical feasibility work flagged unreliable handling of mental-health concerns as a risk. That exclusion is not timidity. It is a boundary marker. The system can converge on findings when an external clinical record can later discipline the conversation. Psychiatry offers fewer external anchors. The chart may not verify the interview; the chart may merely preserve what a previous interview managed to elicit.

Here's why the failure mode changes shape rather than disappearing. A model trained on human feedback learns that a good conversation reaches a conclusion - clear, efficient, resolved. That instinct serves a system triaging a sore throat well. It is close to the opposite of what Nordgaard's semi-structured interview requires: a tolerance for staying in ambiguity, treating a patient's hedge or inconsistency as data rather than as noise to clear out of the way before arriving at an answer. The risk with a well-trained conversational AI conducting a psychiatric interview isn't that it would sound robotic. It's that, trained to be helpful and decisive, it would behave like the SCID's most articulate descendant - reaching a clean answer efficiently, at exactly the cost Nordgaard's data says that efficiency extracts.

There's an old philosophical argument that names the deeper version of this gap, and it predates anything resembling machine learning by decades. Maurice Merleau-Ponty, writing about how people perceive each other's emotional states, rejected the idea that anger is a hidden interior fact inferred secondhand from a set of external clues - a raised voice here, a clenched jaw there, evidence assembled after the fact. He argued that anger is given directly, as a whole, in the bodily comportment of the angry person: it inhabits the gesture and the reddened face rather than hiding behind them.

"I could not imagine the malice and cruelty which I discern in my opponent's looks separated from his gestures, speech and body, none of this takes place in some otherworldly realm, in some shrine located beyond the body of the angry man [...] anger inhabits him and blossoms on the surface of his pale or purple cheeks, his blood-shot eyes..."

Phenomenological psychiatry built a technical vocabulary around the same intuition and called it intersubjectivity - the capacity to grasp another person's mental state directly, through pre-thematic contact with how they express themselves, rather than by inference from a checklist of external signs. Josef Parnas, one of the field's central figures and Nordgaard's own collaborator, pushed the idea into diagnosis itself, describing the clinical picture of schizophrenia as a "core Gestalt": an irreducible, prototypical whole that an experienced clinician perceives directly in a patient's expressive and existential pattern, not a sum arrived at by adding up separable symptoms. That's the same claim Nordgaard's data makes empirically, said from the opposite side of the room. A structured interview doesn't fail because it asks the wrong questions. It fails because it can only ever hand you fragments, and the thing it's trying to diagnose was never available in fragment form to begin with.

If a patient's guardedness, or the particular texture of a hesitation, only becomes what it is in the intersubjective back-and-forth of being asked about it, then stripping the interviewer's turns out of the transcript isn't cleaning the signal. It's discarding half of the encounter in which the signal was ever going to become legible, on the theory that it was never really there.

It would be too convenient to let the explanation stop here, though, because there's a less romantic story that covers a good deal of the same ground: regulation. The FDA has never authorized a generative-AI-based device for any clinical purpose. Its November 2025 committee on AI in mental health signaled that diagnostic claims would face the slow De Novo pathway while treatment-support tools can ride the faster 510(k) route. Companies optimize for what sells. That alone could account for a good portion of the lopsidedness Alotaibi and Sas documented, without needing any theory about interviews and interaction at all.

But the regulatory story doesn't compete with the argument so much as sit underneath it. Regulators aren't arbitrarily harder on psychiatric diagnostic claims; they're harder on claims whose evidentiary basis is hardest to pin down, and psychiatric diagnosis is hard to pin down for exactly the reason the interaction-dependence argument gives: there's no independent test to check the claim against. The regulatory caution and the evidentiary problem look, on closer inspection, like the same fact viewed from two different desks.

It's tempting to land this somewhere comforting - to say that diagnosis, in psychiatry, is simply the kind of thing only a human can do, the way people sometimes say this about art after watching what a generative model can paint. I don't think that frame earns its keep here. Nothing in the argument above requires interaction to be a uniquely human capacity, forever closed to machines. The systems already exist that can hold a multi-turn conversation, track their own uncertainty, and adapt their questions on the fly. What's missing is data built and labeled around the idea that a clinician's move is part of the evidence rather than noise around it, and a way to check the resulting diagnosis against something other than an instrument designed to avoid ambiguity in the first place. That's an unsolved data and evaluation problem, not a metaphysical one. "No one has built it" does not mean "no one can," and it would be premature at best and presumptuous at worst to rule the technology out this early. Building it would require data nobody has collected - the kind that's hard to gather at scale and harder still to clear for use - scored by a framework for grading questions over responses that no one has created yet.

Until something like that exists, the asymmetry Alotaibi and Sas found will keep looking like a verdict on psychiatry's resistance to technology, when it's closer to a verdict on what current technology was built to look for. The evidence shows up, when it shows up at all, only after someone has asked exactly the right question, in exactly the right way, at exactly the right moment - and so far, nobody has figured out how to put that into a training set. Diagnostic AI is strongest where evidence has already become an object. Perhaps in psychiatry, all that remains is to hand it the right one.


Further reading

  • Abeer Alotaibi & Corina Sas, Practitioners' Use of Mental Health Technologies across Clinical Practice Stages, Lancaster University (2025)
  • Julie Nordgaard, Rasmus Revsbech, Ditte Søbye & Josef Parnas, "Assessing the diagnostic validity of a structured psychiatric interview in a first-admission hospital sample," World Psychiatry (2012)
  • "Does method matter? Assessing the validity and clinical utility of structured diagnostic interviews among a clinical sample of first-admitted patients with psychosis: A replication study," Frontiers in Psychiatry (2023)
  • Julie Nordgaard, Louis A. Sass & Josef Parnas, "The psychiatric interview: validity, structure, and subjectivity," European Archives of Psychiatry and Clinical Neuroscience (2013)
  • Josef Parnas, "The Core Gestalt of Schizophrenia," World Psychiatry (2012)
  • Paul E. Meehl, Clinical versus Statistical Prediction: A Theoretical Analysis and a Review of the Evidence (1954)
  • Guilherme Messas, Melissa Tamelini, Milena Mancini & Giovanni Stanghellini, "New Perspectives in Phenomenological Psychopathology: Its Use in Psychiatric Treatment," Frontiers in Psychiatry (2018)
  • Rose McCabe & Patrick G. T. Healey, "Miscommunication in Doctor-Patient Communication," Topics in Cognitive Science (2018)
  • Sergio Burdisso et al., "DAIC-WOZ: On the Validity of Using the Therapist's Prompts in Automatic Depression Detection from Clinical Interviews," Idiap Research Institute (2024)
  • "Towards conversational diagnostic artificial intelligence," Nature (2025), Google DeepMind's AMIE
  • "A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic" (2026)
  • Maurice Merleau-Ponty, The World of Perception (Routledge, 2008)