In today’s post, Mateo Boberg, Julie Nordgaard and Mads Gram Henriksen (University of Copenhagen) explore how reliable psychiatric diagnoses are in the 21st century.
What happens if two psychiatrists assess the same patient? Most people, including many clinicians, would probably expect them to reach the same diagnosis. After all, modern psychiatry has spent decades developing operational diagnostic criteria intended to improve diagnostic reliability. But how reliable are psychiatric diagnoses today?
![]() |
| Mads Gram Henriksen |
![]() |
| Mateo Boberg |
![]() |
| Julie Nordgaard |
What happens if two psychiatrists assess the same patient? Most people, including many clinicians, would probably expect them to reach the same diagnosis. After all, modern psychiatry has spent decades developing operational diagnostic criteria intended to improve diagnostic reliability. But how reliable are psychiatric diagnoses today?
This question has troubled psychiatry for more than fifty years. The famous US–UK Diagnostic Project in the 1970s revealed striking differences between American and British psychiatrists, helping motivate the development of operational diagnostic systems such as DSM-III and later ICD-10. The hope was that explicit diagnostic criteria and guidelines would make psychiatric diagnosis much more reliable.
To examine whether this hope has been fulfilled, we asked more than 1,000 medical doctors working in psychiatry across 19 countries on two continents to diagnose standardised clinical case vignettes. Because every participant received exactly the same clinical information, differences in diagnosis could not be attributed to differences in interviewing style or patient presentation, only to how doctors interpreted the same material.
The overall picture was striking. Two doctors agreed on the diagnosis only 55% of the time, corresponding to a Krippendorff's alpha of 0.48, a value that must be called ‘modest’ at best. For context, agreeing by chance alone would happen only about 13% of the time, so real clinical judgment is clearly doing work here.
Agreement varied a great deal from case to case. Some cases produced strong consensus, with most doctors landing on the same diagnosis. Other cases split doctors almost evenly between two or more very different diagnoses — one case, for instance, was diagnosed as OCD by roughly two in five doctors and as schizophrenia by a similar proportion. The cases that produced the most disagreement all involved schizophrenia spectrum presentations, and this pattern held regardless of doctors' experience level.
Importantly, the disagreements were not random. Doctors disagreed in systematic ways. The same handful of cases repeatedly split between the same small set of alternative diagnoses: most often OCD, personality disorder, or bipolar disorder, but never toward, say, eating disorder or substance use disorder. One likely mechanism is that clinicians anchored on salient but diagnostically secondary symptoms, such as obsessive features, while underweighting features that should have taken precedence in the ICD-10 hierarchy, such as psychotic symptoms. In other words, clinicians appeared to be applying different interpretive frameworks to the same clinical material.
Our findings do not imply that psychiatric diagnosis is arbitrary. Rather, they suggest that diagnosis depends on more than checking whether a list of criteria is fulfilled. Psychiatric assessment requires clinicians to interpret complex patterns of experience and behaviour, and different clinicians may weigh the same phenomena differently.
Nevertheless, the modest diagnostic reliability may have consequences both for the individual patient and for psychiatric research, e.g., if research samples are less homogeneous than we assume, it could help explain why symptoms often cross diagnostic boundaries, and why trans-diagnostic approaches have gained ground in recent years. A deeper understanding of psychopathology and clinical reasoning is needed to address these issues.


