Condition A + Condition B
- Finding 1
- Present
- Finding 2
- Present
- Finding 3
- Present
A Clinical Multimorbidity Benchmark for Diagnosing Cooccurring Conditions through Multiturn Conversations
1EPFL · 2Aarhus University · †equal contribution · ‡equal supervision · *corresponding author
Models interview simulated patients and must identify every condition, with no extra diagnoses.
10%
Best exact score after an interview · ePOCT+
26%
Best exact score with the full record · DDXPlus
38 points
Mean recall drop from one to four conditions · DDXPlus interviews
Patients often have several cooccurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multiturn interaction and multilabel diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of cooccurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions cooccur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single hypothesis tracking.
Scene 1
A child aged two has a cough and three hidden conditions. The model must ask questions to identify them.
Scene 2
The model suspects a common cold. The child does not have one.
Scene 3
An itchy rash leads it to add scabies, while keeping the cold diagnosis.
Scene 4
It stops without identifying pneumonia or anemia.
Illustration from Figure 1, not a recorded consultation.
1 to 4 compatible diagnoses.
Combine their clinical findings.
Add unrelated complaints.
Show only the opening complaint.
Finding 1 Present
Finding 2 Present
Finding 1 AbsentPresent
Finding 3 Present
Compare shared answers× Conflict. Reject this draw.✓ Compatible paths
Condition A + Condition B
Example records One record per condition
Condition A
Female · 36 years
Condition B
Female · 42 years
Condition A + Condition B
Female · 40 years
Climbing score (%)
Climbing score = (DDXPlus exact + ePOCT+ exact) / 2.
Exact means all diagnoses correct, with no extras.
Interactive interviews, up to 20 questions. Each dataset has 800 cases, balanced across 1 to 4 conditions per patient.
| Rank | Model | DDXPlus | ePOCT+ | Climbing score |
|---|
Source: Table 1. Published scores are rounded; small differences may not be significant (Appendix E.1).
5.1 · Multiplicity
The question count stays near 15 while exact set accuracy falls.
Models return more diagnoses as conditions per patient (k) increase, but the number of correct diagnoses grows more slowly.
5.2 · Single answer gap
With four conditions, about 80% of answers include a true diagnosis. Exact set accuracy is 0%.
Most medical benchmarks ask for a single diagnosis. Here we compare finding any true diagnosis with naming every condition and no extras.
6.1 · Acquiring and using evidence
Its questions collect less useful evidence, and it makes errors even with the full record.
The model and a trained decision tree choose from the same 49 questions. Answers come directly from records of patients with two respiratory conditions.
The same trained classifier reads both sets of answers. It is more accurate with the answers collected by the tree, which asks fewer questions.
The tree gains about twice as much information per question. Information gain measures how much an answer reduces uncertainty about the diagnosis set.
With the full record, exact set accuracy is 26% for the model and 79% for the trained classifier.
Both receive the same findings, so missing information explains only part of the model's errors.
6.2.1 · Opening findings
The model is more likely to name the condition whose symptoms appear in the opening.
Each patient has two equally severe conditions. We test each case twice, changing only which condition's findings appear in the opening.
65%
condition with symptoms in the opening
32%
other condition
6.2.2 · Shared symptoms
The model finds the second condition more often when the opening includes symptoms shared by both.
Each patient has conditions A and B. We replace two opening symptoms of A with symptoms shared by A and B. The increase in naming both conditions is not statistically significant.
6.3 · Forced continuation
Forcing the model to keep asking questions adds mostly incorrect diagnoses.
After 20 questions, the model reports an average of 1.3 correct diagnoses out of two.
Tracks support for each condition separately, allowing several diagnoses to remain supported at the same time.
Diagnoses compete for support, and questions focus on the current leading hypothesis. A second condition can remain unexplored.
@article{kesmen2026climb,
title = {{CLIMB}: A Clinical Multimorbidity Benchmark for Diagnosing
Co-occurring Conditions through Multiturn Conversations},
author = {Kesmen, Yusuf and Mukherjee, Aniruddha and Chang, Yena and
Sasu, David and Brokowski, Trevor and Kulinkina, Alexandra V.
and Keitel, Kristina and Arora, Akhil and Klein, Lars Henning
and Hartley, Mary-Anne},
journal = {arXiv preprint arXiv:2609.35462},
year = {2026},
url = {https://arxiv.org/abs/2609.35462}
}