CLIMB

A Clinical Multimorbidity Benchmark for Diagnosing Cooccurring Conditions through Multiturn Conversations

  • Yusuf Kesmen1*
  • Aniruddha Mukherjee1†
  • Yena Chang1†
  • David Sasu1
  • Trevor Brokowski1
  • Alexandra V. Kulinkina1
  • Kristina Keitel1
  • Akhil Arora2‡
  • Lars Henning Klein1‡
  • Mary-Anne Hartley1‡

1EPFL  ·  2Aarhus University  ·  †equal contribution  ·  ‡equal supervision  ·  *corresponding author

The model follows one diagnosis and misses the others.

Models interview simulated patients and must identify every condition, with no extra diagnoses.

10%

Best exact score after an interview · ePOCT+

26%

Best exact score with the full record · DDXPlus

38 points

Mean recall drop from one to four conditions · DDXPlus interviews

Read the abstract

Patients often have several cooccurring clinical conditions, and the findings needed to identify and disambiguate them emerge over the course of a consultation. Evaluating clinical reasoning in this setting requires both multiturn interaction and multilabel diagnosis. We introduce CLIMB, a benchmark in which a doctor model interviews a simulated patient to recover a ground truth set of cooccurring clinical conditions. Cases are synthesized from clinical decision algorithms and diagnostic datasets, grounding multimorbid presentations in structured clinical knowledge. Across six frontier and open models, none recovers the exact set of conditions in more than 10% of interactive cases. Diagnostic performance declines when conditions cooccur, even when models receive the full clinical record and the true number of conditions. Interaction reduces performance further. In controlled experiments, models behave like single hypothesis trackers: they anchor on the diagnosis suggested by the opening findings, keep questioning around it, and recover a second condition mainly when a finding in view points to it. Questioning them further does not complete the set but adds mostly wrong diagnoses. We formalise this pattern with a theoretical reference model of single hypothesis tracking.

Example consultation

Dr. Agent Common cold? Common cold + Scabies My son has a cough. 3 days ago. Oh yes, he has a rash on hisfingers. Seems itchy. When did it start? Anything else? Diagnosis Ready:Common cold; Scabies Pneumonia · missed Scabies · found ✓ Anemia · missed Common cold · not his ✗
1/3 Conditions found. Two missed; one incorrect diagnosis added.

Scene 1

A child aged two has a cough and three hidden conditions. The model must ask questions to identify them.

Scene 2

The model suspects a common cold. The child does not have one.

Scene 3

An itchy rash leads it to add scabies, while keeping the cold diagnosis.

Scene 4

It stops without identifying pneumonia or anemia.

Illustration from Figure 1, not a recorded consultation.

How patients are generated

CLIMB What brings you in today? Patient record HIDDEN SET size unknown
  1. 1

    Draw conditions

    1 to 4 compatible diagnoses.

  2. 2

    Merge findings

    Combine their clinical findings.

  3. 3

    Add noise

    Add unrelated complaints.

  4. 4

    Hide the diagnoses

    Show only the opening complaint.

See how cases are built
Schematic decision paths Condition A + Condition B
Finding 1 Present Absent Finding 2 Finding 3 Finding 3 PresentPresentPresent Condition A Condition B Condition B
Path A

Finding 1 Present

Finding 2 Present

Path B

Finding 1 AbsentPresent

Finding 3 Present

Compare shared answers× Conflict. Reject this draw.✓ Compatible paths

Combined patient✓

Condition A + Condition B

Finding 1
Present
Finding 2
Present
Finding 3
Present

Leaderboard

Climbing score (%)

Scale: 0 to 10%. Perfect score: 100%.

Results

5.1 · Multiplicity

Do models ask more questions when patients have more conditions?

The question count stays near 15 while exact set accuracy falls.

Models return more diagnoses as conditions per patient (k) increase, but the number of correct diagnoses grows more slowly.

Exact set accuracy falls

DDXPlus interviews · mean across six models · Table 2

Number of questions asked

Same cases and models · maximum 20 questions · Figure 3c

5.2 · Single answer gap

Why is one correct diagnosis not enough?

With four conditions, about 80% of answers include a true diagnosis. Exact set accuracy is 0%.

Most medical benchmarks ask for a single diagnosis. Here we compare finding any true diagnosis with naming every condition and no extras.

At least one correct diagnosis versus the exact set

DDXPlus · six models · Figure 4a and Table 2. The upper curve counts any true diagnosis in the answer, not top 1 accuracy; its values are approximate.

6.1 · Acquiring and using evidence

Why does the model miss a second condition?

Its questions collect less useful evidence, and it makes errors even with the full record.

The model and a trained decision tree choose from the same 49 questions. Answers come directly from records of patients with two respiratory conditions.

Which questions provide better evidence?

The same trained classifier reads both sets of answers. It is more accurate with the answers collected by the tree, which asks fewer questions.

Same classifier, different questions

GPT 5.6 versus decision tree · exact set accuracy · Table 12

How informative is each question?

The tree gains about twice as much information per question. Information gain measures how much an answer reduces uncertainty about the diagnosis set.

Information gained per question

Both systems evaluated with the same statistical model over the tree's question budget · Figure 5c and Appendix D.7

What happens when the model receives all the findings?

With the full record, exact set accuracy is 26% for the model and 79% for the trained classifier.

Both receive the same findings, so missing information explains only part of the model's errors.

Exact set accuracy with different evidence

GPT 5.6 versus classifier · same cases · findings supplied as a list · classifier trained on generated cases · Table 12

6.2.1 · Opening findings

Does the opening affect which condition is diagnosed?

The model is more likely to name the condition whose symptoms appear in the opening.

Each patient has two equally severe conditions. We test each case twice, changing only which condition's findings appear in the opening.

65%

condition with symptoms in the opening

32%

other condition

Each icon ≈ 1% · GPT 5.6 · 195 pairs, two openings each · Table 20. Answers may include extra diagnoses.

6.2.2 · Shared symptoms

Does a clue to the second condition help?

The model finds the second condition more often when the opening includes symptoms shared by both.

Each patient has conditions A and B. We replace two opening symptoms of A with symptoms shared by A and B. The increase in naming both conditions is not statistically significant.

Change only the opening symptoms

GPT 5.6 · 97 patients, equally severe conditions · Table 21. Naming both may include extra diagnoses.

6.3 · Forced continuation

Do more questions help?

Forcing the model to keep asking questions adds mostly incorrect diagnoses.

After 20 questions, the model reports an average of 1.3 correct diagnoses out of two.

Correct and incorrect diagnoses as questions continue

GPT 5.6 · 50 DDXPlus cases · Figure 7 · approximate curves. Shading starts at the median stop. Diagnoses are checked after each turn without feeding them back into the interview.

Theoretical model

Set tracker

Tracks support for each condition separately, allowing several diagnoses to remain supported at the same time.

Single hypothesis tracker

Diagnoses compete for support, and questions focus on the current leading hypothesis. A second condition can remain unexplored.

BibTeX

@article{kesmen2026climb,
  title   = {{CLIMB}: A Clinical Multimorbidity Benchmark for Diagnosing
             Co-occurring Conditions through Multiturn Conversations},
  author  = {Kesmen, Yusuf and Mukherjee, Aniruddha and Chang, Yena and
             Sasu, David and Brokowski, Trevor and Kulinkina, Alexandra V.
             and Keitel, Kristina and Arora, Akhil and Klein, Lars Henning
             and Hartley, Mary-Anne},
  journal = {arXiv preprint arXiv:2609.35462},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.35462}
}