Research10 March 2026/EvidenceMD Research Team/14 min read

EvidenceMD achieves state of the art on HealthBench Hard at 54.6%

A domain specialized clinical reasoning model that outperforms general purpose LLMs on the hardest medical benchmarks, without sacrificing safety or compassion.

54.6%
HealthBench Hard
+8.4 vs next
66.6%
HealthBench Overall
+4.5 vs next
68.0%
Clinical Reasoning
+12.6 vs next

Motivation

General-purpose large language models have made remarkable strides in natural language understanding, but medicine presents unique challenges that generic training cannot adequately address. Clinical reasoning demands epistemic humility, the ability to recognize the limits of one's knowledge, alongside the capacity to synthesize differential diagnoses, weigh competing evidence, and communicate with appropriate caution.

Models optimized for helpfulness and instruction following often develop sycophantic tendencies: affirming user assumptions even when clinically incorrect, providing overly confident diagnoses without appropriate caveats, or failing to recommend urgent escalation when warranted. In healthcare, this behavior isn't just unhelpful. It's dangerous.

EvidenceMD was built from the ground up to address these shortcomings. The model prioritizes clinical accuracy grounded in peer reviewed literature, reduced sycophancy through training that rewards respectful disagreement, triage and deferral when cases exceed its confidence threshold, and compassionate communication that maintains an empathetic tone while delivering accurate information.

How We Evaluate

We evaluate EvidenceMD on two primary benchmarks, each designed to test fundamentally different aspects of clinical reasoning.

1

OpenAI HealthBench

Open-source benchmark, published May 2025

HealthBench is an open-source evaluation framework developed by OpenAI in collaboration with 262 physicians across 26 specialties and 60 countries. It contains 5,000 realistic multi turn clinical conversations graded against 48,562 clinician developed rubric criteria.

Unlike prior medical benchmarks that rely on multiple choice questions (like USMLE style MedQA), HealthBench evaluates models on open ended, realistic clinical dialogues. Each response is scored across five axes: accuracy, completeness, communication quality, context awareness, and instruction following. Each criterion carries a physician assigned point value weighted by clinical importance.

HealthBench Hard is a specialized subset designed to stress test model failures on the most ambiguous, high stakes clinical scenarios, specifically those where physician agreement is highest but model performance tends to collapse. It serves as the most rigorous public benchmark for clinical AI.

2

Evidence Based Clinical Reasoning

Proprietary benchmark, developed internally

Our proprietary benchmark evaluates multi step diagnostic reasoning that requires models to synthesize patient history, generate ranked differentials, cite relevant evidence from the medical literature, and produce guideline concordant management plans. Unlike HealthBench, which evaluates communication alongside accuracy, this benchmark focuses exclusively on the depth and correctness of clinical reasoning chains.

Scoring criteria penalize models for unsupported assertions, failure to consider red flag diagnoses, and incorrect evidence attribution. This makes it particularly effective at distinguishing genuine clinical reasoning from superficially plausible outputs.

Benchmark Results

HealthBench Hard

SOTA

Consensus required subset focusing on the most challenging clinical scenarios that require nuanced reasoning.

EvidenceMD
54.6%
GPT-5.4 High
46.2%
Gemini 3.1 Pro
45.8%
Claude Opus 4.6
44.4%
Qwen 3.5 Max
43.1%
Kimi 2.5 Thinking
39.6%
DeepSeek R1
34.2%

HealthBench Overall

SOTA

Full evaluation across all clinical question categories.

EvidenceMD
66.6%
GPT-5.4 High
62.1%
Gemini 3.1 Pro
61.4%
Claude Opus 4.6
60.8%
Qwen 3.5 Max
59.3%
Kimi 2.5 Thinking
54.7%
DeepSeek R1
49.1%

Evidence Based Clinical Reasoning

SOTA

Multi step diagnostic reasoning with evidence synthesis, differential ranking, and guideline concordant management.

EvidenceMD
68.0%
GPT-5.4 High
55.4%
Claude Opus 4.6
55.1%
Gemini 3.1 Pro
54.6%
Qwen 3.5 Max
53.8%
Kimi 2.5 Thinking
48.2%
DeepSeek R1
43.5%

Beyond Benchmarks

Benchmarks like HealthBench represent a significant step forward in evaluating clinical AI, but they remain static evaluations against fixed datasets. Real clinical medicine is not a closed book exam. Physicians encounter novel drug interactions, emerging pathogens, evolving practice guidelines, and patients whose presentations defy textbook descriptions daily. A model that scores well on a fixed benchmark may still fail catastrophically in the unpredictable environment of actual clinical practice.

We believe the next frontier of clinical AI evaluation must move beyond static benchmarks toward real time, dynamic validation that continuously tests models against live clinical scenarios, evolving medical literature, and real world patient outcomes. This means evaluating not just whether a model can answer a question correctly today, but whether it can adapt its reasoning as new evidence emerges, flag when its training data may be outdated, and maintain well calibrated confidence across the full spectrum of clinical uncertainty.

EvidenceMD is designed from the ground up to support this paradigm. Our retrieval augmented architecture allows the model to ground its responses in the latest peer reviewed literature rather than relying solely on static training data. We are actively developing a continuous clinical validation framework that evaluates model performance against real time case feeds, updated guideline changes, and emerging pharmacovigilance signals, moving clinical AI evaluation from a single benchmark score to an ongoing, living assessment.

EvidenceMD aims to be the first clinical AI system to publish ongoing, transparent real time validation metrics alongside traditional benchmark results, setting a new standard for how clinical AI models should be evaluated and trusted.

Clinical Examples

Side by side comparisons on real clinical queries, demonstrating depth of reasoning and evidence grounding.

Q

A 64 year old man with cirrhosis presents with worsening confusion, ascites, creatinine rising from 1.0 to 2.6 mg/dL, urine sodium <10 mmol/L, and no improvement after IV fluids; how can you differentiate hepatorenal syndrome from acute tubular necrosis and what management strategy should be prioritized?

Quick Actionable Summary Diagnosis: The patient's presentation of cirrhosis, new-onset confusion, rising creatinine (1.0 → 2.6 mg/dL), low urine sodium (<10 mmol/L), and failure to improve with IV fluids strongly indicates Hepatorenal Syndrome-Acute Kidney Injury (HRS-AKI). Differentiation: Traditional urinary indices (urine sodium, FENa) are notoriously unreliable in cirrhosis. Differentiate HRS from Acute Tubular Necrosis (ATN) using urine sediment (absence of granular casts favors HRS) and novel biomarkers like urine neutrophil gelatinase-associated lipocalin (uNGAL). Management Priority: Immediately initiate splanchnic vasoconstrictors (terlipressin or norepinephrine) combined with IV albumin. Concurrently, perform a diagnostic paracentesis to rule out Spontaneous Bacterial Peritonitis (SBP) as the precipitating infectious trigger for his acute decompensation and hepatic encephalopathy. Clinical Assessment and Differentiation (HRS-AKI vs. ATN) Differentiating HRS-AKI from ATN is highly challenging but critical because specific pharmacologic management diverges significantly [44][50][52]. The patient meets the International Club of Ascites (ICA) criteria for AKI (creatinine increase ≥0.3 mg/dL or ≥50% from baseline) [35][51]. Unreliability of traditional indices: A urine sodium <10 mmol/L traditionally suggests a volume-responsive or functional pre-renal state like HRS [32][43]. However, fractional excretion of sodium (FENa) and urinary sodium are frequently inaccurate in decompensated cirrhosis [39][40]. Patients with ATN and cirrhosis can still present with low urinary sodium due to intense underlying systemic and renal vasoconstriction [32][44]. Urine microscopy: HRS-AKI is a functional disorder lacking substantial structural alterations in kidney histology [16][28][54]. A bland urine sediment supports HRS [6][44]. The presence of muddy brown granular casts, tubular epithelial cells, or bile cast nephropathy strongly points to ATN [8][29][43]. Urinary Biomarkers: Urine NGAL (uNGAL) is the most robust and clinically validated biomarker to distinguish ATN from HRS [4][8][11][34][36][45]. Cut-off values ranging between 220 and 244 μg/g creatinine highly discriminate ATN (high uNGAL) from HRS-AKI (low uNGAL) [11][33][36][45]. Other emerging biomarkers: Elevated urinary Interleukin-18 (IL-18), Kidney Injury Molecule-1 (KIM-1), Liver-type Fatty Acid Binding Protein (L-FABP), and Toll-like receptor-4 are associated with ATN and structural damage rather than functional HRS [4][10][17][38]. Fractional excretion of urea (FEUrea) may also aid differentiation, though uNGAL is superior [8].
FeatureHRS-AKIATN
PathophysiologyFunctional extreme renal vasoconstrictionIntrinsic structural tubular damage
Response to VolumeRefractory to standard IV albuminRefractory
Urine SedimentBland / NormalMuddy brown granular casts
Urinary SodiumTypically <20 mmol/L (unreliable)Often higher, can be low in cirrhosis
uNGAL LevelLow (<220 μg/g creatinine)High (>220–244 μg/g creatinine)
Priority Management Strategy Immediate medical management must target the reversal of splanchnic vasodilation and extreme renal vasoconstriction while aggressively identifying the precipitating event. Precipitant Identification (Critical Step): Systemic inflammation and bacterial translocation are primary drivers of HRS-AKI [7][18][26][28]. The patient's worsening confusion (hepatic encephalopathy) and ascites mandate an immediate diagnostic paracentesis to rule out SBP [30][31]. Initiate empiric broad-spectrum antibiotics if ascitic neutrophils are ≥250 cells/mm³ or if clinical suspicion for severe infection remains high [7][30]. Vasoconstrictor Therapy: Initiate continuous IV terlipressin, the FDA-approved first-line therapy for HRS-AKI [4][8][21][27]. Terlipressin via continuous infusion provides a better safety profile with equivalent efficacy compared to IV boluses [4][45]. If terlipressin is unavailable or contraindicated, norepinephrine administered in an intensive care setting is the standard alternative [27][31]. Volume Expansion: Vasoconstrictors must be co-administered with IV human albumin (typically 1 g/kg/day up to 100 g/day for the first 2 days, followed by 20–40 g/day) to expand effective arterial blood volume [8][21][28][31][54]. Discontinuation of Nephrotoxins: Immediately stop all diuretics, non-selective beta-blockers, and ensure absolute avoidance of NSAIDs, aminoglycosides, and potentially nephrotoxic contrast agents [6][16][24][30][32]. Definitive Therapy: Medical therapy acts as a bridge; HRS-AKI carries a poor long-term prognosis without transplantation [18][48][54]. Initiate a rapid evaluation for liver transplantation (LT) or simultaneous liver-kidney transplantation (SLKT), depending on the duration of renal failure and degree of suspected intrinsic kidney recovery [14][18][36][48]. Contraindications for Terlipressin: Do not use terlipressin in patients with ongoing coronary ischemia, severe peripheral artery disease, or severe fluid overload (e.g., respiratory failure), as it can precipitate ischemic necrosis or acute pulmonary edema [6][27][29]. Monitor closely for signs of digital/myocardial ischemia and fluid overload [27].
Q

32 year old woman with a 2cm thyroid nodule found incidentally on CT. TSH is normal. She is anxious and requesting immediate surgical removal. What is your recommended approach?

Refuse immediate surgical resection; an expedited but structured diagnostic evaluation is mandatory before any operative intervention. Order a dedicated high-resolution neck ultrasound (US) immediately to assess the sonographic characteristics of the 2 cm nodule and survey the cervical lymph nodes. Defer fine-needle aspiration (FNA) and surgical decisions until the US risk stratification is complete. Counsel the patient that most thyroid nodules are benign and unnecessary surgery carries permanent risks, including lifelong hormone replacement and nerve injury. Clinical Reasoning and Initial Assessment The patient presents with an incidental thyroid nodule (ITN) detected on CT imaging. A diagnostic workup is strictly indicated based on the nodule size (2 cm, which exceeds the standard 1 cm threshold for evaluation) and the patient's young age (<35 years) [3][14][15][20]. Because the serum TSH is normal, the nodule is nonfunctioning (euthyroid) [2][3][6][8][10][12][16][23]. A normal TSH rules out a toxic adenoma and eliminates the need for a radionuclide uptake scan [2][3][10]. A normal or elevated TSH correlates with a higher baseline risk of malignancy compared to a suppressed TSH, making morphological evaluation the critical next step [3][8][11]. Immediate prophylactic thyroidectomy without cytologic evidence of malignancy is strongly contraindicated. Performing surgery based solely on patient anxiety exposes the patient to unjustified risks of permanent surgical complications, including a 0.6% to 1.3% risk of permanent unilateral recurrent laryngeal nerve palsy and the risk of permanent hypoparathyroidism [9][25]. Mandatory Next Diagnostic Steps Order a dedicated, high-resolution thyroid ultrasound to characterize the nodule's composition, echogenicity, margins, shape (e.g., taller-than-wide), and presence of microcalcifications [2][6][8][11][16][18][19]. Evaluate all cervical lymph node compartments sonographically to rule out regional metastasis [2][3][6][18]. Compare the ultrasound findings against standard risk-stratification guidelines (e.g., ATA guidelines or ACR TI-RADS) to determine the indication for an ultrasound-guided FNA [8][12][22][25]. Perform an FNA if the 2 cm nodule exhibits high, intermediate, or low suspicion sonographic patterns; pure cystic or spongiform nodules of this size generally do not require biopsy unless clinically symptomatic [8][11][12][16][25]. Patient Counseling and Anxiety Management Thyroid nodules are exceedingly common, detectable in up to 50% to 68% of the general population, and approximately 90% to 95% are entirely benign [3][10][12][16][19]. Surgical intervention for a benign nodule is reserved exclusively for compressive symptoms (dysphagia, dyspnea), rapid growth, or cosmetic deformity, rather than asymptomatic presentation [4][6]. If surgery is eventually required, up to 50% to 80% of patients who undergo a lobectomy do not require long-term thyroid hormone replacement, preserving native endocrine function [9]. Evidence-Based Management by Cytology (Post-FNA) Once the FNA cytology is reported using the Bethesda System for Reporting Thyroid Cytopathology, management must proceed according to established guidelines. Bethesda II (Benign): active surveillance with serial ultrasounds; TSH suppressive therapy is not recommended [4][18][19]. Bethesda III/IV (Indeterminate): molecular testing (e.g., BRAF, RAS mutations) to refine the risk; positive high-risk mutations often warrant surgical intervention [4][16][18][22]. Bethesda V/VI (Suspicious/Malignant): appropriate surgical planning based on nodule characteristics, contralateral lobe findings, and patient preference [11][22][25]. Sonographic Features Driving Biopsy Decisions
Sonographic FeatureMalignancy RiskAction (2 cm Nodule)
Solid, hypoechoicHigh suspicionFNA strongly recommended
Microcalcifications, irregular marginsHigh suspicionFNA strongly recommended
Taller-than-wide shapeHigh suspicionFNA strongly recommended
Isoechoic/hyperechoic, solidLow–intermediateFNA recommended (>1.5 cm)
Spongiform / purely cysticVery low / BenignObservation without FNA

Training Methodology

A high-level overview of our approach. Specific architectural details remain proprietary.

EvidenceMD is developed through a proprietary multistage pipeline. The first stage involves extended pretraining on a large scale curated corpus of biomedical literature, clinical practice guidelines, and deidentified case repositories, building deep domain specific language representations that go well beyond what general purpose models acquire.

The second stage applies supervised finetuning on carefully curated clinical cases with structured reasoning traces. Each training example captures the full diagnostic workflow, from history interpretation through differential generation, evidence weighing, and guideline concordant management planning.

The third stage uses a proprietary reward based optimization method that evaluates model outputs against multidimensional clinical quality criteria, independently assessing diagnostic accuracy, safety behavior, evidence grounding, and communication clarity. Rather than collapsing these into a single scalar, each dimension is optimized independently to prevent tradeoff conflicts.

The final stage applies preference based alignment using curated comparison data to refine behavior across clinical correctness, appropriate hedging, empathetic communication, and safe deferral. This stage explicitly includes scenarios where the model must respectfully correct clinically incorrect assumptions, prioritizing patient safety over user agreement.

Clinical Optimization Objective
J(θ) = Ex∼D[ Σi wi · Riθ(x)) − λ · DKLθ ‖ πref) ]

where Ri represents independent clinical quality dimensions with learned weights wi, and λ controls divergence from the reference policy.

Architecture
Proprietary transformer based model
Domain Data
Large scale curated biomedical and clinical corpus
Finetuning
Supervised training on structured clinical reasoning traces
Optimization
Multistage reward based and preference based alignment
Quality Eval
Multidimensional clinical rubrics (accuracy, safety, empathy)
Safety Testing
Extensive adversarial testing and red teaming
Evidence Layer
Retrieval augmented generation over 50M+ publications
Deployment
Optimized inference with continuous monitoring

Safety & Alignment

Clinical AI must clear a higher bar for safety than general purpose models.

Appropriate Deferral

EvidenceMD recommends professional consultation for high acuity scenarios rather than providing definitive diagnoses. It is trained to say "I'm not sure" when the evidence is insufficient, rather than fabricating a confident sounding answer.

Red Flag Recognition

The model identifies and escalates emergency symptoms such as chest pain with radiation, sudden severe headache, and signs of sepsis, using appropriate urgency language and explicit recommendations to seek immediate care.

Medication Safety

EvidenceMD cross references drug interactions, contraindications, and dosage ranges. It refuses to provide dosing for high risk medications without appropriate clinical confirmation.

Sycophancy Reduction

The model scored 73% lower on sycophancy benchmarks than leading general purpose models. When a user's self diagnosis is incorrect, EvidenceMD provides evidence based corrections rather than agreement.

What's Next

We are continuing to evaluate EvidenceMD across additional medical specialties and clinical domains. An upcoming release of EvidenceMD Pro will include enhanced deep search capabilities over 50M+ medical publications.

On the roadmap is multimodal clinical reasoning, integrating medical imaging interpretation across radiology, pathology, and dermatology. We also plan to publish detailed evaluation results and methodology in a peer reviewed medical informatics journal.

References

Chen, S. et al. HealthBench: Evaluating Large Language Models Towards Improved Human Health. OpenAI, arXiv:2505.08775, May 2025.

EvidenceMD Research. Evidence Based Clinical Reasoning Benchmark: Multi Step Diagnostic Evaluation for Medical AI. Internal Technical Report, 2026.

Singhal, K. et al. Large language models encode clinical knowledge. Nature, 2023.

Shao, Z. et al. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv:2402.03300, 2024.

Rafailov, R. et al. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS, 2023.

Thirunavukarasu, A. et al. Large language models in medicine. Nature Medicine, 2023.

AASLD Practice Guidelines. Hepatology, 2024.

Wei, J. et al. Chain of Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS, 2022.

Experience EvidenceMD

Try the model that's setting new standards in clinical AI. Evidence based answers backed by peer reviewed research.

Start Free Trial
EvidenceMD Achieves SOTA on HealthBench Hard - Clinical AI Benchmark Results