Medical AI Diagnosis: Research Strategies, Clinical Trials, Accuracy Metrics, and LLM Evaluation Trends in 2026
Executive Summary
Medical AI diagnosis has evolved into a mature field employing three primary strategies: image-based classification achieving 86-98% accuracy, text-based clinical reasoning using Large Language Models (LLMs) with 52-86% performance, and multimodal systems combining both modalities for superior results8(https://therightgpt.com/ai-medical-diagnosis-tools-in-2026/)[[9]](https://www.scispot.com/blog/ai-diagnostics-revolutionizing-medical-diagnosis-in-2025)[[17]](https://www.nature.com/articles/s41746-025-01543-z). As of February 2026, the field demonstrates both remarkable technical capabilities and significant real-world implementation challenges.
Key Findings:
-
Text-based approaches substantially outperform image-only analysis when performed by LLMs, achieving 98% accuracy on radiology reports versus 84% for direct chest X-ray image analysis, though multimodal integration provides incremental improvements47(https://formative.jmir.org/2025/1/e77482).
-
Layman vs. expert query handling reveals a critical gap: While AI systems achieve 94.9% accuracy on medical exams when operating alone, laypeople using the same LLMs achieve only <34.5% accuracy due to interaction failures, non-clinical language reducing accuracy by 7-9%, and over-reliance on AI recommendations28(https://www.nature.com/articles/s41591-025-04074-y)[[44]](https://news.mit.edu/2025/llms-factor-unrelated-information-when-recommending-medical-treatments-0623).
-
Real-world deployment has surged dramatically: 66% of U.S. physicians now use AI tools (up from 38% in 2023), but massive disparities exist with large hospitals at 90-96% adoption versus small hospitals at 53-59%, and workflow integration—not technology—remains the primary barrier105(https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf)[[104]](https://intuitionlabs.ai/articles/ai-adoption-us-hospitals-2025).
-
Patient outcome evidence shows measurable benefits: Sepsis detection AI reduces mortality by 20-33.5% and detects deterioration 6 hours earlier than standard care, while breast cancer screening AI increases detection rates by 20-27% without increasing false alarms110(https://hub.jhu.edu/2022/07/21/artificial-intelligence-sepsis-detection/)[[111]](https://www.biorxiv.org/content/10.1101/224014)[[114]](https://www.breastcancer.org/screening-testing/artificial-intelligence).
-
LLM benchmark evolution for 2025-2026: DeepSeek-R1 and Claude 3.5 Sonnet dominate recent evaluations, with reasoning models systematically outperforming general-purpose LLMs by 23+ percentage points on rare disease diagnosis, though generative AI still performs 15.8% worse than expert physicians overall37(https://www.nature.com/articles/s41586-025-10097-9)[[17]](https://www.nature.com/articles/s41746-025-01543-z).
1. Current Medical Research Strategies — Three Distinct Diagnostic Modalities
Image-Based Classification
Image-based diagnostic systems represent the most clinically advanced AI application, analyzing X-rays, MRIs, CT scans, ultrasounds, pathology slides, and dermatological images using convolutional neural networks (CNNs) and transformer architectures8(https://therightgpt.com/ai-medical-diagnosis-tools-in-2026/)[[9]](https://www.scispot.com/blog/ai-diagnostics-revolutionizing-medical-diagnosis-in-2025)[[13]](https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=).
Radiology Performance Benchmarks:
| Application | Sensitivity | Specificity | AUC | Source |
|---|---|---|---|---|
| Lung nodule detection (MGH-MIT) | 94% AI vs 65% radiologists | — | — | 9(https://www.scispot.com/blog/ai-diagnostics-revolutionizing-medical-diagnosis-in-2025) |
| Breast cancer detection (South Korea) | 90% AI vs 78% radiologists | — | — | 9(https://www.scispot.com/blog/ai-diagnostics-revolutionizing-medical-diagnosis-in-2025) |
| Mammography (Transpara v.1.7.0) | 8.4% increase vs double-reading | Stable across densities | — | 2(https://www.crescendo.ai/news/ai-in-healthcare-news) |
| Diabetic retinopathy (meta-analysis) | Comparable to expert ophthalmologists | Comparable to expert ophthalmologists | — | 13(https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=) |
| Colorectal cancer (real-time endoscopy) | 97.3% | 99.0% | 0.975 | 3(https://pmc.ncbi.nlm.nih.gov/articles/PMC12455834/) |
Blood Cell Analysis Breakthrough: CytoDiffusion, published in Nature Machine Intelligence (November 2025), employs diffusion-based generative classification rather than conventional discriminative models34(https://www.nature.com/articles/s42256-025-01122-7). Trained on over 500,000 blood smear images from Addenbrooke’s Hospital, the system achieved 90.5% sensitivity and 96.2% specificity for blast cell detection (leukemia-associated cells)34(https://www.nature.com/articles/s42256-025-01122-7)[[33]](https://scienceblog.com/ai-fools-blood-experts-with-fake-cells-then-outdiagnoses-them/). Remarkably, ten experienced hematologists with 1-34 years of experience achieved only 52.3% accuracy in distinguishing CytoDiffusion-generated synthetic images from real images—essentially random guessing—demonstrating the system’s superior metacognitive awareness31(https://gigazine.net/gsc_news/en/20251130-ai-blood-analyzer-leukemia/)[[32]](https://news.ssbcrack.com/new-ai-system-enhances-blood-cell-analysis-for-improved-leukemia-diagnosis/).
Text-Based Clinical Reasoning Systems
Large Language Models applied to medical diagnosis process structured text from patient records, clinical histories, and symptom descriptions using natural language processing to extract information from electronic health records (EHRs)10(https://news.cuanschutz.edu/dbmi/health-ai-tools-support-clinicians)[[11]](https://www.nature.com/articles/d41586-026-00290-9)[[13]](https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=).
DeepRare: Multi-Agent System for Rare Disease Diagnosis
Published in Nature (February 18, 2026), DeepRare represents a breakthrough in LLM-driven medical reasoning37(https://www.nature.com/articles/s41586-025-10097-9). The system employs a three-tier hierarchical architecture integrating over 40 specialized tools with web-scale medical knowledge:
- Architecture: Tier 1 central host powered by DeepSeek-V3 LLM; Tier 2 comprises six specialized agent servers (Phenotype Extractor, Knowledge Searcher, Case Searcher, Phenotype Analyzer, Genotype Analyzer, Disease Normalizer); Tier 3 integrates heterogeneous external medical knowledge sources36(https://lifespan.io/news/ai-tool-sets-new-standard-in-diagnosing-rare-diseases/)[[37]](https://www.nature.com/articles/s41586-025-10097-9)
DeepRare Performance Metrics:
Tested across 9 datasets spanning 6,401 clinical cases, 2,919 distinct rare diseases, and 14 medical specialties35(https://www.sixthtone.com/news/1018223/china%E2%80%99s-deeprare-ai-aims-to-speed-rare-disease-diagnosis)[[36]](https://lifespan.io/news/ai-tool-sets-new-standard-in-diagnosing-rare-diseases/):
- HPO-based evaluation: Average Recall@1 of 57.18% and Recall@3 of 65.25%, outperforming the second-best method (Claude-3.7-Sonnet-thinking) by 23.79 and 18.65 percentage-point margins35(https://www.sixthtone.com/news/1018223/china%E2%80%99s-deeprare-ai-aims-to-speed-rare-disease-diagnosis)[[37]](https://www.nature.com/articles/s41586-025-10097-9)
- Multimodal (phenotype + genetic data): Achieved Recall@1 of 69.1% on Xinhua Hospital dataset compared with Exomiser’s 55.9%37(https://www.nature.com/articles/s41586-025-10097-9)
- Physician comparison: In direct comparison with rare disease physicians (163 cases), DeepRare achieved Recall@1 of 64.4% versus physicians’ 54.6%, and Recall@5 of 78.5% versus 65.6%35(https://www.sixthtone.com/news/1018223/china%E2%80%99s-deeprare-ai-aims-to-speed-rare-disease-diagnosis)[[36]](https://lifespan.io/news/ai-tool-sets-new-standard-in-diagnosing-rare-diseases/)
- Reasoning chain validation: Experts reviewing reasoning chains found 95.4% agreement between system’s evidence assessments and clinical experts’ evaluations35(https://www.sixthtone.com/news/1018223/china%E2%80%99s-deeprare-ai-aims-to-speed-rare-disease-diagnosis)
Multimodal Integration
Rather than analyzing single data modalities, cutting-edge systems integrate heterogeneous data sources: medical images, clinical records, laboratory results, genetic information, and demographic variables9(https://www.scispot.com/blog/ai-diagnostics-revolutionizing-medical-diagnosis-in-2025)[[13]](https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=). However, research reveals a critical finding: multimodal models tend to rely more heavily on textual cues than visual features when both are available46(https://www.medrxiv.org/content/10.1101/2024.08.31.24312878).
GPT-4V Performance on NEJM Image Challenges (93 cases)48(https://www.medrxiv.org/content/10.1101/2023.11.01.23297938):
- Multimodal (text + images): 80.6% raw accuracy, 70.5% adjusted
- Text only: 66.7% raw, 54.3% adjusted
- Image only: 45.2% raw, 29.3% adjusted
Multimodal input significantly outperformed text-only (p=.046) and image-only (p<.001), with similar performance across radiological, clinical, and pathological image categories (p=.91)48(https://www.medrxiv.org/content/10.1101/2023.11.01.23297938).
2. Major Clinical Trials and Studies — Evidence from Real-World Implementation
Randomized Controlled Trials
Active RCTs as of February 2026:
| Trial Name | Focus | Primary Outcome | Status | Trial ID |
|---|---|---|---|---|
| DOLCE | Lung cancer prediction AI on 5-30mm pulmonary nodules | Discharge decisions + health economics | Enrolling at 10 UK sites through Aug 2025 | NCT0538977482(https://clinicaltrials.gov/study/NCT05389774) |
| VeriSee AI-assisted screening | Diabetic retinopathy and age-related macular degeneration | Detection rates and ophthalmology referral | Recruiting April 2025–Dec 2027 (Taiwan) | NCT0706964783(https://clinicaltrials.gov/study/NCT07069647) |
| IDEAL Study | Intracranial aneurysm detection | Coprimary: superior sensitivity (0.75 vs 0.65) + non-inferior specificity | 6,450 participants, 25 Chinese hospitals | 95(https://link.springer.com/article/10.1186/s13063-024-08184-9) |
| EMPOWER-PLUS | Type 2 diabetes management with AI-powered nudges | HbA1c level difference at 9 months | 3-arm RCT, Singapore | NCT0621452093(https://clinicaltrials.gov/study/NCT06214520) |
Meta-Analysis of AI Early Warning Systems (2025):
A systematic meta-analysis of five prospective clinical validation studies found AI-powered early warning systems for predicting in-hospital clinical deterioration achieved 31% relative mortality reduction (OR 0.69, 95% CI 0.60–0.79)80(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-025-03048-x). Additionally, hospital length of stay was significantly shortened by an average of 0.35 days (95% CI [-0.68, -0.01], p = 0.04)80(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-025-03048-x).
Critical Context: The meta-analysis included only five prospective studies after screening 3,787 articles, indicating extremely limited high-quality evidence for mortality reduction despite decades of AI development80(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-025-03048-x).
Imaging Accuracy in Large-Scale Clinical Studies
Lung Cancer Detection Meta-Analysis (209 diagnostic studies)115(https://www.nature.com/articles/s41698-025-01095-1):
- Pooled sensitivity: 0.86 (95% CI: 0.84–0.87)
- Pooled specificity: 0.86 (95% CI: 0.84–0.87)
- AUC: 0.92 (95% CI: 0.90–0.94)
Breast Cancer Screening Clinical Trials:
- Swedish study (>80,000 women): AI group detected 20% more cancers than radiologist-only group114(https://www.breastcancer.org/screening-testing/artificial-intelligence)
- German and U.S. study (nearly 1.2 million mammograms): Radiologist-AI collaboration was 2.6% better at detecting breast cancer than radiologist alone114(https://www.breastcancer.org/screening-testing/artificial-intelligence)
- Korea national screening study (24,543 women, 140 detected cancers): AI-CAD increased cancer detection rate (CDR) by 13.8% (5.70‰ versus 5.01‰, p<0.001) for breast-imaging subspecialists without increasing recall rate, additionally detecting 6 ductal carcinoma in situ and 11 invasive cancers113(https://pubmed.ncbi.nlm.nih.gov/PMC12463597)
Text-Based Diagnosis Performance in Clinical Settings
Rheumatologic Disease Prediction (8,454 patients’ self-written symptom descriptions, maximum 200 words)49(https://pmc.ncbi.nlm.nih.gov/articles/PMC12456275/):
| Disease | AUC-ROC | Optimized Specificity | Misdiagnosis Rate |
|---|---|---|---|
| Osteoarthritis | 0.68 (CI: 0.65-0.70) | 0.82 | 17% |
| Fibromyalgia | 0.75 (CI: 0.72-0.78) | 0.92 | 8% |
| Immune-mediated rheumatic diseases | 0.69 (CI: 0.60-0.69) | Sensitivity: 0.92, NPV: 0.77 | — |
Acute Respiratory Infection Symptoms (multi-center emergency department study)50(https://www.medrxiv.org/content/10.1101/2024.12.16.24319044):
GPT-4 Turbo achieved F1-scores of 91.8% at Site 1 and 94.0% at independent Site 2 validation, substantially outperforming conventional ICD-10 coding (F1-score of 45.1% at Site 1 and 27.4% at Site 2)50(https://www.medrxiv.org/content/10.1101/2024.12.16.24319044).
3. Layman vs Expert Query Performance — A Critical Gap in Real-World Utility
The Paradox: High Exam Scores, Low Real-World Success
Research reveals a striking paradox: AI models perform exceptionally well on standardized medical exams yet struggle significantly in realistic clinical conversations. The Conversational Reasoning Assessment Framework for Testing in Medicine (CRAFT-MD) study, published January 2, 2025 in Nature Medicine, found that all four large language models tested performed well on medical exam-style questions but showed significantly degraded performance in conversational interactions mimicking actual patient-clinician exchanges19(https://catalyst.harvard.edu/news/article/how-good-are-ai-clinicians-at-medical-conversations/).
Real-World Performance Gap Discovery (Nature Medicine randomized controlled trial, February 2026, 1,298 participants)28(https://www.nature.com/articles/s41591-025-04074-y):
- LLMs tested alone: 94.9% condition identification accuracy, 56.3% disposition accuracy
- Laypeople using the same LLMs: <34.5% condition identification, <44.2% disposition accuracy (no better than control group using traditional resources)
Key finding: Standard benchmarks and LLM-simulated user interactions DO NOT predict real-world layperson-LLM interaction failures. Human-LLM interaction failures emerge as the primary challenge: participants provided incomplete information, LLMs misinterpreted queries, and participants inconsistently followed recommendations28(https://www.nature.com/articles/s41591-025-04074-y).
Impact of Non-Clinical Language on Accuracy
MIT Research Findings (June 2025)44(https://news.mit.edu/2025/llms-factor-unrelated-information-when-recommending-medical-treatments-0623)[[45]](https://www.nature.com/articles/s43856-024-00717-2):
Nonclinical information—including typos, extra white space, colorful language, uncertain phrasing, and informal tone—reduce the accuracy of treatment recommendations in LLM systems by 7-9 percent compared to formally structured clinical text44(https://news.mit.edu/2025/llms-factor-unrelated-information-when-recommending-medical-treatments-0623). These variations disproportionately affect recommendations for female patients, with models making approximately 7% more errors in treatment recommendations for women despite removing gender cues from the clinical context44(https://news.mit.edu/2025/llms-factor-unrelated-information-when-recommending-medical-treatments-0623).
When patients describe symptoms in colloquial language rather than clinical terminology, LLMs tend to recommend less aggressive intervention (self-management at home rather than clinic visits), potentially representing a safety concern44(https://news.mit.edu/2025/llms-factor-unrelated-information-when-recommending-medical-treatments-0623)[[45]](https://www.nature.com/articles/s43856-024-00717-2).
Patient-Generated Questions vs. Expert Queries
NYU Langone Study (July 16, 2024, JAMA Network Open)18(https://nyulangone.org/news/ai-tool-successfully-responds-patient-questions-electronic-health-record):
GPT-4 tested against primary care physicians responding to 344 real patient Electronic Health Record (EHR) queries showed:
- Accuracy, completeness, and relevance: No statistical difference between AI and human providers
- Understandability and tone: AI outperformed by 9.5 percent
- Empathy: AI responses 125% more likely to be considered empathetic
- Positivity and affiliation: AI 62% more likely to use positive language
Critical limitation: AI responses were 38% longer and 31% more likely to use complex language than human responses, with AI writing at an eighth-grade reading level compared to physicians’ sixth-grade level, suggesting AI may over-explain for layperson audiences18(https://nyulangone.org/news/ai-tool-successfully-responds-patient-questions-electronic-health-record).
Expert-Level Performance Comparison
Cross-sectional study (Italy, France, Spain, Portugal, December 2022–February 2024)20(https://pmc.ncbi.nlm.nih.gov/articles/PMC12190018/):
17,144 physicians compared with GPT-4-turbo on national medical exam questions:
- AI virtual assistant: 72–96% accuracy
- Physician accuracy: 46–62%
- Statistically significant difference (logistic regression, p < 0.001; adjusted OR = 7.46, 95% CI 6.43–8.66)
Domain-Specific Variation: In paediatrics specifically, physicians tended to outperform AI (52% vs. 45% correct answers), though not statistically significant20(https://pmc.ncbi.nlm.nih.gov/articles/PMC12190018/). This suggests domain-specific limitations in AI performance, particularly in clinical areas requiring nuanced developmental understanding.
4. Accuracy Quantification Methods — Comprehensive Evaluation Frameworks
Core Metrics for Medical AI Evaluation
Test-Based Metrics (Prevalence-Independent)16(https://link.springer.com/article/10.1007/s00330-025-11890-w)[[5]](https://pmc.ncbi.nlm.nih.gov/articles/PMC8993826/):
| Metric | Formula | Clinical Application |
|---|---|---|
| Sensitivity (Recall) | TP/(TP+FN) | Critical in screening to avoid missing disease |
| Specificity | TN/(TN+FP) | Important for confirming diseases with severe consequences |
| F1-Score | Harmonic mean of precision and recall | Strongly recommended for class imbalance (disease prevalence <40%) |
| Matthews Correlation Coefficient (MCC) | Robust statistical metric | Valuable for imbalanced datasets, invariant to class swapping |
Outcome-Based Metrics (Prevalence-Dependent)16(https://link.springer.com/article/10.1007/s00330-025-11890-w)[[5]](https://pmc.ncbi.nlm.nih.gov/articles/PMC8993826/):
- Precision (PPV): TP/(TP+FP) — proportion of positive predictions that are truly positive
- Negative Predictive Value (NPV): TN/(TN+FN) — essential for screening scenarios
- Threat Score: Well-suited for detecting rare events, excluding correctly classified negatives
Multi-Threshold Metrics16(https://link.springer.com/article/10.1007/s00330-025-11890-w):
- AUROC (Area Under ROC Curve): Measures ability to distinguish positive from negative cases; however, AUROC can be misleadingly high in low-prevalence settings
- AUPRC (Area Under Precision-Recall Curve): Prevalence-dependent, far more informative than AUROC in low-prevalence settings, offering superior real-world performance evaluation
Critical Pitfall: Low-Prevalence Performance Degradation
A CE-marked algorithm with 94% sensitivity and 95% specificity at 50% disease prevalence performs very differently at 3% actual prevalence: 63% of flagged cases are false alarms (false discovery rate = 1 − precision), potentially overwhelming clinical workflows with unnecessary follow-ups and overtreatment16(https://link.springer.com/article/10.1007/s00330-025-11890-w).
This pitfall is addressed through simultaneous reporting of F1-score, MCC, and precision, which would have revealed an F1-score of 53%, MCC of 57%, and precision of 37%—clearly indicating the algorithm was not advisable for deployment16(https://link.springer.com/article/10.1007/s00330-025-11890-w).
Novel Evaluation Frameworks for Medical AI
Relative Expert-Informed Metrics (RPAD/RRAD)6(https://arxiv.org/html/2509.11941v2):
These metrics compare AI outputs against multiple expert opinions rather than a single reference standard, normalizing performance against inter-expert disagreement:
- DeepSeek-V3: Achieved relative precision and recall metrics of 1.15 and 1.16 (average variants), surpassing other models
- A value above 1.0 indicates the algorithm performs at least as well as expert consensus
Clinically-Informed Weighted Evaluation (MAX-EVAL-11 benchmark for ICD-11)4(https://www.medrxiv.org/content/10.1101/2025.10.30.25339130v1.full.pdf):
- 5-level hierarchical matching: Exact match (1.0), parent (0.9), grandparent (0.8), great-grandparent (0.7), chapter (0.6)
- Clinical context weighting: Primary diagnoses (1.0), procedures (0.9), secondary diagnoses (0.8), comorbidities (0.6), external causes (0.4)
- Final weighted score: 0.5·EM + 0.3·CP + 0.15·PE + 0.05·HM
Conformal Prediction for Uncertainty Quantification39(https://www.nature.com/articles/s41467-022-34945-8)[[38]](https://pmc.ncbi.nlm.nih.gov/articles/PMC12092326/):
For prostate cancer diagnosis, conformal prediction at 99.9% confidence achieved 1 error in 794 predictions (0.1%) compared to 14 errors (2%) without CP, with AUC of 99.9% when combining AI predictions with human review of flagged uncertain cases39(https://www.nature.com/articles/s41467-022-34945-8).
STARD-AI Reporting Guidelines
Published in Nature Medicine (September 2025), STARD-AI represents the first comprehensive, globally-consensus standards for transparent reporting of AI-centered diagnostic test accuracy studies40(https://pubmed.ncbi.nlm.nih.gov/40954311/)[[41]](https://pmc.ncbi.nlm.nih.gov/articles/PMC8240576/). Developed through a multistage process involving over 240 international stakeholders, the guideline includes 18 new or modified items addressing:
- Dataset practices
- AI index test evaluation methods
- Algorithmic bias and fairness considerations
Current Reporting Gaps (51 AI diagnostic accuracy abstracts from international pathology conferences)42(https://pmc.ncbi.nlm.nih.gov/articles/PMC9576989/):
- Registration number and registry name: 0 of 51 (0%) — most critical gap
- Estimates with precision (95% CI): 8 of 51 (16%)
- Type of participant series: 8 of 51 (16%)
- Eligibility criteria and data collection settings: 9 of 51 (18%)
5. Current Trends in Comparing LLMs for Medical Queries — 2025-2026 Benchmarking Evolution
Emerging Evaluation Benchmarks
MedR-Bench (Nature Communications, November 2025)29(https://www.nature.com/articles/s41467-025-64769-1):
Comprises 1,453 structured patient cases from clinical case reports, spanning 13 body systems and 10 specialties with 656 cases dedicated to rare diseases. Rather than multiple-choice questions, it evaluates three critical stages:
- Examination Recommendation: DeepSeek-R1 achieves highest recall at 43.61% in 1-turn settings, while precision scores range from 22%-41%
- Diagnostic Decision-Making: Models exceed 85% accuracy when sufficient examination results are available
- Treatment Planning: The most challenging task with all models achieving only 27%-30% accuracy (versus 36.67% human physician baseline)
Reasoning Quality Metrics29(https://www.nature.com/articles/s41467-025-64769-1):
- Efficiency: Proportion of reasoning steps contributing new insights (DeepSeek-R1: 98.59%)
- Factuality: Alignment with established medical knowledge (most LLMs: 90%-98%, no model achieves perfect 100%)
- Completeness: Reference case reasoning steps included in generated output (typically 66%-78%, indicating critical omissions)
Clinical Safety-Effectiveness Dual-Track Benchmark (CSEDB) (npj Digital Medicine, December 2025)27(https://www.nature.com/articles/s41746-025-02277-8):
Encompasses 2,069 open-ended clinical scenario Q&A items developed by 32 specialist physicians across 26 clinical departments and 11 priority patient populations:
- Average total score: 57.2% (±24.5%)
- Safety scoring lower at 54.7% (±26.1%) than effectiveness at 62.3% (±22.3%)
- 13.3% performance drop in high-risk scenarios (risk levels 4-5) compared to ordinary-risk scenarios (p<0.0001)
Domain-Specific vs. General-Purpose Models: MedGPT (medical-specific) scored 15.3% higher than the second-best model overall and 19.8% higher in safety dimensions27(https://www.nature.com/articles/s41746-025-02277-8).
Head-to-Head LLM Comparisons (2025-2026)
Cardiology Multiple-Choice Performance (83 French national questions, May 2025)22(https://pmc.ncbi.nlm.nih.gov/articles/PMC12802300/):
| Model | Accuracy | Strengths |
|---|---|---|
| Claude | 78.31% (highest) | Heart failure (100%), Arrhythmias (90.9%) |
| ChatGPT-4 | 75.90% | Diagnostic investigations (87.5%) |
| Gemini | 75.90% | — |
| Mistral | 72.29% | — |
| Perplexity | 68.67% | — |
Thailand National Medical Licensing Examination24(https://www.medrxiv.org/content/10.1101/2024.12.20.24319441):
| Model | Overall Accuracy |
|---|---|
| GPT-4o | 88.9% (superior) |
| GPT-4 | 83.3% |
| Claude-3.5-Sonnet | 80.1% |
| Claude-3-Opus | 77.8% |
| Gemini-1.5-Pro | 72.8% |
| Gemini-1.0-Pro | 61.4% |
Radiology Board Examinations (150 questions)26(https://pubmed.ncbi.nlm.nih.gov/PMC11756834):
- GPT-4: 83.3% accuracy (125/150 questions), substantially outperforming competitors
- Tongyi Qianwen: 70.7%
- Claude: 62.0%
- Gemini Pro: 55.3%
- Bard: 54.7%
GPT-4 demonstrated exceptional proficiency in neurology (100% accuracy on 11/11 questions) and genitourinary (90.5% on 19/21 questions) but decreased in musculoskeletal (72.7%) and digestive (66.7%) categories26(https://pubmed.ncbi.nlm.nih.gov/PMC11756834).
Radiology “Diagnosis Please” Cases25(https://pmc.ncbi.nlm.nih.gov/articles/PMC11522128/):
- Claude 3 Opus: 62.0% accuracy for top three differential diagnoses (highest)
- GPT-4o: 49.4%
- Gemini 1.5 Pro: 41.0% (declined to respond to 6 of 324 questions due to safety concerns)
USMLE Performance Benchmarking
2025 Comprehensive USMLE Evaluation (376 publicly accessible questions)15(https://www.nature.com/articles/s41598-025-31010-4):
| Model | Step 1 | Step 2 CK | Step 3 |
|---|---|---|---|
| DeepSeek V3 | 89% | 93% (USMLE score 261 with semantic augmentation) | 84% (USMLE score 253 with semantic augmentation) |
| ChatGPT (GPT-4o Mini) | 87% | 85% | 80% |
| Grok 3 | 76% | 77% | 73% |
| Qwen (Qwen2.5-Max) | 71% | 78% | 72% |
Semantic Augmentation (RAG-Enhanced) Performance14(https://pmc.ncbi.nlm.nih.gov/articles/PMC12015668/):
Llama 405B Model with SCAI RAG:
- Step 1: 90.8% (vs 88.5% native)
- Step 2: 84.5% (vs 78.6% native)
- Step 3: 95.1% (vs 87.0% native) — only model achieving >95% on Step 3
SCAI confabulation (error) rates were as low as 4.9% for the 405B model on Step 3, compared to a 23.5% median error rate of clinically missed diagnoses at autopsy14(https://pmc.ncbi.nlm.nih.gov/articles/PMC12015668/).
Meta-Analysis of Generative AI Diagnostic Performance
Comprehensive meta-analysis (March 2025, 83 studies published June 2018-June 2024)17(https://www.nature.com/articles/s41746-025-01543-z):
- Generative AI models pooled overall accuracy: 52.1% (95% CI: 47.0–57.1%)
Comparing Expertise Levels:
- No significant difference vs. non-expert physicians (0.6% higher, p = 0.93)
- Significantly inferior to expert physicians (15.8% lower, 95% CI: 4.4–27.1%, p = 0.007)
Most-Evaluated Models: GPT-4 (54 articles), GPT-3.5 (40 articles), GPT-4V (9 articles)17(https://www.nature.com/articles/s41746-025-01543-z)
Critical Finding: 76% of studies at high risk of bias, primarily due to small test sets and inability to prove external evaluation17(https://www.nature.com/articles/s41746-025-01543-z).
6. Real-World Hospital Deployment — Proven Success Cases and Implementation Barriers
Physician Adoption Surge and Institutional Disparities
Physician Usage Statistics105(https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf):
- 66% of U.S. physicians using AI tools by 2024 (up from 38% in 2023) — 78% year-over-year increase
- 71% of non-federal acute-care hospitals using predictive AI integrated into EHRs in 2024104(https://intuitionlabs.ai/articles/ai-adoption-us-hospitals-2025)
Massive Geographic and Size-Based Disparities104(https://intuitionlabs.ai/articles/ai-adoption-us-hospitals-2025):
| Hospital Type | AI Usage Rate |
|---|---|
| Large (>400 beds) | 90-96% |
| Small (<100 beds) | 53-59% |
| Urban | 77-81% |
| Rural | 48-56% |
| System-affiliated | 81-86% |
| Independent | 31-37% |
Geographic Clustering79(https://www.nature.com/articles/s44360-025-00016-7):
- Pacific region: 27.4 hospitals/cluster (largest)
- East South Central: 4.6 hospitals/cluster (smallest)
- State-level: South Carolina and South Dakota highest (1.8), Idaho and Montana lowest (0.6)
Proven Success Cases
Sepsis Detection - Johns Hopkins TREWS System110(https://hub.jhu.edu/2022/07/21/artificial-intelligence-sepsis-detection/):
Published in Nature Medicine (July 2022), TREWS achieved:
- 20% reduction in sepsis mortality
- 6 hours earlier detection than standard care
- 82% accuracy vs. traditional tools <50%
- Interpretable recommendations — clinicians can see why the tool makes specific suggestions
Cabell Huntington Hospital (prospective before-and-after study, 2,296 sepsis cases)111(https://www.biorxiv.org/content/10.1101/224014):
- 33.5% reduction in sepsis-related in-hospital mortality (3.97% → 2.64%, P=0.038)
- 17.1% reduction in sepsis-related length of stay
- ML alerts fired 2 hours earlier than conventional Sepsis-1 alerts
Radiology AI - Annalise.ai in NHS England54(https://digitaldefynd.com/IQ/ai-in-healthcare-case-studies/):
Operating across six imaging networks covering approximately 2.8 million chest X-rays annually (35% of all UK chest X-rays):
- 45% improvement in diagnostic accuracy
- 12% increase in diagnostic efficiency
- 9 days reduction in average lung cancer treatment start time
- 27% increase in early-stage cancer detection rates
Massachusetts General Hospital-MIT System54(https://digitaldefynd.com/IQ/ai-in-healthcare-case-studies/):
- 94% diagnostic accuracy in detecting lung nodules
- Significantly outperformed human radiologists at 65% accuracy on the same task
Clinical Documentation - Highest Success Rate108(https://pmc.ncbi.nlm.nih.gov/articles/PMC12202002/):
- 100% of health systems have begun developing or piloting
- 60% deploying in at least limited areas, 14% with full deployment
- 53% report high degree of success — the only use case showing universal adoption and strong success rates
Stanford Health Care: AI-powered documentation tools achieved 96% physician satisfaction with time savings averaging two hours per day per physician78(https://www.johnsnowlabs.com/preparing-hospitals-for-large-scale-ai-deployments-in-2026/).
Implementation Barriers and Critical Challenges
Workflow Integration - The Primary Barrier75(https://pmc.ncbi.nlm.nih.gov/articles/PMC11393514/)[[72]](https://logicon.tech/healthcare-ai-workflow-adoption-fails/):
The most critical barrier is workflow integration, not the technology itself. A 2024 JAMA study demonstrated that when an AI-driven sepsis alert was embedded directly into EHR triage workflows, clinician response time improved by 35%, while portal-based versions of the same algorithm showed no significant improvement72(https://logicon.tech/healthcare-ai-workflow-adoption-fails/).
Technical Friction Points73(https://pubmed.ncbi.nlm.nih.gov/PMC12729787)[[74]](https://pubmed.ncbi.nlm.nih.gov/PMC11933996):
- Excessive clicks
- Alert pop-ups causing fatigue
- Lack of seamless EHR integration
- Inadequate technical support
- Among ambient AI scribe users: 81% of comments about linguistic and device accessibility were negative
Trust and Transparency Barriers75(https://pmc.ncbi.nlm.nih.gov/articles/PMC11393514/)[[105]](https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf):
Specific trust facilitators physicians require:
- 88% citing need for designated feedback channel for issues
- 85% emphasizing data privacy assurance
- 84% requiring seamless EHR integration and workflow support
- 82% requiring AI safety validated by trusted entity with monitoring over time
- 47% ranked increased regulatory oversight as #1 action needed
Data Infrastructure and Interoperability79(https://www.nature.com/articles/s44360-025-00016-7):
Feature importance analysis of 3,560 U.S. hospitals revealed that interoperability measures—specifically barriers to health information exchange—were the strongest predictors of predictive AI adoption. Hospitals with higher interoperability capabilities and larger bed capacities showed increased AI adoption, while hospitals with greater information exchange barriers were less likely to implement predictive AI79(https://www.nature.com/articles/s44360-025-00016-7).
Validation and Governance Gaps79(https://www.nature.com/articles/s44360-025-00016-7):
Among 3,560 U.S. hospitals implementing AI:
- 47.4% reported NO evaluation of model accuracy
- 51.4% reported NO evaluation of bias
- Among hospitals implementing predictive models, 56.9% were developed by hospitals’ electronic health record vendors
Physician Knowledge Gaps107(https://pmc.ncbi.nlm.nih.gov/articles/PMC9472134/)[[106]](https://hitconsultant.net/2025/12/30/physician-assistant-trends-2025-ai-adoption-title-changes-and-workforce-growth/):
- Only 13% agreed they had good knowledge of clinical AI
- 77% expressed high willingness to learn and 78% demanded training from hospitals or schools
- Among Physician Assistants: 56% now use AI daily in practice, yet 87% report needing more formal training
7. Patient Outcomes and Clinical Utility — Evidence from Rigorous Trials
Mortality Reduction Evidence
AI Early Warning Systems Meta-Analysis (2025)80(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-025-03048-x):
Systematic meta-analysis of five prospective clinical validation studies:
- 31% relative mortality reduction (OR 0.69, 95% CI 0.60–0.79)
- Significant reductions in 30-day mortality
- Hospital length of stay shortened by average of 0.35 days (95% CI [-0.68, -0.01], p = 0.04)
Critical Context: Only five prospective studies identified from 3,787 articles screened, indicating extremely limited high-quality evidence despite decades of AI development80(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-025-03048-x).
Diagnostic Accuracy Meta-Analyses
Lung Cancer Detection (209 studies)115(https://www.nature.com/articles/s41698-025-01095-1):
| Metric | Internal Validation | External Validation |
|---|---|---|
| Sensitivity | 0.86 | 0.82 |
| Specificity | 0.86 | 0.84 |
| AUC | 0.93 | 0.90 |
Deep learning algorithms achieved:
- Sensitivity: 0.87 (95% CI: 0.85–0.89)
- Specificity: 0.87 (95% CI: 0.85–0.89)
- AUC: 0.94 (95% CI: 0.91–0.95)
3D deep learning outperformed 2D approaches with 0.94 AUC vs. 0.93 AUC115(https://www.nature.com/articles/s41698-025-01095-1).
Breast Cancer Screening114(https://www.breastcancer.org/screening-testing/artificial-intelligence)[[113]](https://pubmed.ncbi.nlm.nih.gov/PMC12463597):
- Swedish study (>80,000 women): AI detected 20% more cancers
- German/U.S. study (1.2M mammograms): Radiologist-AI 2.6% better at detection
- Korea national screening (24,543 women): AI-CAD increased detection rate by 13.8% (5.70‰ vs. 5.01‰, p<0.001) without increasing recall rate
Cost-Effectiveness Evidence
Systematic Review (19 economic studies, August 2025)88(https://www.nature.com/articles/s41746-025-01722-y):
| Application | ICER/Cost Savings |
|---|---|
| Atrial fibrillation screening | £4,847–£5,544 per QALY (well below NHS £20,000 threshold) |
| Diabetic retinopathy screening | 14–19.5% per-patient cost reduction; ICERs as low as $1,107.63/QALY |
| Breast cancer screening | $23,755 per QALY (below $100,000 benchmark) |
| Colonoscopy AI | $149.2M annual savings (Japan), $85.2M (USA) |
| Medication management | ROI of 12.4:1 |
Macro-Level Economic Impact90(https://www.aei.org/health-care/ais-uncertain-cost-effects-in-health-care/)[[91]](https://www.medrxiv.org/content/10.1101/2025.10.05.25337345):
- Existing AI platforms could deliver up to $360 billion in annual cost reductions in U.S. healthcare
- Harvard-based economic assessment: $200–360 billion annual net savings over next five years (2024-2029), representing a 5-10% reduction in total U.S. healthcare spending
- Global AI healthcare market projected to expand at 38.5% CAGR from 2024 through 203089(https://aijourn.com/economicimpacthealthcare/)
NHS England Cost Savings: Documented £250 million savings between 2022 and 2024 after adopting AI chatbots and automation tools for clinical coding92(https://graphlogic.ai/blog/ai-chatbots/ai-use-cases-by-industry/ai-reduces-costs-healthcare/).
External Validation and Generalization Challenges
External Validation Deficiency81(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-024-02830-7):
Of 572 identified studies developing ML-based ICU prediction models:
- Only 84 (14.7%) were externally validated
- When ML models were applied to new hospital data, performance consistently declined:
- Average AUROC reduction of -0.037 (95% credible interval -0.052 to -0.027)
- Constituting 7–23% relative performance decrease
- 49.5% of validated studies showed substantial performance loss (>0.05 AUROC reduction)
Trial Generalizability Study (oncology, 152,402 real-world patients)109(https://www.nature.com/articles/s41591-024-03352-5):
The TrialTranslator framework evaluated generalizability of 11 landmark phase 3 RCTs:
- High-risk phenotypes: Treatment mOS point estimate averaged 62% lower than RCT results
- Full cohort emulation: 9 of 11 emulated trials had hazard ratios significantly favoring treatment, but only 5 of 11 had HRs within 95% CI of RCT estimates
- Emulated trial HRs were attenuated by average of 35% compared to RCTs
8. Major Benchmark Datasets — The Foundation for Medical AI Training
Clinical Notes and Critical Care Datasets
MIMIC Family (MIT/Beth Israel Deaconess Medical Center)76(https://pubmed.ncbi.nlm.nih.gov/PMC9810617)[[77]](https://www.medrxiv.org/content/10.1101/2023.05.18.23290207):
| Dataset | Size | Composition |
|---|---|---|
| MIMIC-III | 58,000+ hospital admissions (June 2001–Oct 2012) | ~2,083,180 deidentified clinical notes across 15 categories |
| MIMIC-IV | 145,915 patients (2008–2019) | 331,794 discharge summaries, 2,321,355 radiology reports |
| MIMIC-CXR | 64,588 patients (2011–2017) | 227,835 X-ray studies, 377,110 total radiographs |
| MIMIC-IV-MM | Multimodal integration | Joins MIMIC-IV, MIMIC-CXR, and MIMIC-IV-Note datasets |
CardioEHR Dataset (published February 14, 2026)102(https://www.nature.com/articles/s41597-026-06855-7):
Longitudinal EHR dataset from Wuhan Union Hospital containing 35,243 patients (2010–2020) and 37,975 patients (2011–2024), including structured clinical information, diagnoses, laboratory results, and socioeconomic indices.
Critical Licensing Constraint: As of September 24, 2025, PhysioNet guidance explicitly prohibits sharing credentialed health data with third-party services, including LLM APIs, because most commercial LLM services retain data by default112(https://physionet.org).
Medical Imaging Datasets
ChestX-ray14 (NIH)57(https://www.biorxiv.org/content/10.1101/841619)[[59]](https://arxiv.org/abs/1803.04565)[[60]](https://arxiv.org/abs/2105.12430):
- 112,120 frontal-view X-ray images from 30,805 unique patients
- Images stored at 1024 × 1024 pixels with 8-bit grayscale
- Published benchmark split: 86,524 training images, 25,596 test images
- Disease identification employed NLP techniques (DNorm and MetaMap) applied to radiological reports
CheXpert Dataset (Stanford)55(https://pubmed.ncbi.nlm.nih.gov/PMC10850044)[[56]](https://pubmed.ncbi.nlm.nih.gov/PMC11101488)[[58]](https://www.medrxiv.org/content/10.1101/19013342):
- 224,316 chest X-rays from 65,240 patients (Oct 2002–July 2017)
- Annotated for 14 clinically relevant observations (13 pathologies plus “No Finding”)
- Official split: 223,414 training studies, 200 validation studies (labeled by 3 board-certified radiologists), 500 test studies (labeled by 5 board-certified radiologists)
State-of-the-art Performance: Mean AUC of 0.841 across all 14 diseases for multi-branch residual attention network61(https://pubmed.ncbi.nlm.nih.gov/PMC11126636). On CheXpert’s validation dataset, methods achieve mean AUC scores ranging from 0.835 to 0.904 depending on handling of uncertain labels61(https://pubmed.ncbi.nlm.nih.gov/PMC11126636).
Dermatology Imaging: ISIC Datasets
ISIC Archive Evolution53(https://challenge.isic-archive.com/data/):
| Year | Training Images | Test Images | Total |
|---|---|---|---|
| 2016 | 900 | 379 | 1,279 |
| 2017 | 2,000 | 600 | 2,600 |
| 2018 | 10,015 | 1,512 | 11,527 |
| 2019 | 25,331 | 8,238 | 33,569 |
| 2020 | 33,126 | 10,982 | 44,108 |
MILK10k Benchmark (launched August 7, 2025)51(https://www.isic-archive.com/):
- 5,240 training lesions with paired dermoscopic and close-up images
- 479 testing lesions
- Multi-modal imaging approach
Data Quality Issues: Analysis revealed substantial duplicate images: 12,039 binary identical duplicate files across all training sets, 1,592 identical duplicates across test sets, and 2,263 downsampled duplicate images, for a total of 14,310 duplicates52(https://e-space.mmu.ac.uk/631025/7/1-s2.0-S1361841521003509-main.pdf).
Licensing: ISIC datasets are publicly accessible with CC-BY-NC (Creative Commons Attribution-Non-Commercial 4.0) licensing for 2018-2020 data53(https://challenge.isic-archive.com/data/).
Vision-Language and Multimodal Datasets
BIOMEDICA (2025)63(https://arxiv.org/abs/2503.22727):
Large-scale vision-language dataset derived from PubMed Central Open Access publications:
- Over 6 million scientific articles
- 24 million image-caption pairs
- 27 metadata fields including expert human annotations
- Spans diverse biomedical domains: clinical radiology, pathology images, research microscopy, immunoassays, chemical structures
- Open access (derived from PubMed Central Open Access)
MedSegBench (August 2024)62(https://www.medrxiv.org/content/10.1101/2024.08.26.24312619):
- 35 distinct 2D medical image segmentation datasets
- Over 60,000 images
- Spanning ultrasound, dermoscopy, MRI, X-ray, OCT, and chest X-ray
- Publicly available in standardized Numpy npz format with pre-defined train/validation/test splits
- Published under Creative Commons Licenses (CC-BY-NC, CC-BY-SA, CC-BY-NC-SA)
Genomic and Integrated Multimodal Datasets
Major Genomic Databases100(https://opendatascience.com/18-open-healthcare-datasets-2025-update/)[[101]](https://www.shaip.com/blog/healthcare-datasets-for-machine-learning-projects/):
| Database | Size | Scope |
|---|---|---|
| TCGA (The Cancer Genome Atlas) | 20,000+ samples | 33 cancer types with genomic, epigenomic, transcriptomic, proteomic data |
| UK Biobank | 500,000 volunteer participants | In-depth genetic and health information |
| 1000 Genomes Project | 2,500 individuals | 26 populations with sequencing data |
| ADNI (Alzheimer’s Disease Neuroimaging Initiative) | 1,000+ participants | MRI and PET neuroimaging, CSF and blood biomarkers, genetic information |
Multimodal Integration Trend: Current research integrates diverse data modalities—genomic, transcriptomic, proteomic, imaging, environmental, and electronic health record data—into unified analytical frameworks, with multimodal approaches enhancing predictive accuracy for critical care outcomes by 12% or more compared to single-modality approaches103(https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1743921/full)[[99]](https://www.shaip.com/blog/multimodal-medical-datasets-for-ai-research/).
9. Medical-Specific LLMs and Foundation Models — Domain-Optimized Architectures
BioGPT: Domain-Specific Generative Model
Architecture and Training68(https://arxiv.org/pdf/2210.10341)[[69]](https://medium.com/@EleventhHourEnthusiast/paper-review-biogpt-631f33f5c4ea):
Built on GPT-2 medium architecture with 347 million parameters (24 layers, 1,024 hidden size, 16 attention heads), BioGPT was pre-trained from scratch on 15 million PubMed abstracts published before 2021 using a domain-specific vocabulary of 42,384 tokens generated through byte pair encoding68(https://arxiv.org/pdf/2210.10341). This specialized vocabulary allows biomedical terms like “thrombin” to be represented as single tokens rather than being fragmented into subword units70(https://arxiv.org/html/2403.18421v1).
Performance on Biomedical Tasks68(https://arxiv.org/pdf/2210.10341):
- Relation Extraction: BC5CDR (chemical-disease) 44.98% F1, KD-DTI (drug-target) 38.42% F1, DDI (drug-drug) 40.76% F1
- Question Answering: PubMedQA 78.2% accuracy; BioGPT-Large (1.5B parameters) 81.0%
- Document Classification: HoC (Hallmarks of Cancer) 85.12% F1, outperforming BioBERT (81.54%) and PubMedBERT (82.32%)
Critical Comparative Finding (Nature Communications 2025)71(https://www.nature.com/articles/s41467-025-56989-2):
Fine-tuned domain-specific models (SOTA: macro-average 0.6536) substantially outperformed zero- and few-shot LLMs (0.4561-0.5131 macro-average) on biomedical NLP tasks, with particularly pronounced gaps in information extraction: relation extraction (SOTA 0.79 vs. best LLM zero-shot 0.33).
GatorTron: Large-Scale Clinical Language Model
Model Variants (University of Florida)86(https://www.medrxiv.org/content/10.1101/2022.02.27.22271257v2.full)[[87]](https://www.nature.com/articles/s41746-022-00742-2):
| Variant | Parameters | Layers | Hidden Dimensions | Attention Heads |
|---|---|---|---|---|
| GatorTron-base | 345M | 24 | 1,024 | 16 |
| GatorTron-medium | 3.9B | 48 | 2,560 | 40 |
| GatorTron-large | 8.9B | 56 | 3,584 | 56 |
Pre-trained on 90+ billion words including:
- 82+ billion words from de-identified University of Florida Health clinical notes (2011-2021)
- 6 billion from PubMed
- 2.5 billion from Wikipedia
- 0.5 billion from MIMIC-III
Training utilized 992 NVIDIA A100 GPUs across 124 supercomputing nodes, with GatorTron-large converging in ~7 epochs and ~6 days87(https://www.nature.com/articles/s41746-022-00742-2).
GatorTron-large Clinical Task Performance85(https://www.emergentmind.com/topics/gatortron-base)[[87]](https://www.nature.com/articles/s41746-022-00742-2):
| Task | F1-Score / Accuracy |
|---|---|
| Clinical NER (2010 i2b2) | F1=0.8996 |
| Clinical NER (2012 i2b2) | F1=0.8091 |
| Clinical NER (2018 n2c2) | F1=0.9000 |
| Medical Relation Extraction (2018 n2c2) | F1=0.9627 |
| Semantic Textual Similarity (2019 n2c2/OHNLP) | Pearson r=0.8903 (GatorTron-medium best) |
| NLI (MedNLI) | 90.20% accuracy (outperforming BioBERT by 9.6%, ClinicalBERT by 7.5%) |
| Medical QA (emrQA) | F1=0.9543 |
Error Analysis Finding: GatorTron successfully identified longer phrases and complex linguistic structures as single entities (e.g., “a mildly dilated ascending aorta” as a complete entity), whereas smaller models like ClinicalBERT identified only “mildly dilated”84(https://medium.com/@EleventhHourEnthusiast/paper-review-gatortron-b38578b077a2)[[87]](https://www.nature.com/articles/s41746-022-00742-2).
Clinical BERT Variants and Evolution
BioClinical ModernBERT (June 2025)66(https://www.lighton.ai/lighton-blogs/announcing-bioclinical-modernbert-a-new-sota-encoder-model-for-medical-nlp)[[67]](https://arxiv.org/abs/2506.10896):
The latest advancement, developed collaboratively by Dana-Farber Cancer Institute, Harvard University, LightOn, MIT, McGill University, Albany Medical College, and Microsoft Research:
- Pre-trained on over 53.5 billion tokens (largest biomedical and clinical corpus to date)
- Leveraging 20 datasets from diverse institutions, domains, and geographic regions
- Available in base (150M parameters) and large (396M parameters) versions
- Extended context length: 8,192 tokens (vs. standard BERT’s 512 tokens)
Clinical ModernBERT Performance65(https://arxiv.org/html/2504.03964v1):
| Task | Performance |
|---|---|
| EHR Classification (AUROC) | 0.9769 (vs BioBERT 0.9680) |
| Medical NER (F1) | 0.766 (BioBERT best at 0.794, competitive) |
| Retrieval (PMC-Patients NDCG@10) | 0.2167 (best performance) |
| i2b2 2006 NER | F1 0.965 |
| i2b2 2012 NER | F1 0.804 (best result) |
| i2b2 2014 NER | F1 0.966 |
Token-Aware Masking Innovation: Removing token-aware masking caused top-1 MLM accuracy to drop from 63.31% to 48.84%, with top-25 accuracy falling below 59%, demonstrating critical importance of semantic-aware masking65(https://arxiv.org/html/2504.03964v1).
Med-PaLM 2: Google’s Medical Reasoning Model
Critical Note: No information about Med-PaLM 3 was found in the research. Available evidence covers Med-PaLM and Med-PaLM 2 as of May 2023 and January 2025.
Med-PaLM vs Med-PaLM 2 Performance Comparison96(https://arxiv.org/abs/2305.09617)[[97]](https://arxiv.org/pdf/2305.09617)[[98]](https://pmc.ncbi.nlm.nih.gov/articles/PMC11922739/):
Med-PaLM was the first model to exceed “passing” score on USMLE-style questions, achieving 67.2% accuracy on MedQA. Med-PaLM 2 substantially improved this to 86.5% accuracy, representing 19.1% absolute improvement.
Med-PaLM 2 USMLE and Benchmark Performance97(https://arxiv.org/pdf/2305.09617):
| Benchmark | Accuracy/Performance |
|---|---|
| MedQA (USMLE-style) | 86.5% with ensemble refinement (79.7% few-shot, 83.7% chain-of-thought+self-consistency) |
| MedMCQA (Indian medical exams) | 72.3% |
| PubMedQA | 81.8% (with self-consistency), 75.0% baseline |
| MMLU Clinical Knowledge | 88.7% |
| MMLU Medical Genetics | 92.0% |
| MMLU Professional Medicine | 95.2% |
| MMLU College Biology | 95.8% |
Physician Evaluation on Consumer Medical Questions (140 questions)30(https://pubmed.ncbi.nlm.nih.gov/PMC10396962)[[96]](https://arxiv.org/abs/2305.09617)[[98]](https://pmc.ncbi.nlm.nih.gov/articles/PMC11922739/):
- Med-PaLM 2 answers aligned with scientific consensus 92.6% of the time (vs. Flan-PaLM 61.9%)
- Knowledge recall: 95.4% evidence of correct knowledge (vs. Flan-PaLM 76.3%, clinician-provided 97.8%)
- Safety: 90.6% of answers rated as low-risk harm (vs. 79.4% for Med-PaLM)
Pairwise Comparison (1,066 consumer medical questions)96(https://arxiv.org/abs/2305.09617)[[98]](https://pmc.ncbi.nlm.nih.gov/articles/PMC11922739/):
Physicians preferred Med-PaLM 2 answers over physician-generated answers on 8 of 9 clinical utility axes (p<0.001):
- Better reflects medical consensus 72.9% of the time compared to physician answers
- Better reading comprehension and knowledge recall
10. Critical Limitations and Future Directions — Bridging the Gap Between Research and Practice
Performance Degradation: Benchmark vs. Real-World Settings
Diagnostic Reasoning Limitations12(https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html):
In controlled research settings, large language models matched or outperformed physicians on diagnostic reasoning and treatment planning. However, when researchers modified standard medical multiple-choice questions so the correct answer became “none of the other answers”—without changing the clinical reasoning required—accuracy dropped sharply across leading AI systems, in some cases by more than a third12(https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html).
When AI systems were tested in settings more closely resembling real clinical work12(https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html):
- When models had to ask follow-up questions, performance fell
- When models had to manage incomplete information, performance fell
- When models had to revise decisions as new details emerged, performance fell
- On tests measuring reasoning under uncertainty, AI systems performed closer to medical students than experienced physicians
External Validation Consistent Degradation21(https://pmc.ncbi.nlm.nih.gov/articles/PMC12689012/)[[81]](https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-024-02830-7):
A systematic review of six recent radiology AI studies (2022-2025) found:
- Internal validation achieved AUCs ranging from 0.76 to 0.95
- External validation showed consistent degradation with median AUC drops of approximately 0.03
- Specificity drops were more pronounced (up to 24 percentage points)
When ML models were applied to new hospital data, average AUROC reduction of -0.037 occurred, constituting 7–23% relative performance decrease, with 49.5% of validated studies showing substantial performance loss81(https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-024-02830-7).
Publication Bias and Evaluation Methodology Gaps
Publication Bias94(https://www.medrxiv.org/content/10.1101/2023.09.12.23295381v1.full):
A scoping evaluation analyzing 84 randomized controlled trials of AI in clinical practice (January 2018–August 2023) found 82.1% reported positive results for their primary endpoints. This success rate is notably higher than historical rates for other medical interventions, suggesting significant publication bias94(https://www.medrxiv.org/content/10.1101/2023.09.12.23295381v1.full).
Evaluation Methodology Gaps12(https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html)[[13]](https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=):
A review of more than 500 medical AI studies found critical gaps:
- Nearly half tested models using medical exam-style questions
- Only 5% used real patient data
- Very few measured whether models recognized uncertainty
- Even fewer examined bias or fairness
Benchmark-Reality Gap: Nearly half of studies tested on curated datasets that underperform on real clinical images with greater variability; clinicians rarely have complete data, and AI performance drops when managing missing information12(https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html)[[13]](https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=).
Algorithmic Bias and Fairness Concerns
Demographic Representation13(https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=)[[43]](https://pubmed.ncbi.nlm.nih.gov/PMC12730494):
Dermatology systems show reduced diagnostic accuracy in patients with darker skin phototypes, reflecting underrepresentation in training datasets13(https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email=). Among 168 FDA-approved devices cleared in 2024, sex data appeared in only 43.5% of devices, while race/ethnicity data appeared in only 15.5%43(https://pubmed.ncbi.nlm.nih.gov/PMC12730494).
Data Scarcity and Quality Issues107(https://pmc.ncbi.nlm.nih.gov/articles/PMC9472134/):
A systematic review found 26 (74.28%) of 35 included studies highlighted lack of AI knowledge as a primary concern among participants, with specific concerns about:
- Data scarcity: Lack of high-quality datasets was the primary concern
- Absence of regulatory standards: Followed by data quality issues
- Model accuracy gaps: 47.4% of hospitals with AI models reported no evaluation of model accuracy, and 51.4% reported no evaluation of bias79(https://www.nature.com/articles/s44360-025-00016-7)
Hallucination Risks in Medical Contexts
Hallucination Rates23(https://wizey.one/alternatives/ai-models-vs-wizey/):
Significant hallucination risks exist across all models:
- GPT-4o demonstrates 15.8% hallucination rate in general contexts, while Claude 3.7 shows 16.0%
- In medical-specific scenarios, GPT-4’s hallucination rate increases to 28.6%
- Cancer information analysis without structured databases shows 19% hallucination rates for GPT-4 and 35% for GPT-3.5
Cost-Performance Trade-offs71(https://www.nature.com/articles/s41467-025-56989-2):
GPT-4 exhibited 60-100x higher costs than GPT-3.5:
- Extractive tasks: GPT-4 $2-$10 per 100 instances vs. GPT-3.5 $0.03-$0.16
- Generative tasks: GPT-4 $84.02 per 100 instances vs. GPT-3.5 $0.71
However, performance differences did not scale proportionally to cost except in reasoning tasks71(https://www.nature.com/articles/s41467-025-56989-2).
Future Research Directions
Multimodal AI Integration1(https://www.nature.com/articles/s44387-025-00011-z):
Future studies should integrate text, imaging, structured clinical data (laboratory results, genomic profiles), and longitudinal patient records into unified frameworks with improved cross-modal alignment. Benchmark datasets should reflect real-world clinical complexity.
Explainability and Clinical Validation1(https://www.nature.com/articles/s44387-025-00011-z):
Development of chain-of-reasoning frameworks visualizing diagnostic pathways, adversarial testing to quantify model uncertainty, and large-scale prospective trials comparing LLM-assisted versus conventional diagnostic workflows are essential.
Uncertainty Quantification10(https://news.cuanschutz.edu/dbmi/health-ai-tools-support-clinicians):
A major research direction addresses AI confidence calibration. The University of Colorado LARK Lab published work on uncertainty estimation methodology, proposing novel approaches for LLMs to quantify prediction confidence accurately. This addresses a critical clinical need: AI systems can sound confident while being profoundly wrong—dangerous in medicine10(https://news.cuanschutz.edu/dbmi/health-ai-tools-support-clinicians).
Real-World Validation7(https://pubmed.ncbi.nlm.nih.gov/PMC12706444):
Only 5% of 761 evaluated LLM evaluation studies assessed performance on real patient care data. The field must transition from isolated benchmarks to realistic clinical workflows, requiring systems to retrieve information, place orders, and complete multi-step workflows12(https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html).
Synthesis and Strategic Implications
As of February 2026, medical AI diagnosis has matured from experimental proof-of-concepts into clinically deployed systems demonstrating measurable patient benefits. However, the field confronts a paradox: technical capabilities far exceed real-world utility when humans interact with these systems.
Three critical insights emerge:
-
Text-based clinical reasoning dominates image interpretation when performed by LLMs (98% vs. 84% accuracy), yet multimodal integration provides incremental improvements through complementary information capture. Structured clinical protocols outperform free-form interpretation, suggesting AI works best when augmenting—not replacing—standardized clinical workflows.
-
The layman-expert performance gap reveals fundamental interaction failures: While AI achieves 95%+ accuracy on medical exams when operating alone, laypeople using the same systems achieve <35% accuracy due to non-clinical language reducing accuracy by 7-9%, incomplete information provision, and misinterpretation of AI recommendations. This gap represents the frontier challenge for clinical AI deployment.
-
Real-world implementation faces infrastructure—not technology—barriers: With 66% physician adoption but only 13% reporting good AI knowledge, and 47% of hospitals conducting no model accuracy evaluation, the field demonstrates a dangerous mismatch between rapid deployment and inadequate governance frameworks. Geographic disparities (96% large hospital adoption vs. 53% small hospital adoption) risk exacerbating healthcare inequities.
Looking forward, success requires shifting from “Where can we put AI?” to “Where will AI make humans better?“64(https://www.healthcareittoday.com/2025/12/23/ai-and-automation-in-healthcare-2026-health-it-predictions/) The evidence shows AI excels at reducing mortality in sepsis detection (20-33% reduction), improving cancer screening (20-27% detection increases), and generating $200-360 billion in annual healthcare savings. However, these benefits materialize only when AI is deeply integrated into clinical workflows with robust validation, continuous monitoring, and physician-in-the-loop oversight.
The medical AI field in 2026 stands at an inflection point: technical capabilities are proven, but sustainable clinical value depends on solving human-AI interaction challenges, closing validation gaps, and building governance frameworks that prioritize safety over deployment speed.
Sources
[1] Large language models for disease diagnosis: a scoping review | npj Artificial Intelligence - https://www.nature.com/articles/s44387-025-00011-z [2] The Latest AI News + Breakthroughs in Healthcare and Medical | News - https://www.crescendo.ai/news/ai-in-healthcare-news [3] Artificial intelligence in healthcare and medicine: clinical applications, therapeutic advances, and future perspectives - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12455834/ [4] MAX-EVAL-11: A Comprehensive Benchmark for Evaluating Large Language Models on Full-Spectrum ICD-11 Medical Coding - https://www.medrxiv.org/content/10.1101/2025.10.30.25339130v1.full.pdf [5] On evaluation metrics for medical applications of artificial intelligence - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC8993826/ [6] How to Evaluate Medical AI - https://arxiv.org/html/2509.11941v2 [7] Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks - https://pubmed.ncbi.nlm.nih.gov/PMC12706444 [8] AI Medical Diagnosis Tools in 2026 - The Right GPT - https://therightgpt.com/ai-medical-diagnosis-tools-in-2026/ [9] AI Diagnostics: Revolutionizing Medical Diagnosis in 2026 | Trends - https://www.scispot.com/blog/ai-diagnostics-revolutionizing-medical-diagnosis-in-2025 [10] Health AI in 2026: CU Researchers are Implementing Trustworthy Tools to Support Clinicians - https://news.cuanschutz.edu/dbmi/health-ai-tools-support-clinicians [11] AI succeeds in diagnosing rare diseases - https://www.nature.com/articles/d41586-026-00290-9 [12] Clinical AI Has Boomed. A New Stanford-Harvard State of Clinical AI Report Shows What Holds Up in Practice. | Department of Medicine News | Stanford Medicine - https://medicine.stanford.edu/news/current-news/standard-news/clinical-ai-has-boomed.html [13] Artificial Intelligence in Healthcare: From Diagnosis to Rehabilitation - https://www.cureus.com/articles/452999-artificial-intelligence-in-healthcare-from-diagnosis-to-rehabilitation.pdf?email= [14] Semantic Clinical Artificial Intelligence vs Native Large Language Model Performance on the USMLE - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12015668/ [15] Benchmarking large language models on the United States medical licensing examination for clinical reasoning and medical licensing scenarios | Scientific Reports - https://www.nature.com/articles/s41598-025-31010-4 [16] ESR Essentials: common performance metrics in AI—practice recommendations by the European Society of Medical Imaging Informatics | European Radiology | Springer Nature Link - https://link.springer.com/article/10.1007/s00330-025-11890-w [17] A systematic review and meta-analysis of diagnostic performance comparison between generative AI and physicians | npj Digital Medicine - https://www.nature.com/articles/s41746-025-01543-z [18] AI Tool Successfully Responds to Patient Questions in Electronic Health Record | NYU Langone News - https://nyulangone.org/news/ai-tool-successfully-responds-patient-questions-electronic-health-record [19] How Good Are AI ‘Clinicians’ at Medical Conversations? - Harvard Catalyst - https://catalyst.harvard.edu/news/article/how-good-are-ai-clinicians-at-medical-conversations/ [20] Artificial Intelligence Outperforms Physicians in General Medical Knowledge, Except in the Paediatrics Domain: A Cross-Sectional Study - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12190018/ [21] Assessing the generalizability of artificial intelligence in radiology: a systematic review of performance across different clinical settings - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12689012/ [22] Comparative study of the performance of ChatGPT-4, Claude, Gemini, Mistral, and perplexity on multiple-choice questions in cardiology - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12802300/ [23] General AI vs Medical AI: ChatGPT, Claude & Gemini vs Wizey 2026 | Wizey - AI Health Assistant - https://wizey.one/alternatives/ai-models-vs-wizey/ [24] Evaluation of Large Language Models in Thailand’s National Medical Licensing Examination - https://www.medrxiv.org/content/10.1101/2024.12.20.24319441 [25] Diagnostic performances of GPT-4o, Claude 3 Opus, and Gemini 1.5 Pro in “Diagnosis Please” cases - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC11522128/ [26] Performance Evaluation and Implications of Large Language Models in Radiology Board Exams: Prospective Comparative Analysis - https://pubmed.ncbi.nlm.nih.gov/PMC11756834 [27] A novel evaluation benchmark for medical LLMs illuminating safety and effectiveness in clinical domains | npj Digital Medicine - https://www.nature.com/articles/s41746-025-02277-8 [28] Reliability of LLMs as medical assistants for the general public: a randomized preregistered study | Nature Medicine - https://www.nature.com/articles/s41591-025-04074-y [29] Quantifying the reasoning abilities of LLMs on clinical cases | Nature Communications - https://www.nature.com/articles/s41467-025-64769-1 [30] Large language models encode clinical knowledge - https://pubmed.ncbi.nlm.nih.gov/PMC10396962 [31] CytoDiffusion, an AI blood cell analysis system, outperforms human experts in detecting leukemia - GIGAZINE - https://gigazine.net/gsc_news/en/20251130-ai-blood-analyzer-leukemia/ [32] New AI System Enhances Blood Cell Analysis for Improved Leukemia Diagnosis - SSBCrack News - https://news.ssbcrack.com/new-ai-system-enhances-blood-cell-analysis-for-improved-leukemia-diagnosis/ [33] AI Fools Blood Experts With Fake Cells, Then Outdiagnoses Them - ScienceBlog.com - https://scienceblog.com/ai-fools-blood-experts-with-fake-cells-then-outdiagnoses-them/ [34] Deep generative classification of blood cell morphology | Nature Machine Intelligence - https://www.nature.com/articles/s42256-025-01122-7 [35] China’s DeepRare AI Aims to Speed Rare Disease Diagnosis - https://www.sixthtone.com/news/1018223/china%E2%80%99s-deeprare-ai-aims-to-speed-rare-disease-diagnosis [36] AI Tool Sets New Standard in Diagnosing Rare Diseases - https://lifespan.io/news/ai-tool-sets-new-standard-in-diagnosing-rare-diseases/ [37] An agentic system for rare disease diagnosis with traceable reasoning | Nature - https://www.nature.com/articles/s41586-025-10097-9 [38] Deep Conformal Supervision: Leveraging Intermediate Features for Robust Uncertainty Quantification - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12092326/ [39] Estimating diagnostic uncertainty in artificial intelligence assisted pathology using conformal prediction | Nature Communications - https://www.nature.com/articles/s41467-022-34945-8 [40] The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence - PubMed - https://pubmed.ncbi.nlm.nih.gov/40954311/ [41] Developing a reporting guideline for artificial intelligence-centred diagnostic test accuracy studies: the STARD-AI protocol - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC8240576/ [42] Reporting of Artificial Intelligence Diagnostic Accuracy Studies in Pathology Abstracts: Compliance with STARD for Abstracts Guidelines - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC9576989/ [43] Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration in 2024: Regulatory Characteristics, Predicate Lineage, and Transparency Reporting - https://pubmed.ncbi.nlm.nih.gov/PMC12730494 [44] LLMs factor in unrelated information when recommending medical treatments | MIT News | Massachusetts Institute of Technology - https://news.mit.edu/2025/llms-factor-unrelated-information-when-recommending-medical-treatments-0623 [45] Current applications and challenges in large language models for patient care: a systematic review | Communications Medicine - https://www.nature.com/articles/s43856-024-00717-2 [46] Visual-Textual Integration in LLMs for Medical Diagnosis: A Quantitative Analysis - https://www.medrxiv.org/content/10.1101/2024.08.31.24312878 [47] JMIR Formative Research - Explainable AI-Driven Analysis of Radiology Reports Using Text and Image Data: Experimental Study - https://formative.jmir.org/2025/1/e77482 [48] Evaluating the Multimodal Capabilities of Generative AI in Complex Clinical Diagnostics - https://www.medrxiv.org/content/10.1101/2023.11.01.23297938 [49] Let’s ask the patient: disease prediction based on patients’ symptom descriptions in free text - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12456275/ [50] Large Language Model Symptom Identification from Clinical Text: A Multi-Center Study - https://www.medrxiv.org/content/10.1101/2024.12.16.24319044 [51] ISIC | International Skin Imaging Collaboration - https://www.isic-archive.com/ [52] Analysis of the ISIC image datasets: Usage, benchmarks and recommendations - https://e-space.mmu.ac.uk/631025/7/1-s2.0-S1361841521003509-main.pdf [53] ISIC Challenge - https://challenge.isic-archive.com/data/ [54] 15 AI in Healthcare Case Studies [2026] - DigitalDefynd Education - https://digitaldefynd.com/IQ/ai-in-healthcare-case-studies/ [55] Enhancing diagnostic deep learning via self-supervised pretraining on large-scale, unlabeled non-medical images - https://pubmed.ncbi.nlm.nih.gov/PMC10850044 [56] CheXmask: a large-scale dataset of anatomical segmentation masks for multi-center chest x-ray images - https://pubmed.ncbi.nlm.nih.gov/PMC11101488 [57] Breaking Medical Data Sharing Boundaries by Employing Artificial Radiographs - https://www.biorxiv.org/content/10.1101/841619 [58] Interpreting chest X-rays via CNNs that exploit disease dependencies and uncertainty labels - https://www.medrxiv.org/content/10.1101/19013342 [59] Learning to recognize Abnormalities in Chest X-Rays with Location-Aware Dense Networks - https://arxiv.org/abs/1803.04565 [60] Weighing Features of Lung and Heart Regions for Thoracic Disease Classification - https://arxiv.org/abs/2105.12430 [61] Automated thorax disease diagnosis using multi-branch residual attention network - https://pubmed.ncbi.nlm.nih.gov/PMC11126636 [62] MedSegBench: A Comprehensive Benchmark for Medical Image Segmentation in Diverse Data Modalities - https://www.medrxiv.org/content/10.1101/2024.08.26.24312619 [63] A Large-Scale Vision-Language Dataset Derived from Open Scientific Literature to Advance Biomedical Generalist AI - https://arxiv.org/abs/2503.22727 [64] AI and Automation in Healthcare – 2026 Health IT Predictions | Healthcare IT Today - https://www.healthcareittoday.com/2025/12/23/ai-and-automation-in-healthcare-2026-health-it-predictions/ [65] Clinical ModernBERT: An efficient and long context encoder for biomedical text - https://arxiv.org/html/2504.03964v1 [66] Announcing BioClinical ModernBERT: a new SOTA encoder model for Medical NLP - LightOn - https://www.lighton.ai/lighton-blogs/announcing-bioclinical-modernbert-a-new-sota-encoder-model-for-medical-nlp [67] [2506.10896] BioClinical ModernBERT: A State-of-the-Art Long-Context Encoder for Biomedical and Clinical NLP - https://arxiv.org/abs/2506.10896 [68] BioGPT:GenerativePre-trainedTransformerfor # BiomedicalTextGenerationandMining - https://arxiv.org/pdf/2210.10341 [69] BioGPT: Generative Pre-trained Transformer for Biomedical Text Generation and Mining | by Eleventh Hour Enthusiast | Medium - https://medium.com/@EleventhHourEnthusiast/paper-review-biogpt-631f33f5c4ea [70] BioMedLM: A 2.7B Parameter Language Model Trained On Biomedical Text - https://arxiv.org/html/2403.18421v1 [71] Benchmarking large language models for biomedical natural language processing applications and recommendations | Nature Communications - https://www.nature.com/articles/s41467-025-56989-2 [72] Why Healthcare AI Fails: The Role of Workflow Adoption - Logicon - https://logicon.tech/healthcare-ai-workflow-adoption-fails/ [73] Translating AI to the Bedside with Physician Buy-In: Recommendations from a Meta-Analysis and Systematic Review of the Literature - https://pubmed.ncbi.nlm.nih.gov/PMC12729787 [74] Physician Perspectives on Ambient AI Scribes Physician Perspectives on Ambient AI Scribes Physician Perspectives on Ambient AI Scribes - https://pubmed.ncbi.nlm.nih.gov/PMC11933996 [75] Barriers to and Facilitators of Artificial Intelligence Adoption in Health Care: Scoping Review - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC11393514/ [76] MIMIC-IV, a freely accessible electronic health record dataset - https://pubmed.ncbi.nlm.nih.gov/PMC9810617 [77] Multimodal Risk Prediction with Physiological Signals, Medical Images and Clinical Notes - https://www.medrxiv.org/content/10.1101/2023.05.18.23290207 [78] Preparing Hospitals for Large-Scale AI Deployments in 2026 - https://www.johnsnowlabs.com/preparing-hospitals-for-large-scale-ai-deployments-in-2026/ [79] The landscape of AI implementation in US hospitals | Nature Health - https://www.nature.com/articles/s44360-025-00016-7 [80] AI-Powered early warning systems for clinical deterioration significantly improve patient outcomes: a meta-analysis | BMC Medical Informatics and Decision Making | Springer Nature Link - https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-025-03048-x [81] External validation of AI-based scoring systems in the ICU: a systematic review and meta-analysis | BMC Medical Informatics and Decision Making | Springer Nature Link - https://bmcmedinformdecismak.biomedcentral.com/articles/10.1186/s12911-024-02830-7 [82] DOLCE: Determining the Impact of Optellum’s Lung Cancer Prediction (LCP) Artificial Intelligence Solution on Service Utilisation, Health Economics and Patient Outcomes - https://clinicaltrials.gov/study/NCT05389774 [83] Artificial Intelligence-Aided Screening for Patients With Diabetic Retinopathy and Age-related Macular Degeneration in Family Medicine and Geriatric Medicine Outpatient Clinics: A Randomized Controlled Clinical Trial - https://clinicaltrials.gov/study/NCT07069647 [84] GatorTron: A Large Clinical Language Model to Unlock Patient Information from Unstructured Electronic Health Records | by Eleventh Hour Enthusiast | Medium - https://medium.com/@EleventhHourEnthusiast/paper-review-gatortron-b38578b077a2 [85] GatorTron-Base: Clinical NLP Transformer - https://www.emergentmind.com/topics/gatortron-base [86] GatorTron: A Large Language Model for Clinical Natural Language Processing | medRxiv - https://www.medrxiv.org/content/10.1101/2022.02.27.22271257v2.full [87] A large language model for electronic health records | npj Digital Medicine - https://www.nature.com/articles/s41746-022-00742-2 [88] Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare | npj Digital Medicine - https://www.nature.com/articles/s41746-025-01722-y [89] The Good, the Bad: Behind the Scenes Economic Impact of AI in Healthcare | The AI Journal - https://aijourn.com/economicimpacthealthcare/ [90] AI’s Uncertain Cost Effects in Health Care | American Enterprise Institute - AEI - https://www.aei.org/health-care/ais-uncertain-cost-effects-in-health-care/ [91] The Impact of Artificial Intelligence on the Health Economy, Workforce Productivity, and Administrative Efficiency: A Systematic Review - https://www.medrxiv.org/content/10.1101/2025.10.05.25337345 [92] How AI Is Cutting Healthcare Costs in 2025 — With Real Results - https://graphlogic.ai/blog/ai-chatbots/ai-use-cases-by-industry/ai-reduces-costs-healthcare/ [93] EMPOWERing Patients with Type 2 Diabetes Mellitus (T2DM) in Primary Care Through App-based Motivational Interviewing PLUS Artificial Intelligence Powered Diabetes Management (EMPOWER-PLUS): Randomised Controlled Trial - https://clinicaltrials.gov/study/NCT06214520 [94] Randomized Controlled Trials Evaluating AI in Clinical Practice: A Scoping Evaluation | medRxiv - https://www.medrxiv.org/content/10.1101/2023.09.12.23295381v1.full [95] Assessing the Impact of an Artificial Intelligence-Based Model for Intracranial Aneurysm Detection in CT Angiography on Patient Diagnosis and Outcomes (IDEAL Study)—a protocol for a multicenter, double-blinded randomized controlled trial | Trials | Springer Nature Link - https://link.springer.com/article/10.1186/s13063-024-08184-9 [96] Towards Expert-Level Medical Question Answering with Large Language Models - https://arxiv.org/abs/2305.09617 [97] TowardsExpert-Level # MedicalQuestionAnswering # withLargeLanguageModels - https://arxiv.org/pdf/2305.09617 [98] Toward expert-level medical question answering with large language models - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC11922739/ [99] Unlocking Healthcare AI Potential with Multimodal Medical Datasets | Shaip - https://www.shaip.com/blog/multimodal-medical-datasets-for-ai-research/ [100] 18 Open Healthcare Datasets – 2025 Update - https://opendatascience.com/18-open-healthcare-datasets-2025-update/ [101] 22 Free and Open Medical Datasets for AI Development in 2025 - https://www.shaip.com/blog/healthcare-datasets-for-machine-learning-projects/ [102] CardioEHR: A longitudinal electronic health record dataset of cardiovascular patients from central China | Scientific Data - https://www.nature.com/articles/s41597-026-06855-7 [103] Frontiers | Multi-modal AI in precision medicine: integrating genomics, imaging, and EHR data for clinical insights - https://www.frontiersin.org/journals/artificial-intelligence/articles/10.3389/frai.2025.1743921/full [104] AI in Hospitals: 2025 Adoption Trends & Statistics | IntuitionLabs - https://intuitionlabs.ai/articles/ai-adoption-us-hospitals-2025 [105] AMA Augmented Intelligence Research | AMA - https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf [106] Physician Assistant Trends 2026: AI Adoption, Title Changes, and Workforce Growth - https://hitconsultant.net/2025/12/30/physician-assistant-trends-2025-ai-adoption-title-changes-and-workforce-growth/ [107] Acceptance of clinical artificial intelligence among physicians and medical students: A systematic review with cross-sectional survey - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC9472134/ [108] Adoption of artificial intelligence in healthcare: survey of health system priorities, successes, and challenges - PMC - https://pmc.ncbi.nlm.nih.gov/articles/PMC12202002/ [109] Evaluating generalizability of oncology trial results to real-world patients using machine learning-based trial emulations | Nature Medicine - https://www.nature.com/articles/s41591-024-03352-5 [110] Sepsis-detection AI has the potential to prevent thousands of deaths | Hub - https://hub.jhu.edu/2022/07/21/artificial-intelligence-sepsis-detection/ [111] Evaluating a sepsis prediction machine learning algorithm in the emergency department and intensive care unit: a before and after comparative study - https://www.biorxiv.org/content/10.1101/224014 [112] PhysioNet - https://physionet.org [113] Artificial intelligence integrates multi-omics data for precision stratification and drug resistance prediction in breast cancer - https://pubmed.ncbi.nlm.nih.gov/PMC12463597 [114] Using AI (Artificial Intelligence) to Detect Breast Cancer - https://www.breastcancer.org/screening-testing/artificial-intelligence [115] Systematic review and meta-analysis of artificial intelligence for image-based lung cancer classification and prognostic evaluation | npj Precision Oncology - https://www.nature.com/articles/s41698-025-01095-1