Artificial intelligence is increasingly used to assist doctors in diagnosing diseases by analyzing complex medical data, such as images and patient records. However, many AI systems focus on simply getting the right answer most of the time, which can be misleading in medicine because some diseases are very rare. A newly published research paper introduces a smarter way to train these AI models that better distinguishes between sick and healthy patients, improving their usefulness in real-world clinical settings.
Key Takeaways
- Traditional AI training methods optimize for accuracy, which can be misleading when diseases are rare and data is imbalanced.
- The new approach optimizes models based on a metric called AUROC, which better reflects how well the model ranks true positive cases above negatives without bias from class imbalance.
- The researchers developed a novel prompt optimization technique called Ranking-PE that improves how multimodal large language models (MLLMs) learn from clinical data, combining text and images.
- Testing on real clinical datasets showed that Ranking-PE significantly outperforms accuracy-based methods, improving diagnostic ranking scores by up to 16 percentage points.
In clinical AI, data imbalance is a major challenge. For example, if only 5% of patients have a certain disease, a model that always predicts “no disease” can still achieve 95% accuracy but is clinically useless. To address this, the researchers focused on optimizing the AI’s ability to rank patients correctly—placing those with disease higher than those without—using a metric called the Area Under the Receiver Operating Characteristic curve (AUROC). Unlike accuracy, AUROC evaluates how well the model separates positive and negative cases across all thresholds, making it more reliable for imbalanced data.
The team worked with multimodal large language models (MLLMs), which are AI systems designed to process and understand multiple types of data simultaneously, such as medical images and clinical notes. These models use “prompts,” or input instructions, to guide their reasoning. Optimizing these prompts is crucial for improving model performance. Previous methods, like GEPA, selected prompts based on simple accuracy scores — checking if the model was correct on each case individually.
The new method, called Pair-level Pareto prompt evolution (Ranking-PE), takes a different approach. Instead of looking at individual correctness, it compares pairs of patient cases—one positive (diseased) and one negative (healthy)—and checks if the model ranks the positive case higher. This pairwise comparison aligns directly with the AUROC metric. By replacing the traditional accuracy-based feedback with this ranking-aware feedback at multiple stages of prompt optimization, Ranking-PE guides the model to better distinguish between sick and healthy patients.
Importantly, this improved prompt optimization does not require additional calls to the AI model or complex surrogate loss functions, making it efficient. The researchers tested Ranking-PE on three diseases using the MIMIC dataset, a large collection of real-world clinical data. They found that while accuracy-based prompt optimization sometimes degraded the model’s ranking ability, Ranking-PE consistently improved it, achieving notable gains in AUROC—up to +5.8 percentage points on one model and +16.2 points on another.
The study also highlights that having a strong medical-grade visual component in the AI system is essential. The visual encoder, which processes medical images, must be well-trained on medical data to provide reliable information—something prompt optimization alone cannot fix.
These findings suggest that focusing on ranking metrics like AUROC and using pairwise prompt optimization could lead to more clinically useful AI diagnostic tools. By better handling class imbalance and multimodal data, such methods bring us closer to AI systems that can support doctors in making more accurate and reliable diagnoses. Future work may explore applying this approach to other diseases and integrating it into clinical workflows to evaluate its impact in real healthcare settings.
Based on research published on arXiv by Tian Xia, Minghao Liu, Yiqing Liang et al..
