From Benchmarks to Bedside: Improving Safety and Reliability of LLM-Based Clinical Decision Support with Evidence-based Medical Reasoning
AI is increasingly permeating many areas of our lives, including medicine- e.g. through wearables. Leading companies are also working to integrate LLMs into hospitals. Yet recent studies suggest that current models are not yet ready for this use and may even pose risks to patients. For this reason, among others, AI applications in high-risk domains such as medicine must therefore be transparent, controllable and open to human intervention to reduce errors. This requirement is also reflected in the EU AI Act. To ensure the safe, transparent, and reliable use of this rapidly evolving technology in the future, further research is therefore essential.
The project integrates different approaches for incorporating medical “knowledge” in current LLMs and evaluates their clinical impact. It aims to strengthen the models’ clinical reasoning, enable robust, data-driven evaluation and support the development of a medically grounded clinical decision-support system. By strengthening the robustness and medical accuracy of LLMs and ensure their adherence to clinical guidelines, the project contributes to ensure patient safety in an increasingly digital healthcare of the coming decades.
Further information: https://radiologie.mri.tum.de/de