Optimization of Multimodal Language Models in Radiology through Weighted Integration of Imaging and Contextual Information: The Example of Neuro-Oncology
Radiological decision-making relies on the integration of visual imaging data and clinical information. Multimodal large language models can process both modalities, but currently lack mechanisms for targeted weighting. The aim of this project is to systematically investigate, using neuro-oncology as an example, how imaging and contextual information can be optimally combined to improve diagnostic accuracy, reduce misinterpretations, and enable structured, visually supported presentation of clinical cases.
The project is structured into three work packages: (1) analysis of how synthetic clinical information with varying specificity influences visual tumor detection and development of weighting strategies; (2) validation of these strategies on real neuro-oncological cases; and (3) implementation of the optimized approaches into a system for automated generation of structured, visually enriched case summaries for tumor boards.
The goal is to derive robust concepts for multimodal information integration and to provide standardized weighting schemes and prompt templates for clinical use.