Research Paper:
Generating Personalized MRI Reports from Key Phrases with Multiple LLM Agents and RAG Technique
Yulia Shichkina
and Aleksandr P. Stepanov

Department of Computer Science and Engineering, Saint Petersburg Electrotechnical University (LETI)
Litera F, Professor Popov 5, Saint Petersburg 197022, Russia
Corresponding author
In the paper, the novel method for generating personalized magnetic resonance imaging (MRI) reports based on key phrases is introduced. In order to demonstrate the method, an application with multiple large language model (LLM) agents was written. For establishing connection between different agents, graph-based software architecture was used. The graph consists of the following nodes. (1) The data retriever node. In this node relevant reports along with the metadata are obtained from the vector database. (2) The findings generation node. This node includes an implementation of two different variations of an MRI report’s findings section generation method. The first is based on merging already existed findings parts of radiology reports that were retrieved by key phrases into the new one. The second is based on reconstructing findings section so that it contains specific paragraphs retrieved by key phrases from the vector store. (3) The impression generation node. This node implements summarization of a findings section into an impression. The performance of a findings part generation was evaluated with BLEU, ROUGE and BERTScore metrics. Experiments have shown promising results in generation of findings parts: for the first approach mean ROUGE-L score was about 0.5, BERTScore was around 0.8, for the second approach mean ROUGE-L score was approximately 0.6, BERTScore was roughly 0.9. The first approach was around three times slower than the second. During the experiments, the dataset of 908 reports gathered from seven radiologists was used. Qualitative analysis was performed with five-point Likert scale questionary and statistically analyzed by means of Mann–Whitney U test.
Findings section generation
- [1] L. Luo, “Combining knowledge graph and artificial intelligence to conduct financial report quality detection research,” J. Adv. Comput. Intell. Intell. Inform., Vol.29, No.4, pp. 787-795, 2025. https://doi.org/10.20965/jaciii.2025.p0787
- [2] M. Fareed, M. Fatima, J. Uddin, A. Ahmed, and M. A. Sattar, “A systematic review of ethical considerations of large language models in healthcare and medicine,” Frontiers in Digital Health, Vol.7, Article No.1653631, 2025. https://doi.org/10.3389/fdgth.2025.1653631
- [3] T. Nakaura et al., “The impact of large language models on radiology: A guide for radiologists on the latest innovations in AI,” Japanese J. of Radiology, Vol.42, No.7, pp. 685-696, 2024. https://doi.org/10.1007/s11604-024-01552-0
- [4] M. P. Hartung, I. C. Bickle, F. Gaillard, and J. P. Kanne, “How to create a great radiology report,” RadioGraphics, Vol.40, No.6, pp. 1658-1670, 2020. https://doi.org/10.1148/rg.2020200020
- [5] Z. Sun et al., “Evaluating GPT-4 on impressions generation in radiology reports,” Radiology, Vol.307, No.5, Article No.e231259, 2023. https://doi.org/10.1148/radiol.231259
- [6] A. Serapio et al., “An open-source fine-tuned large language model for radiological impression generation: A multi-reader performance study,” BMC Medical Imaging, Vol.24, Article No.254, 2024. https://doi.org/10.1186/s12880-024-01435-w
- [7] L. Zhang et al., “Constructing a large language model to generate impressions from findings in radiology reports,” Radiology, Vol.312, No.3, Article No.e240885, 2024. https://doi.org/10.1148/radiol.240885
- [8] Ș.-V. Voinea et al., “GPT-driven radiology report generation with fine-tuned Llama 3,” Bioengineering, Vol.11, No.10, Article No.1043, 2024. https://doi.org/10.3390/bioengineering11101043
- [9] F. Zeng, Z. Lyu, Q. Li, and X. Li, “Enhancing LLMs for impression generation in radiology reports through a multi-agent system,” arXiv:2412.06828, 2024. https://doi.org/10.48550/arXiv.2412.06828
- [10] P. Lewis et al., “Retrieval-augmented generation for knowledge-intensive NLP tasks,” Proc. of the 34th Int. Conf. on Neural Information Processing Systems, pp. 9459-9474, 2020.
- [11] M. L. Saini, V. M. Shrimal, A. Garg, D. C. Sati, and S. P. K. Mygapula, “Medical diagnosis with RAG-LLMs: A hybrid approach for AI-driven healthcare,” IEEE 4th Int. Conf. for Advancement in Technology, 2025. https://doi.org/10.1109/ICONAT66879.2025.11362545
- [12] H. M. T. Alam, D. Srivastav, M. A. Kadir, and D. Sonntag, “Towards interpretable radiology report generation via concept bottlenecks using a multi-agentic RAG,” Proc. of 47th European Conf. on Information Retrieval, Part 3, pp. 201-209, 2025. https://doi.org/10.1007/978-3-031-88714-7_18
- [13] Ç. U. Öğdü, K. Arslanoğlu, and M. Karaköse, “An adaptive multi-agent LLM-based clinical decision support system integrating biomedical RAG and web intelligence,” IEEE Access, Vol.13, pp. 167390-167404, 2025. https://doi.org/10.1109/ACCESS.2025.3613340.
- [14] J. Burrows, “‘Delta’: A measure of stylistic difference and a guide to likely authorship,” Literary and Linguistic Computing, Vol.17, No.3, pp. 267-287, 2002. https://doi.org/10.1093/llc/17.3.267
- [15] FRIDA. https://huggingface.co/ai-forever/FRIDA [Accessed November 17, 2025]
- [16] Chroma: Open-source search infrastructure for AI. https://www.trychroma.com/ [Accessed November 17, 2025]
- [17] A. Yang et al., “Qwen3 technical report,” arXiv:2505.09388, 2025. https://doi.org/10.48550/arXiv.2505.09388
- [18] A. Grattafiori et al., “The Llama 3 herd of models,” arXiv:2407.21783, 2024. https://doi.org/10.48550/arXiv.2407.21783
- [19] C.-Y. Lin, “ROUGE: A package for automatic evaluation of summaries,” Text Summarization Branches Out (Proc. of the ACL-04 Workshop), pp. 74-81, 2004.
- [20] K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: A method for automatic evaluation of machine translation,” Proc. of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311-318, 2002. https://doi.org/10.3115/1073083.1073135
- [21] Y. Wu et al., “Google’s neural machine translation system: Bridging the gap between human and machine translation,” arXiv:1609.08144, 2016. https://doi.org/10.48550/arXiv.1609.08144
- [22] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi, “BERTScore: Evaluating text generation with BERT,” arXiv:1904.09675, 2019. https://doi.org/10.48550/arXiv.1904.09675
- [23] J. A. Omiye, H. Gui, S. J. Rezaei, J. Zou, and R. Daneshjou, “Large language models in medicine: The potentials and pitfalls: A narrative review,” Annals of Internal Medicine, Vol.177, No.2, pp. 210-220, 2024. https://doi.org/10.7326/M23-2772
- [24] S. A. Kostrov and M. P. Potapov, “Large language models in medicine: Current ethical challenges,” Medical Ethics, Vol.2025, No.2, pp. 22-31, 2025. https://doi.org/10.24075/medet.2025.008
- [25] M. Cascella et al., “The breakthrough of large language models release for medical applications: 1-year timeline and perspectives,” J. of Medical Systems, Vol.48, Article No.22, 2024. https://doi.org/10.1007/s10916-024-02045-3
- [26] S. M. C. A. Negri, “Robot as legal person: Electronic personhood in robotics and artificial intelligence,” Frontiers in Robotics and AI, Vol.8, Article No.789327, 2021. https://doi.org/10.3389/frobt.2021.789327
- [27] S. Schmidgall et al., “Evaluation and mitigation of cognitive biases in medical language models,” npj Digital Medicine, Vol.7, Article No.295, 2024. https://doi.org/10.1038/s41746-024-01283-6
- [28] N. Yadav et al., “Data privacy in healthcare: In the era of artificial intelligence,” Indian Dermatology Online J., Vol.14, No.6, pp. 788-792, 2023. https://doi.org/10.4103/idoj.idoj_543_23
- [29] I. C. Wiest et al., “Anonymizing medical documents with local, privacy preserving large language models: The LLM-anonymizer,” medRxiv, 2024. https://doi.org/10.1101/2024.06.11.24308355
This article is published under a Creative Commons Attribution-NoDerivatives 4.0 Internationa License.