MERGE-Med: Multi-Expert Retrieval-Grounded Engine for Medical Report Summarization Using Hierarchical RAG and Contrastive Learning
DOI:
https://doi.org/10.68337/cpsm.v1.i1.2026-008Keywords:
Medical report summarization, mixture-of-experts, retrieval-augmented generation, contrastive learning, hallucination mitigation, clinical NLP, LLaMA-3, electronic health recordsAbstract
Electronic health records (EHRs) have created large volumes of clinical data that require automated and accurate summarization that patients can understand. Current large language models (LLMs) are prone to hallucination and show limited task specialization and inadequate clinical grounding. MERGE-Med is a hybrid multi-expert architecture that consists of (i) a mixture-of-experts (MoE) layer with three domain-specialized LLaMA-3 8B models for diagnostic reasoning, treatment protocol generation, and patient-friendly communication; (ii) a hierarchical retrieval-augmented generation (RAG) pipeline that combines FAISS dense retrieval, BM25 sparse retrieval, cross-encoder reranking, and Unified Medical Language System (UMLS) knowledge-graph traversal over 35 million PubMed abstracts, 50,000 clinical guidelines, and 15,000 pharmaceutical records; and (iii) a contrastive fine-tuning framework validated by an ensembled DeBERTa-v3-large fact-checking layer that achieves more than 95% consistency. MERGE-Med achieved accuracies of 92.45%, 82.34%, and 98.76% on MedQA, MedMCQA, and Clinical Knowledge, respectively, improvements of 5.18, 6.44, and 1.78 percentage points (pp) over the baseline LLaMA-3 8B. The hallucination rate decreased by 67% (from 18.3% to 6.1%), ROUGE-L improved by 22.3%, and in an evaluation by board-certified physicians of 100 de-identified reports, 92% of summaries were judged appropriate for patients, with an inference latency of 2.3 s.
References
[1] Ilona Kickbusch, Jürgen M. Pelikan, Franklin Apfel, and Agis D. Tsouros, Eds., Health Literacy: The Solid Facts. Copenhagen, Denmark: World Health Organization Regional Office for Europe, 2013, ISBN: 978-92-890-0015-4. [Online]. Available: https://iris.who.int/handle/10665/326432
[2] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. 2019 Conf. North Amer. Chapter Assoc. Comput. Linguist.: Human Lang. Technol. (NAACL-HLT), 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.
[3] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei, "Language models are few-shot learners," in Adv. Neural Inf. Process. Syst., vol. 33, 2020, pp. 1877-1901, doi: 10.48550/arXiv.2005.14165.
[4] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample, "LLaMA: Open and efficient foundation language models," arXiv:2302.13971, 2023, doi: 10.48550/arXiv.2302.13971.
[5] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela, "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Adv. Neural Inf. Process. Syst., vol. 33, 2020, pp. 9459-9474, doi: 10.48550/arXiv.2005.11401.
[6] Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu, "BioGPT: Generative pre-trained transformer for biomedical text generation and mining," Brief. Bioinform., vol. 23, no. 6, Art. no. bbac409, 2022, doi: 10.1093/bib/bbac409.
[7] Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang, "BioBERT: A pre-trained biomedical language representation model for biomedical text mining," Bioinformatics, vol. 36, no. 4, pp. 1234-1240, 2020, doi: 10.1093/bioinformatics/btz682.
[8] Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R. Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H. Chen, Nigam H. Shah, Sami Lachgar, Philip Andrew Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Agüera y Arcas, Nenad Tomašev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara Mahdavi, Joelle K. Barral, Dale R. Webster, Greg S. Corrado, Yossi Matias, Shekoofeh Azizi, Alan Karthikesalingam, and Vivek Natarajan, "Toward expert-level medical question answering with large language models," Nat. Med., vol. 31, no. 3, pp. 943-950, 2025, doi: 10.1038/s41591-024-03423-7.
[9] Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, Eduardo Pontes Reis, Anna Seehofnerová, Nidhi Rohatgi, Poonam Hosamani, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S. Chaudhari, "Adapted large language models can outperform medical experts in clinical text summarization," Nat. Med., vol. 30, no. 4, pp. 1134-1142, 2024, doi: 10.1038/s41591-024-02855-5.
[10] Qiao Jin, Won Kim, Qingyu Chen, Donald C. Comeau, Lana Yeganova, W. John Wilbur, and Zhiyong Lu, "MedCPT: Contrastive pre-trained transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval," Bioinformatics, vol. 39, no. 11, Art. no. btad651, 2023, doi: 10.1093/bioinformatics/btad651.
[11] Lameck Mbangula Amugongo, Pietro Mascheroni, Steven Brooks, Stefan Doering, and Jan Seidel, "Retrieval augmented generation for large language models in healthcare: A systematic review," PLOS Digit. Health, vol. 4, no. 6, Art. no. e0000877, 2025, doi: 10.1371/journal.pdig.0000877.
[12] William Fedus, Barret Zoph, and Noam Shazeer, "Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity," J. Mach. Learn. Res., vol. 23, no. 120, pp. 1-39, 2022. [Online]. Available: https://jmlr.org/papers/v23/21-0998.html
[13] Tianyu Gao, Xingcheng Yao, and Danqi Chen, "SimCSE: Simple contrastive learning of sentence embeddings," in Proc. 2021 Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2021, pp. 6894-6910, doi: 10.18653/v1/2021.emnlp-main.552.
[14] Ray Smith, "An overview of the Tesseract OCR engine," in Proc. 9th Int. Conf. Document Anal. Recognit. (ICDAR), Curitiba, Brazil, 2007, pp. 629-633, doi: 10.1109/ICDAR.2007.4376991.
Downloads
Published
Versions
- 2026-10-01 (2)
- 2026-09-30 (1)
Data Availability Statement
This study used the publicly available MedQA, MedMCQA, and MMLU benchmark datasets; the data from the custom clinical evaluation set are available from the corresponding author upon reasonable request.Conference Proceedings Volume
Section
Categories
License
Copyright (c) 2026 The Authors

This work is licensed under a Creative Commons Attribution 4.0 International License.
Articles published in Conference Proceedings in Science and Management are made available under the Creative Commons Attribution 4.0 International License (CC BY 4.0). Authors retain copyright in their work and grant the journal first publication rights. Readers are free to share and adapt the published material for any purpose, including commercial use, provided appropriate credit is given to the author, a link to the license is provided, and any changes are indicated.
