MERGE-Med: Multi-Expert Retrieval-Grounded Engine for Medical Report Summarization Using Hierarchical RAG and Contrastive Learning

Authors

  • Sai Tejaswi D. School of Engineering, Anurag University, Hyderabad, India Author
  • Raja Sekhar Reddy P. School of Engineering, Anurag University, Hyderabad, India Author
  • Shailaja K. School of Engineering, Anurag University, Hyderabad, India Author
  • Vishnu Murthy G. School of Engineering, Anurag University, Hyderabad, India Author

DOI:

https://doi.org/10.68337/cpsm.v1.i1.2026-008

Keywords:

Medical report summarization, mixture-of-experts, retrieval-augmented generation, contrastive learning, hallucination mitigation, clinical NLP, LLaMA-3, electronic health records

Abstract

Electronic health records (EHRs) have created large volumes of clinical data that require automated and accurate summarization that patients can understand. Current large language models (LLMs) are prone to hallucination and show limited task specialization and inadequate clinical grounding. MERGE-Med is a hybrid multi-expert architecture that consists of (i) a mixture-of-experts (MoE) layer with three domain-specialized LLaMA-3 8B models for diagnostic reasoning, treatment protocol generation, and patient-friendly communication; (ii) a hierarchical retrieval-augmented generation (RAG) pipeline that combines FAISS dense retrieval, BM25 sparse retrieval, cross-encoder reranking, and Unified Medical Language System (UMLS) knowledge-graph traversal over 35 million PubMed abstracts, 50,000 clinical guidelines, and 15,000 pharmaceutical records; and (iii) a contrastive fine-tuning framework validated by an ensembled DeBERTa-v3-large fact-checking layer that achieves more than 95% consistency. MERGE-Med achieved accuracies of 92.45%, 82.34%, and 98.76% on MedQA, MedMCQA, and Clinical Knowledge, respectively, improvements of 5.18, 6.44, and 1.78 percentage points (pp) over the baseline LLaMA-3 8B. The hallucination rate decreased by 67% (from 18.3% to 6.1%), ROUGE-L improved by 22.3%, and in an evaluation by board-certified physicians of 100 de-identified reports, 92% of summaries were judged appropriate for patients, with an inference latency of 2.3 s.

References

[1] Ilona Kickbusch, Jürgen M. Pelikan, Franklin Apfel, and Agis D. Tsouros, Eds., Health Literacy: The Solid Facts. Copenhagen, Denmark: World Health Organization Regional Office for Europe, 2013, ISBN: 978-92-890-0015-4. [Online]. Available: https://iris.who.int/handle/10665/326432

[2] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova, "BERT: Pre-training of deep bidirectional transformers for language understanding," in Proc. 2019 Conf. North Amer. Chapter Assoc. Comput. Linguist.: Human Lang. Technol. (NAACL-HLT), 2019, pp. 4171-4186, doi: 10.18653/v1/N19-1423.

[3] Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei, "Language models are few-shot learners," in Adv. Neural Inf. Process. Syst., vol. 33, 2020, pp. 1877-1901, doi: 10.48550/arXiv.2005.14165.

[4] Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample, "LLaMA: Open and efficient foundation language models," arXiv:2302.13971, 2023, doi: 10.48550/arXiv.2302.13971.

[5] Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela, "Retrieval-augmented generation for knowledge-intensive NLP tasks," in Adv. Neural Inf. Process. Syst., vol. 33, 2020, pp. 9459-9474, doi: 10.48550/arXiv.2005.11401.

[6] Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu, "BioGPT: Generative pre-trained transformer for biomedical text generation and mining," Brief. Bioinform., vol. 23, no. 6, Art. no. bbac409, 2022, doi: 10.1093/bib/bbac409.

[7] Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang, "BioBERT: A pre-trained biomedical language representation model for biomedical text mining," Bioinformatics, vol. 36, no. 4, pp. 1234-1240, 2020, doi: 10.1093/bioinformatics/btz682.

[8] Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R. Pfohl, Heather Cole-Lewis, Darlene Neal, Qazi Mamunur Rashid, Mike Schaekermann, Amy Wang, Dev Dash, Jonathan H. Chen, Nigam H. Shah, Sami Lachgar, Philip Andrew Mansfield, Sushant Prakash, Bradley Green, Ewa Dominowska, Blaise Agüera y Arcas, Nenad Tomašev, Yun Liu, Renee Wong, Christopher Semturs, S. Sara Mahdavi, Joelle K. Barral, Dale R. Webster, Greg S. Corrado, Yossi Matias, Shekoofeh Azizi, Alan Karthikesalingam, and Vivek Natarajan, "Toward expert-level medical question answering with large language models," Nat. Med., vol. 31, no. 3, pp. 943-950, 2025, doi: 10.1038/s41591-024-03423-7.

[9] Dave Van Veen, Cara Van Uden, Louis Blankemeier, Jean-Benoit Delbrouck, Asad Aali, Christian Bluethgen, Anuj Pareek, Malgorzata Polacin, Eduardo Pontes Reis, Anna Seehofnerová, Nidhi Rohatgi, Poonam Hosamani, William Collins, Neera Ahuja, Curtis P. Langlotz, Jason Hom, Sergios Gatidis, John Pauly, and Akshay S. Chaudhari, "Adapted large language models can outperform medical experts in clinical text summarization," Nat. Med., vol. 30, no. 4, pp. 1134-1142, 2024, doi: 10.1038/s41591-024-02855-5.

[10] Qiao Jin, Won Kim, Qingyu Chen, Donald C. Comeau, Lana Yeganova, W. John Wilbur, and Zhiyong Lu, "MedCPT: Contrastive pre-trained transformers with large-scale PubMed search logs for zero-shot biomedical information retrieval," Bioinformatics, vol. 39, no. 11, Art. no. btad651, 2023, doi: 10.1093/bioinformatics/btad651.

[11] Lameck Mbangula Amugongo, Pietro Mascheroni, Steven Brooks, Stefan Doering, and Jan Seidel, "Retrieval augmented generation for large language models in healthcare: A systematic review," PLOS Digit. Health, vol. 4, no. 6, Art. no. e0000877, 2025, doi: 10.1371/journal.pdig.0000877.

[12] William Fedus, Barret Zoph, and Noam Shazeer, "Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity," J. Mach. Learn. Res., vol. 23, no. 120, pp. 1-39, 2022. [Online]. Available: https://jmlr.org/papers/v23/21-0998.html

[13] Tianyu Gao, Xingcheng Yao, and Danqi Chen, "SimCSE: Simple contrastive learning of sentence embeddings," in Proc. 2021 Conf. Empirical Methods Natural Lang. Process. (EMNLP), 2021, pp. 6894-6910, doi: 10.18653/v1/2021.emnlp-main.552.

[14] Ray Smith, "An overview of the Tesseract OCR engine," in Proc. 9th Int. Conf. Document Anal. Recognit. (ICDAR), Curitiba, Brazil, 2007, pp. 629-633, doi: 10.1109/ICDAR.2007.4376991.

Downloads

Published

2026-09-30 — Updated on 2026-10-01

Versions

Data Availability Statement

This study used the publicly available MedQA, MedMCQA, and MMLU benchmark datasets; the data from the custom clinical evaluation set are available from the corresponding author upon reasonable request.

How to Cite

MERGE-Med: Multi-Expert Retrieval-Grounded Engine for Medical Report Summarization Using Hierarchical RAG and Contrastive Learning. (2026). Conference Proceedings in Science and Management, 1(1), 32-35. https://doi.org/10.68337/cpsm.v1.i1.2026-008 (Original work published 2026)