Large language models for intelligent healthcare and medicine: A systematic survey

Authors

  • Chaolei Wu College of Information Engineering, Hubei University of Chinese Medicine, Wuhan 430065, P. R. China
  • Yanhui Zhu School of Computing, Montclair State University, Montclair NJ 07043, United States (Email: zhuya@montclair.edu)

Abstract

Against the backdrop of healthcare digitalization, intelligent medical systems, and the rapid expansion of multi-source medical data, large language models are reshaping medical artificial intelligence from a traditional task-driven paradigm into a generative, interactive, and multimodal intelligence paradigm. This paper systematically reviews recent progress in large language models and multimodal large language models for intelligent healthcare and medicine, with emphasis on technological evolution, model architectures, training data, evaluation systems, medical applications, and trustworthy deployment. The review first traces the development of medical artificial intelligence from supervised learning and pre-trained language models to large language models and multimodal large language models. It then summarizes the backbone architectures and domain adaptation methods of medical large language models, as well as modality encoding, alignment mechanisms, and task interfaces in multimodal systems. In addition, we discuss medical data resources, fine-tuning strategies, evaluation frameworks, and their practical value in diagnosis augmentation, clinical document generation, medical education, mental health services, and specialized clinical scenarios. Although existing studies have demonstrated substantial potential, medical large language models still face critical challenges, including hallucinations, insufficient factual reliability, privacy and data security risks, fairness and bias concerns, limited interpretability, and barriers to clinical deployment. Finally, we outline promising future directions, including retrieval-augmented generation, multimodal fusion, privacy-preserving training, human-machine collaboration, and general medical artificial intelligence. This survey aims to provide a comprehensive reference for trustworthy research, engineering implementation, and intelligent healthcare applications of medical large language models.

Document Type: Invited review

Cited as: Wu, C., Zhu, Y. Large language models for intelligent healthcare and medicine: A systematic survey. Advanced IntelliEngineering, 2026, 1(1): 21-44. https://doi.org/10.46690/aie.2026.01.03

DOI:

https://doi.org/10.46690/aie.2026.01.03

Keywords:

Large language models, medical artificial intelligence, intelligent healthcare, medical engineering, intelligent medical systems

References

[1] Vollset, S. E., Ababneh, H. S., Abate, Y. H., et al. Burden of disease scenarios for 204 countries and territories, 2022-2050: A forecasting analysis for the global burden of disease study 2021. The Lancet, 2024, 403(10440): 2204-2256.

[2] Boniol, M., Kunjumen, T., Nair, T. S., et al. The global health workforce stock and distribution in 2020 and 2030: A threat to equity and 'universal' health coverage? BMJ Global Health, 2022, 7(6): e009316.

[3] Huesmann, L., Sudacka, M., Durning, S. J., et al. Clinical reasoning: What do nurses, physicians, and students reason about. Journal of Interprofessional Care, 2023, 37(6): 990-998.

[4] Schuler, K., Jung, I. C., Zerlik, M., et al. Context factors in clinical decision-making: A scoping review. BMC Medical Informatics and Decision Making, 2025, 25(1): 133.

[5] Sokol, K., Fackler, J., Vogt, J. E. Artificial intelligence should genuinely support clinical reasoning and decision making to bridge the translational gap. npj Digital Medicine, 2025, 8(1): 345.

[6] Fahim, Y. A., Hasani, I. W., Kabba, S., et al. Artificial intelligence in healthcare and medicine: Clinical applications, therapeutic advances, and future perspectives. European Journal of Medical Research, 2025, 30(1): 848.

[7] Narayan, S. M., Chung, M. K., Adedinsewo, D., et al. Access to digital health technologies: Personalized framework and global perspectives. Nature Reviews Cardiology, 2026, 23(1): 9-22.

[8] AlSaad, R., Abd-Alrazaq, A., Boughorbel, S., et al. Multimodal large language models in health care: Applications, challenges, and future outlook. Journal of Medical Internet Research, 2024, 26: e59505.

[9] Jandoubi, B., Akhloufi, M. A. Multimodal artificial intelligence in medical diagnostics. Information, 2025, 16(7): 591.

[10] Buess, L., Keicher, M., Navab, N., et al. From large language models to multimodal AI: A scoping review on the potential of generative AI in medicine. Biomedical Engineering Letters, 2025, 15(5): 845-863.

[11] Chen, S. F., Alyakin, A., Seas, A., et al. LLM-assisted systematic review of large language models in clinical medicine. Nature Medicine, 2026, 32(3): 1152-1159.

[12] Singhal, K., Azizi, S., Tu, T., et al. Large language models encode clinical knowledge. Nature, 2023, 620(7972): 172-180.

[13] Busch, F., Hoffmann, L., Rueger, C., et al. Current applications and challenges in large language models for patient care: A systematic review. Communications Medicine, 2025, 5(1): 26.

[14] Kung, T. H., Cheatham, M., Medenilla, A., et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLOS Digital Health, 2023, 2(2): e0000198.

[15] Kyu-Hwan, J. Large language models in medicine: Clinical applications, technical challenges, and ethical considerations. Healthcare Informatics Research, 2025, 31(2): 114-124.

[16] Bedi, S., Liu, Y., Orr-Ewing, L., et al. Testing and evaluation of health care applications of large language models: A systematic review. The Journal of the American Medical Association, 2025, 333(4): 319-328.

[17] Fareed, M., Fatima, M., Uddin, J., et al. A systematic review of ethical considerations of large language models in healthcare and medicine. Frontiers in Digital Health, 2025, 7: 1653631.

[18] Pantanowitz, L., Pearce, T., Abukhiran, I., et al. Non-generative artificial intelligence in medicine: Advancements and applications in supervised and unsupervised machine learning. Modern Pathology, 2025, 38(3): 100680.

[19] Shaikh, M. R., Jeyabose, A., Arjunan, R. V. Deep learning for Alzheimer's disease: Advances in classification, segmentation, subtyping, and explainability. BioMedical Engineering OnLine, 2025, 24(1): 150.

[20] Cao, L., Wu, C., Luo, G., et al. Online biomedical named entities recognition by data and knowledge-driven model. Artificial Intelligence in Medicine, 2024, 150: 102813.

[21] Huang, M., Han, J., Lin, P., et al. Surveying biomedical relation extraction: A critical examination of current datasets and the proposal of a new resource. Briefings in Bioinformatics, 2024, 25(3): bbae132.

[22] Prinzi, F., Currieri, T., Galio, S., et al. Shallow and deep learning classifiers in medical image analysis. European Radiology Experimental, 2024, 8(1): 26.

[23] Sahoo, S. S., Plasek, J. M., Xu, H., et al. Large language models for biomedicine: Foundations, opportunities, challenges, and best practices. Journal of the American Medical Informatics Association, 2024, 31(9): 2114-2124.

[24] De Santis, E., Martino, A., Ronci, F., et al. From bag-of-words to transformers: A comparative study for text classification in healthcare discussions in social media. IEEE Transactions on Emerging Topics in Computational Intelligence, 2025, 9(1): 1063-1077.

[25] Liang, Z., Zhao, Y., Xu, H., et al. A hybrid model integrating RoBERTa, TF-IDF, and attention mechanism for medical query intent classification. Scientific Reports, 2025, 15(1): 42847.

[26] Wang, C., Pan, D., Kumari, S., et al. Large language models driven health text information analysis in consumer electronics. IEEE Transactions on Consumer Electronics, 2025, 71(2): 3510-3521.

[27] Shortliffe, E. H., Davis, R., Axline, S. G., et al. Computer-based consultations in clinical therapeutics: Explanation and rule acquisition capabilities of the mycin system. Computers and Biomedical Research, 1975, 8(4): 303-320.

[28] Kononenko, I. Machine learning for medical diagnosis: History, state of the art and perspective. Artificial Intelligence in Medicine, 2001, 23(1): 89-109.

[29] Lu, C., Zhang, J., Liu, R. Deep learning-based image classification for integrating pathology and radiology in AI-assisted medical imaging. Scientific Reports, 2025, 15(1): 27029.

[30] Kumar, M., Shubham, Saini, L. M., et al. Transformers for enhanced biomedical text classification. Paper Presented at the 4th International Conference on Innovative Mechanisms for Industry Applications (ICIMIA), Tirupur, India, 3-5 September, 2025.

[31] Ahmed, S. F., Alam, M. S. B., Hassan, M., et al. Deep learning modelling techniques: Current progress, applications, advantages, and challenges. Artificial Intelligence Review, 2023, 56(11): 13521-13617.

[32] Zhou, J., Li, H., Chen, S., et al. Large language models in biomedicine and healthcare. npj Artificial Intelligence, 2025, 1(1): 44.

[33] Gu, Y., Tinn, R., Cheng, H., et al. Domain-specific language model pretraining for biomedical natural language processing. ACM Transactions on Computing for Healthcare, 2021, 3(1): 2.

[34] Lee, J., Yoon, W., Kim, S., et al. BioBERT: A pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 2020, 36(4): 1234-1240.

[35] Fraile Navarro, D., Ijaz, K., Rezazadegan, D., et al. Clinical named entity recognition and relation extraction using natural language processing of medical free text: A systematic review. International Journal of Medical Informatics, 2023, 177: 105122.

[36] Xie, Q., Chen, Q., Chen, A., et al. Medical foundation large language models for comprehensive text analysis and beyond. npj Digital Medicine, 2025, 8(1): 141.

[37] Wei, J., Wang, X., Schuurmans, D., et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 2022, 35: 24824-24837.

[38] Singhal, K., Tu, T., Gottweis, J., et al. Toward expert-level medical question answering with large language models. Nature Medicine, 2025, 31(3): 943-950.

[39] Tu, T., Azizi, S., Driess, D., et al. Towards generalist biomedical AI. The New England Journal of Medicine Artificial Intelligence, 2024, 1(3): AIoa2300138.

[40] Team, G., Georgiev, P., Lei, V. I., et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. ArXiv Preprint ArXiv: 2403.05530, 2024.

[41] Saab, K., Tu, T., Weng, W., et al. Capabilities of Gemini models in medicine. ArXiv Preprint ArXiv: 2404.18416, 2024.

[42] Wang, X., Wei, J., Schuurmans, D., et al. Self-consistency improves chain of thought reasoning in language models. ArXiv Preprint ArXiv: 2203.11171, 2022.

[43] Lewis, P., Perez, E., Piktus, A., et al. Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 2020, 33: 9459-9474.

[44] Jaech, A., Kalai, A., Lerer, A., et al. OpenAI o1 system card. ArXiv Preprint ArXiv: 2412.16720, 2024.

[45] Guo, D., Yang, D., Zhang, H., et al. Deepseek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 2025, 645(8081): 633-638.

[46] Luo, R., Sun, L., Xia, Y., et al. BioGPT: Generative pre-trained transformer for biomedical text generation and mining. Briefings in Bioinformatics, 2022, 23(6): bbac409.

[47] Raffel, C., Shazeer, N., Roberts, A., et al. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 2020, 21(1): 140.

[48] Lewis, M., Liu, Y., Goyal, N., et al. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. Paper Presented at the 58th Annual Meeting of the Association for Computational Linguistics, Virtual, 5-10 July, 2020.

[49] Phan, L., Anibal, J. T., Tran, H. T., et al. SciFive: A text-to-text transformer model for biomedical literature. ArXiv Preprint ArXiv: 2106.03598, 2021.

[50] Yuan, H., Yuan, Z., Gan, R., et al. BioBART: Pretraining and evaluation of a biomedical generative language model. Paper Presented at the 21st Workshop on Biomedical Language Processing, Dublin, Ireland, 26 May, 2022.

[51] WHO. Ethics and Governance of Artificial Intelligence for Health. World Health Organization, Geneva, Switzerland, 2021.

[52] Welch, M. L., Grant, B., Deutschman, C., et al. A practical framework for operationalising responsible and equitable artificial intelligence in health care: Tackling bias, inequity, and implementation challenges. The Lancet Digital Health, 2026, 8(3): e100957.

[53] Asgari, E., Montaña-Brown, N., Dubois, M., et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine, 2025, 8(1): 274.

[54] Farquhar, S., Kossen, J., Kuhn, L., et al. Detecting hallucinations in large language models using semantic entropy. Nature, 2024, 630(8017): 625-630.

[55] Bedi, S., Cui, H., Fuentes, M., et al. Holistic evaluation of large language models for medical tasks with Med-HELM. Nature Medicine, 2026, 32(3): 943-951.

[56] Christophe, C., Raha, T., Maslenkova, S., et al. Beyond fine-tuning: Unleashing the potential of continuous pre-training for clinical LLMs. Paper Presented at Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, USA, 12-16 November, 2024.

[57] Tran, H., Yang, Z., Yao, Z., et al. Bioinstruct: Instruction tuning of large language models for biomedical natural language processing. Journal of the American Medical Informatics Association, 2024, 31(9): 1821-1832.

[58] Wu, C., Qiu, P., Liu, J., et al. Towards evaluating and building versatile large language models for medicine. npj Digital Medicine, 2025, 8(1): 58.

[59] Giuffrè, M., You, K., Pang, Z., et al. Expert of experts verification and alignment (EVAL) framework for large language models safety in gastroenterology. npj Digital Medicine, 2025, 8(1): 242.

[60] Wu, C., Zhang, X., Zhang, Y., et al. Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data. Nature Communications, 2025, 16(1): 7866.

[61] Li, C., Chang, K., Yang, C., et al. Towards a holistic framework for multimodal LLM in 3D brain CT radiology report generation. Nature Communications, 2025, 16(1): 2258.

[62] Liu, C., Ouyang, C., Wan, Z., et al. Knowledge-enhanced multimodal ECG representation learning with arbitrary-lead inputs. Paper Presented at the Association for Computational Linguistics, Vienna, Austria, 27 July-1 August, 2025.

[63] Hu, Y., Li, T., Lu, Q., et al. OmniMedVQA: A new large-scale comprehensive evaluation benchmark for medical LVLM. Paper Presented at the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, United States, 16-22 June, 2024.

[64] Tanno, R., Barrett, D. G. T., Sellergren, A., et al. Collaboration between clinicians and vision-language models in radiology report generation. Nature Medicine, 2025, 31: 599-608.

[65] Hu, Y., Chen, Q., Du, J., et al. Improving large language models for clinical named entity recognition via prompt engineering. Journal of the American Medical Informatics Association, 2024, 31(9): 1812-1820.

[66] Johnson, A. E., Pollard, T. J., Shen, L., et al. MIMIC-III, a freely accessible critical care database. Scientific Data, 2016, 3(1): 160035.

[67] Johnson, A. E. W., Bulgarelli, L., Shen, L., et al. MIMIC-IV, a freely accessible electronic health record dataset. Scientific Data, 2023, 10(1): 1.

[68] Herrett, E., Gallagher, A. M., Bhaskaran, K., et al. Data resource profile: Clinical practice research datalink (CPRD). International Journal of Epidemiology, 2015, 44(3): 827-836.

[69] Jin, Q., Dhingra, B., Liu, Z., et al. PubMedQA: A dataset for biomedical research question answering. Paper Presented at the Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, Hong Kong, China, 3-7 November, 2019.

[70] Jin, D., Pan, E., Oufattole, N., et al. What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Applied Sciences, 2021, 11(14): 6421.

[71] Pal, A., Umapathi, L. K., Sankarasubbu, M. MedMCQA: A large-scale multi-subject multi-choice dataset for medical domain question answering. Paper Presented at the Conference on Health, Inference, and Learning, Virtual, 7-8 April, 2022.

[72] Yang, S., Zhao, H., Zhu, S., et al. Zhongjing: Enhancing the Chinese medical capabilities of large language model through expert feedback and real-world multi-turn dialogue. Paper Presented at the 38th Annual AAAI Conference on Artificial Intelligence, Vancouver, Canada, 20-27 February, 2024.

[73] Zeng, G., Yang, W., Ju, Z., et al. Meddialog: Large-scale medical dialogue datasets. Paper Presented at the Conference on Empirical Methods in Natural Language Processing, Virtual, 16-20 November, 2020.

[74] Li, Y., Li, Z., Zhang, K., et al. ChatDoctor: A medical chat model fine-tuned on a large language model Meta-AL (LLaMA) using medical domain knowledge. Cureus, 2023, 15(6): e40895.

[75] Wu, C., Lin, W., Zhang, X., et al. PMC-LLaMA: Toward building open-source language models for medicine. Journal of the American Medical Informatics Association, 2024, 31(9): 1833-1843.

[76] Zhang, X., Tian, C., Yang, X., et al. Alpacare: Instruction-tuned large language models for medical application. ArXiv Preprint ArXiv: 2310.14558, 2023.

[77] Bodenreider, O. The unified medical language system (UMLS): Integrating biomedical terminology. Nucleic Acids Research, 2004, 32(suppl_1): D267-D270.

[78] Ao, D., Yang, Y., Sui, Z., et al. Preliminary study on the construction of Chinese medical knowledge graph. Journal of Chinese Information Processing, 2019, 33(10): 9. (in Chinese)

[79] Basaldella, M., Liu, F., Shareghi, E., et al. COMETA: A corpus for medical entity linking in the social media. Paper Presented at the Conference on Empirical Methods in Natural Language Processing, Virtual, 16-20 November, 2020.

[80] Xiao, H., Zhou, F., Liu, X., et al. A comprehensive survey of large language models and multimodal large language models in medicine. Information Fusion, 2025, 117: 102888.

[81] Zhang, X., Wu, C., Zhao, Z., et al. PMC-VQA: Visual instruction tuning for medical visual question answering. ArXiv Preprint ArXiv: 2305.10415, 2023.

[82] He, X., Zhang, Y., Mou, L., et al. PathVQA: 30000+ questions for medical visual question answering. ArXiv Preprint ArXiv: 2003.10286, 2020.

[83] Hu, X., Gu, L., Kobayashi, K., et al. Medical-CXR-VQA dataset: A large-scale LLM-enhanced medical dataset for visual question answering on chest X-ray images, 2025.

[84] Johnson, A. E. W., Pollard, T. J., Berkowitz, S. J., et al. MIMIC-CXR, a de-identified publicly available database of chest radiographs with free-text reports. Scientific Data, 2019, 6(1): 317.

[85] Irvin, J. Rajpurkar, P., Ko, M., et al. ChexPert: A large chest radiograph dataset with uncertainty labels and expert comparison. Paper Presented at the AAAI Conference on Artificial Intelligence, Hawaii, United States, 27 January-1 February, 2019.

[86] Zhang, D., Lan, X., Geng, S., et al. MEETI: A multimodal ECG dataset from MIMIC-IV-ECG with signals, images, features and interpretations. Scientific Data, 2026, 13(1): 527.

[87] Sun, Y., Zhu, C., Zheng, S., et al. PathAsst: A generative foundation AI assistant towards artificial general intelligence of pathology. Paper Presented at the 38th Annual AAAI Conference on Artificial Intelligence, Vancouver, Canada, 20-27 February, 2024.

[88] Lin, W., Zhao, Z., Zhang, X., et al. PMC-CLIP: Contrastive language-image pre-training using biomedical documents. Paper Presented at International Conference on Medical Image Computing and Computer-Assisted Intervention, Vancouver, Canada, 8-12 October, 2023.

[89] Xie, Y., Zhou, C., Gao, L., et al. Medtrinity-25m: A large-scale multimodal dataset with multigranular annotations for medicine. Paper Presented at the International Conference on Learning Representations (ICLR), Singapore, 24-28 April, 2025.

[90] Yan, S., Hu, M., Jiang, Y., et al. Derm1m: A million-scale vision-language dataset aligned with clinical ontology knowledge for dermatology. Paper Presented at the IEEE/CVF International Conference on Computer Vision, Hawaii, United States, 19-23 October, 2025.

[91] Chen, Z., Varma, M., Delbrouck, J.-B., et al. Chexagent: Towards a foundation model for chest X-ray interpretation. Paper Presented at the Association for the Advancement of Artificial Intelligence’s 2024 Spring Symposium Series, California, United States, 25-27 March, 2024.

[92] Jonnagaddala, J., Wong, Z. S.-Y. Privacy preserving strategies for electronic health records in the era of large language models. npj Digital Medicine, 2025, 8(1): 34.

[93] Zhong, X., Li, S., Chen, Z., et al. Considerations for patient privacy of large language models in health care: Scoping review. Journal of Medical Internet Research, 2025, 27: e76571.

[94] Chen, R. J., Wang, J. J., Williamson, D. F. K., et al. Algorithmic fairness in artificial intelligence for medicine and healthcare. Nature Biomedical Engineering, 2023, 7(6): 719-742.

[95] McDuff, D., Schaekermann, M., Tu, T., et al. Towards accurate differential diagnosis with large language models. Nature, 2025, 642(8067): 451-457.

[96] Gema, A., Minervini, P., Daines, L., et al. Parameter-efficient fine-tuning of LLaMA for the clinical domain. Paper Presented at the 6th Clinical Natural Language Processing Workshop, Mexico City, Mexico, 21 June, 2024.

[97] Shi, W., Xu, R., Zhuang, Y., et al. Medadapter: Efficient test-time adaptation of large language models towards medical reasoning. Paper Presented at Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, USA, 12-16 November, 2024.

[98] Rafailov, R., Sharma, A., Mitchell, E., et al. Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 2023, 36: 53728-53741.

[99] Kim, D., Lee, J., Yun, J., et al. Benchmarking direct preference optimization for medical large vision-language models. ArXiv Preprint ArXiv: 2601.17918, 2026.

[100] Zhang, M., Shen, Y., Li, Z., et al. LLMEval-Med: A real-world clinical benchmark for medical LLMs with physician validation. ArXiv Preprint ArXiv: 2506.04078, 2025.

[101] Tam, T. Y. C., Sivarajkumar, S., Kapoor, S., et al. A framework for human evaluation of large language models in healthcare derived from literature review. npj Digital Medicine, 2024, 7(1): 258.

[102] Artsi, Y., Sorin, V., Glicksberg, B. S., et al. Large language models in real-world clinical workflows: A systematic review of applications and implementation. Frontiers in Digital Health, 2025, 7: 1659134.

[103] Gong, L., Fang, W., Yang, T., et al. Meddialogrubrics: A comprehensive benchmark and evaluation framework for multi-turn medical consultations in large language models. ArXiv Preprint ArXiv: 2601.03023, 2026.

[104] Guo, Z., Lai, A., Thygesen, J. H., et al. Large language models for mental health applications: Systematic review. JMIR Mental Health, 2024, 11(1): e57400.

[105] Li, L., Huang, X., Ma, R., et al. LLM use for mental health: Crowdsourcing users’ sentiment-based perspectives and values from social discussions. Paper Presented at the Web Conference, Dubai, The United Arab Emirates, 13-17 April, 2026.

[106] Vasey, B., Nagrendran, M., Campbell, B., et al. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ-British Medical Journal, 2022, 377: e070904.

[107] Wiest, I. C., Ferber, D., Zhu, J., et al. Privacy-preserving large language models for structured medical information retrieval. npj Digital Medicine, 2024, 7(1): 257.

[108] Zhao, J., Xu, L., Tan, M., et al. Rxsafe-bench: Identifying medication safety issues of large language models in simulated consultations. ArXiv Preprint ArXiv: 2511.04328, 2025.

[109] Kazemzadeh, H., Dizaji, K. M., Tavakoli, S. R., et al. Druggrag: Enhancing pharmacy LLM performance through a novel retrieval-augmented generation pipeline. ArXiv Preprint ArXiv: 2512.14896, 2025.

[110] Xu, S., Yan, Z., Dai, C., et al. MEGA-RAG: A retrieval-augmented generation framework with multi-evidence guided answer refinement for mitigating hallucinations of LLMs in public health. Frontiers in Public Health, 2025, 13: 1653381.

[111] Carlini, N., Tramèr, F., Wallace, E., et al. Extracting training data from large language models. Paper Presented at the 30th USENIX Security Symposium, Vancouver, Canada, 8-10 August, 2021.

[112] Yang, X., Wu, C., Yan, X., et al. Blockchain-based healthcare and medicine data sharing and service system. Paper Presented at the International Conference on Blockchain and Trustworthy Systems, Chengdu, China, 4-5 August, 2022.

[113] Cross, J. L., Choma, M. A., Onofrey, J. A. Bias in medical AI: Implications for clinical decision-making. PLOS Digital Health, 2024, 3(11): e0000651.

[114] Pfohl, S. R., Cole-Lewis, H., Sayres, R., et al. A toolbox for surfacing health equity harms and biases in large language models. Nature Medicine, 2024, 30(12): 3590-3600.

[115] Ji, Y., Ma, W., Sivarajkumar, S., et al. Mitigating the risk of health inequity exacerbated by large language models. npj Digital Medicine, 2025, 8(1): 246.

[116] Celis, E., Keswani, V., Straszak, D., et al. Fair and diverse DPP-based data summarization. Paper Presented at the Thirty-Fifth International Conference on Machine Learning, Stockholm, Sweden, 10-15 July, 2018.

[117] El Halabi, M., Mitrović, S., Norouzi-Fard, A., et al. Fairness in streaming submodular maximization: Algorithms and hardness. Advances in Neural Information Processing Systems, 2020, 33: 13609-13622.

[118] Zhu, Y., Basu, S., Pavan, A. Fairness in monotone k-submodular maximization: Algorithms and applications. Paper Presented at the IEEE International Conference on Big Data, Washington, USA, 15-18 December, 2024.

[119] Mesinovic, M., Watkinson, P., Zhu, T. Explainability in the age of large language models for healthcare. Communications Engineering, 2025, 4(1): 128.

[120] Liu, X., Rivera, S. C., Moher, D., et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. The Lancet Digital Health, 2020, 2(10): e537-e548.

[121] Ribeiro, M. T., Singh, S., Guestrin, C. "Why should I trust you?" Explaining the predictions of any classifier. Paper Presented at the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, San Francisco, USA, 13-17 August, 2016.

[122] Lin, H., Bilmes, J. A class of submodular functions for document summarization. Paper Presented at the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, Oregon, USA, 19-24 June, 2011.

[123] Zhu, Y. Submodular optimization: Variants, theory and applications. Paper Presented at the 33rd ACM International Conference on Information and Knowledge Management, Idaho, USA, 21-25 October, 2024.

[124] Zhu, Y., Basu, S., Pavan, A. Improved evolutionary algorithms for submodular maximization with cost constraints. ArXiv Preprint ArXiv: 2405.05942, 2024a.

[125] Zhu, Y., Basu, S., Pavan, A. Regularized unconstrained weakly submodular maximization. Paper Presented at the 33rd ACM International Conference on Information and Knowledge Management, Idaho, USA, 21-25 October, 2024b.

[126] Jin, W., Li, X., Fatehi, M., et al. Guidelines and evaluation of clinical explainable AI in medical image analysis. Medical Image Analysis, 2023, 84: 102684.

Downloads

Download data is not yet available.

Downloads

Published

2026-06-05

How to Cite

Wu, C., & Zhu, Y. (2026). Large language models for intelligent healthcare and medicine: A systematic survey. Advanced IntelliEngineering, 1(1), 21–44. https://doi.org/10.46690/aie.2026.01.03

Issue

Section

Articles