USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES
https://doi.org/10.17802/2306-1278-2026-15-4-153-163
Abstract
Highlights
- Large Language Models (LLMs) enable automation of initial abstract screening in systematic reviews, significantly reducing manual workload.
- The effectiveness of LLMs heavily depends on prompt engineering, which must clearly translate inclusion and exclusion criteria into actionable instructions.
- Two prompt design strategies are presented and illustrated: a “lenient” approach (maximizing recall) and a “strict” approach (reducing false positives).
- A practical workflow for downloading and processing PubMed abstracts using LLMs (e.g., DeepSeek, GPT‑4o, Perplexity) is demonstrated.
- Key limitations of LLMs – hallucinations, sensitivity to phrasing, and training‑data biases — are discussed, emphasizing the need for expert validation.
- A two‑stage screening strategy is recommended: first a lenient pass, followed by a stricter filter, with subsequent manual verification of relevant articles.
- Prompt engineering is framed as an iterative, domain‑specific art rather than a one‑size‑fits‑all procedure, requiring continuous refinement for each research question.
Abstract
Large Language Models have evolved into powerful tools for automating the primary screening of articles’ abstracts in systematic reviews, enabling a significant reduction in manual labor. The article presents a comprehensive review of prompt engineering principles, demonstrating how traditional meta-analysis criteria can be transformed into clear instructions for artificial intelligence. Using case studies in coronary artery bypass grafting and congenital heart disease surgery, we illustrate the impact of prompt formulation on the comprehensiveness, accuracy, and overall efficiency of literature screening. Furthermore, typical errors are discussed, and the ongoing necessity of expert oversight to minimize hallucinations and biases inherent in artificial intelligence conclusions is emphasized. Ultimately, systematic prompt engineering combined with expert evaluation allows researchers in cardiovascular surgery to optimize the search for meta-analysis sources based on Large Language Models, ensuring a faster and more comprehensive synthesis of evidence and methods for processing primary material.
About the Authors
Alexander S. ShatskiyRussian Federation
Master of Science (University of Bologna), PhD, Doctoral Candidate at the Institute of Coronary and Vascular Surgery of the Federal State Budgetary Institution “National Medical Research Center for Cardiovascular Surgery named after A.N. Bakulev” of the Ministry of Health of the Russian Federation, Moscow, Russian Federation; Deputy Director for Research at the Small Innovative Enterprise “Laboratory of Advanced Technologies” at Innopolis University, Innopolis city, Republic of Tatarstan, Russian Federation
Ehab M. Deigheidy
Russian Federation
PhD, Doctoral Candidate at the V.I. Burakovsky Institute of Cardiac Surgery, Federal State Budgetary Institution “National Medical Research Center for Cardiovascular Surgery named after A. N. Bakulev” of the Ministry of Health of the Russian Federation, Moscow, Russian Federation
Stefaniya E. Masyutina
Italy
Master of Science (Politecnico di Milano), Leading Researcher at the Small Innovative Enterprise “Laboratory of Advanced Technologies” at Innopolis University, Innopolis city, Republic of Tatarstan, Russian Federation; Researcher at Polytechnic University of Milan, Milano, Italy
Maxim L. Mamalyga
Russian Federation
PhD, MD, Leading Researcher at the Department of Surgical Treatment of Coronary Heart Disease, Institute of Coronary and Vascular Surgery, Federal State Budgetary Institution “National Medical Research Center for Cardiovascular Surgery named after A. N. Bakulev” of the Ministry of Health of the Russian Federation, Moscow, Russian Federation
References
1. Parums, D. V. Review Articles, Systematic Reviews, Meta-Analysis, and the Updated Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 Guidelines // Medical Science Monitor. 2021. Vol. 27. P. e934475.
2. Waffenschmidt S., Knelangen M., Sieben W., Bühn S., Pieper D. Single Screening versus Conventional Double Screening for Study Selection in Systematic Reviews: A Methodological Systematic Review // BMC Medical Research Methodology. 2019. Vol. 19. Art. No. 132.
3. Bramer W. M., Rethlefsen M. L., Kleijnen J., Franco O. H. Optimal Database Combinations for Literature Searches in Systematic Reviews: A Prospective Exploratory Study // Systematic Reviews. 2017. Vol. 6. Art. No. 245.
4. Cooper C., Booth A., Varley-Campbell J., Britten N., Garside R. Defining the Process to Literature Searching in Systematic Reviews: A Literature Review of Guidance and Supporting Studies // BMC Medical Research Methodology. 2018. Vol. 18. Art. No. 85.
5. Chai K. E. K., Lines R. L. J., Gucciardi D. F., Ng L. Research Screener: A Machine Learning Tool to Semi-Automate Abstract Screening for Systematic Reviews // Systematic Reviews. 2021. Vol. 10. Art. No. 93.
6. Gates A., Johnson C., Hartling L. Technology-Assisted Title and Abstract Screening for Systematic Reviews: A Retrospective Evaluation of the Abstrackr Machine Learning Tool // Systematic Reviews. 2018. Vol. 7. Art. No. 45.
7. Khraisha Q., Put S., Kappenberg J., Warraitch A., Hadfield K. Can Large Language Models Replace Humans in the Systematic Review Process? Evaluating GPT-4’s Efficacy in Screening and Extracting Data from Peer-Reviewed and Grey Literature in Multiple Languages. arXiv:2310.17526 [Preprint]. 2023.
8. Zaghir J., Naguib M., Bjelogrlic M., Névéol A., Tannier X., Lovis C. Prompt Engineering Paradigms for Medical Applications: Scoping Review // Journal of Medical Internet Research. 2024. Vol. 26. P. e60501.
9. Wang Z., Chu Z., Doan T. V., Ni S., Yang M., Zhang W. History, Development, and Principles of Large Language Models: An Introductory Survey // AI Ethics. 2024. P. 1–17.
10. Shiri F. M., Perumal T., Mustapha N., Mohamed R. A Comprehensive Overview and Comparative Analysis on Deep Learning Models. arXiv:2305.17473 [Preprint]. 2023.
11. Gue C. C. Y., Rahim N. D. A., Rojas-Carabali W., Agrawal R., Palvannan R. K., Abisheganaden J., Yip W. F. Evaluating the OpenAI’s GPT-3.5 Turbo’s Performance in Extracting Information from Scientific Articles on Diabetic Retinopathy // Systematic Reviews. 2024. Vol. 13. Art. No. 135.
12. Wang J., Yang Z., Yao Z., Yu H. JMLR: Joint Medical LLM and Retrieval Training for Enhancing Reasoning and Professional Question Answering Capability. arXiv:2402.17887 [Preprint]. 2024.
13. Huang L., Yu W., Ma W., Zhong W., Feng Z., Wang H., Chen Q., Peng W., Feng X., Qin B., et al. A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions // ACM Transactions on Information Systems. 2025. Vol. 43. P. 1–55.
14. Sclar M., Choi Y., Tsvetkov Y., Suhr A. Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I Learned to Start Worrying about Prompt Formatting. arXiv:2310.11324 [Preprint]. 2023.
15. Saloojee H., Pettifor J. M. Maximizing Access and Minimizing Barriers to Research in Low- and Middle-Income Countries: Open Access and Health Equity // Global Health, Science and Practice. 2022. Vol. 10, No. 1. P. e2100403.
16. Ye A., Maiti A., Schmidt M., Pedersen S. J. A Hybrid Semi-Automated Workflow for Systematic and Literature Review Processes with Large Language Model Analysis // Future Internet. 2024. Vol. 16, No. 5. P. 167.
17. Liu P., Yuan W., Fu J., Jiang Z., Hayashi H., Neubig G. Pre-Train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing // ACM Computing Surveys. 2023. Vol. 55, No. 9. P. 1–35.
18. Eriksen M. B., Frandsen T. F. The Impact of Patient, Intervention, Comparison, Outcome (PICO) as a Search Strategy Tool on Literature Search Quality: A Systematic Review // Journal of the Medical Library Association. 2018. Vol. 106, No. 4. P. 420–431.
19. Kusano G., Akimoto K., Takeoka K. Are Longer Prompts Always Better? Prompt Selection in Large Language Models for Recommendation Systems. arXiv:2412.14454 [Preprint]. 2024.
20. Heston T. F., Khun C. Prompt Engineering in Medical Education // International Medical Education. 2023. Vol. 2, No. 4. P. 198–205.
21. Lee J., Hicke Y., Yu R., Brooks C., Kizilcec R. F. The Life Cycle of Large Language Models in Education: A Framework for Understanding Sources of Bias // British Journal of Educational Technology. 2024. Vol. 55, No. 4. P. 1982–2002.
22. Huang D., Bu Q., Zhang J., Xie X., Chen J., Cui H. Bias Testing and Mitigation in LLM-Based Code Generation. arXiv:2309.14345 [Preprint]. 2023.
23.
Review
For citations:
Shatskiy A.S., Deigheidy E.M., Masyutina S.E., Mamalyga M.L. USING LARGE LANGUAGE MODELS FOR LITERATURE SEARCH IN CARDIOVASCULAR SURGERY SYSTEMATIC REVIEWS AND META-ANALYSES. Complex Issues of Cardiovascular Diseases. 2026;15(4):153-163. (In Russ.) https://doi.org/10.17802/2306-1278-2026-15-4-153-163
JATS XML

































