PubMed دسترسی آزاد

Fine-Tuning, Retrieval-Augmented Generation, and Hybrid Adaptation of Language Models for Clinical Decision-Making in Health Care: Systematic Review.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Large language models (LLMs) demonstrate strong performance on medical knowledge benchmarks, but their safe and effective use in clinical practice depends on posttraining adaptation rather than raw model capability. Fine-tuning, retrieval-augmented generation (RAG), and hybrid approaches are principal strategies for grounding language models in clinical evidence, yet their comparative effectiveness remains unclear. OBJECTIVE: This systematic review aims to synthesize evidence on fine-tuning, RAG, and hybrid posttraining strategies for clinical diagnosis and decision-support tasks and to identify strategy-task alignments and methodological features associated with improved performance. METHODS: We conducted a systematic review in accordance with PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. PubMed/MEDLINE, Scopus, and Web of Science were searched from January 2018 through May 2026. Eligible studies evaluated transformer-based language models that underwent posttraining adaptation, retrieval augmentation, or both for clinical decision support, diagnosis, triage, risk stratification, or related health care applications. Studies evaluating nonadapted models, non-language-model AI systems, prompt engineering without performance evaluation, or nonclinical applications were excluded. Data extracted included model architecture, adaptation strategy, clinical domain, validation approach, and performance outcomes. Risk of bias was assessed using PROBAST+AI (Prediction model Risk of Bias Assessment Tool for AI). Studies were grouped according to the primary enhancement strategy (fine-tuning or parameter-efficient fine-tuning, RAG, or hybrid approaches), and findings were synthesized descriptively. RESULTS: Of 1890 identified records, 35 studies published between 2024 and 2026 met eligibility criteria. Enhancement strategies included RAG (17/35, 48.6%), fine-tuning or parameter-efficient fine-tuning (7/35, 20%), and hybrid approaches (11/35, 31.4%). Studies included diverse specialties from oncology, neurology, radiology, mental health, cardiology, ophthalmology, and surgical care. Fine-tuning demonstrated strong performance for task-specific applications, achieving area under the receiver operating characteristic curve values up to 0.912 for cancer detection and area under curve of 0.892 for major depressive disorder prediction, while matching clinician-level diagnostic performance in several studies. RAG improved guideline adherence and diagnostic accuracy, with increases from 71.1% to 92.1% and from 78.9% to 94.7% in guideline-based decision-support tasks. However, benefits were inconsistent across larger reasoning-capable models. Hybrid systems generally achieved the strongest performance in complex clinical workflows, with external validation accuracies exceeding 90% in stroke triage, dermatology, multimodal imaging, and oncology applications. Risk-of-bias assessment identified substantial methodological limitations, with 25 studies judged as high risk, 9 as unclear risk, and only 1 as low risk overall. Common concerns included inadequate external validation, lack of calibration assessment, nonrepresentative participant selection, and insufficient reporting of analytical methods. CONCLUSIONS: Adaptation strategies should align with task needs, using fine-tuning for narrow classification, RAG for guideline-grounded reasoning, and hybrid approaches for complex multimodal tasks. However, the evidence base remains largely retrospective or benchmark-based. Prospective studies with external validation, calibration, and standardized safety reporting are needed before broader clinical use.

نتیجه فارسی

این مرور سیستماتیک ۳۵ مطالعه را در مورد استراتژی‌های تنظیم دقیق، RAG و هیبریدی برای پشتیبانی بالینی بررسی می‌کند. نتایج نشان می‌دهد که تنظیم دقیق برای طبقه‌بندی‌های خاص و RAG برای استدلال مبتنی بر دستورالعمل‌ها مؤثر است، در حالی که سیستم‌های هیبریدی برای وظایف چندوجهی پیچیده عملکرد بهتری دارند. با این حال، شواهد عمدتاً بازتابی یا مبتنی بر معیار هستند و ارزیابی‌های خارجی و گزارش‌سازی استاندارد ایمنی هنوز ناکافی است.

  • تنظیم دقیق برای طبقه‌بندی‌های خاص و RAG برای استدلال مبتنی بر دستورالعمل‌ها مناسب‌ترند.
  • سیستم‌های هیبریدی در جریان‌های کاری بالینی پیچیده عملکرد بهتری دارند.
  • ارزیابی ریسک سوگیری روش‌شناختی نشان‌دهنده محدودیت‌های جدی در مطالعات است.
  • شواهد عمدتاً بازتابی یا مبتنی بر معیار هستند و نیاز به مطالعات آینده‌نگرانه با ارزیابی خارجی دارند.

ترجمه فارسی چکیده

مدل‌های زبانی بزرگ (LLM) عملکرد قوی در معیارهای دانش پزشکی نشان می‌دهند، اما استفاده ایمن و مؤثر آن‌ها در عمل بالینی به جای توانایی خام مدل، به سازگاری پس از آموزش بستگی دارد. تنظیم دقیق (Fine-tuning)، تولید با تقویت بازیابی (RAG) و رویکردهای هیبریدی استراتژی‌های اصلی برای ریشه‌دار کردن مدل‌های زبانی در شواهد بالینی هستند، اما اثربخشی مقایسه‌ای آن‌ها همچنان نامشخص است. این مرور سیستماتیک به هدف جمع‌آوری شواهد در مورد استراتژی‌های پس از آموزش تنظیم دقیق، RAG و هیبریدی برای وظایف تشخیص و پشتیبانی تصمیم‌گیری بالینی است. بر اساس دستورالعمل‌های PRISMA 2020، جستجو در پایگاه‌های PubMed/MEDLINE، Scopus و Web of Science از ژانویه ۲۰۱۸ تا مه ۲۰۲۶ انجام شد. ۳۵ مطالعه منتشر شده بین ۲۰۲۴ و ۲۰۲۶ شامل شدند. استراتژی‌های بهبود شامل RAG (۴۸.۶٪)، تنظیم دقیق یا تنظیم دقیق کارآمد پارامتری (۲۰٪) و رویکردهای هیبریدی (۳۱.۴٪) بود. تنظیم دقیق عملکرد قوی برای کاربردهای خاص وظیفه را نشان داد، در حالی که RAG به اطاعت از دستورالعمل‌ها و دقت تشخیصی افزایش قابل توجهی داد. سیستم‌های هیبریدی عملکرد قوی‌تری در جریان‌های کاری پیچیده بالینی داشتند. ارزیابی ریسک سوگیری روش‌شناختی محدودیت‌های جدی را شناسایی کرد.

روش پژوهش

این مرور بر اساس دستورالعمل‌های PRISMA 2020 انجام شد. جستجو در پایگاه‌های PubMed/MEDLINE، Scopus و Web of Science از ژانویه ۲۰۱۸ تا مه ۲۰۲۶ انجام شد. مطالعات شامل مدل‌های زبانی مبتنی بر ترنسفورمر بودند که پس از آموزش سازگاری، تقویت بازیابی یا هر دو را تجربه کرده بودند. ارزیابی ریسک سوگیری با استفاده از PROBAST+AI انجام شد.

محدودیت‌ها

ارزیابی ریسک سوگیری روش‌شناختی محدودیت‌های جدی را شناسایی کرد. ۲۵ مطالعه دارای ریسک بالا، ۹ مطالعه دارای ریسک نامشخص و تنها ۱ مطالعه دارای ریسک پایین بودند. نگرانی‌های رایج شامل ارزیابی خارجی ناکافی، فقدان ارزیابی کالیبراسیون، انتخاب ناکافی شرکت‌کنندگان و گزارش ناکافی روش‌های تحلیلی بود.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
مدل‌های زبانی مبتنی بر ترنسفورمر که برای پشتیبانی تصمیم‌گیری بالینی، تشخیص، اولویت‌بندی، ریسک‌سنجی یا کاربردهای مرتبط بهداشتی سازگاری یافته بودند.
مداخله/مواجهه
استراتژی‌های پس از آموزش شامل تنظیم دقیق یا تنظیم دقیق کارآمد پارامتری، تولید با تقویت بازیابی (RAG) و رویکردهای هیبریدی بود.
مقایسه
مدل‌های زبانی غیرسازگاری‌یافته یا سیستم‌های هوش مصنوعی غیرزبانی.
حجم نمونه
۳۵ مطالعه منتشر شده بین ۲۰۲۴ و ۲۰۲۶.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

AIAI agentsLLMsRAGSLMsclinical decision supportevidence-based medicinefine-tuninggenerative AIhealth carelanguage modelslarge language modelsnatural language processingposttraining methodsreinforcement learningretrieval-augmented generationsmall language models
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Computational Network Analysis for Defining Transcriptional Programs.

Cancer cell identity is governed by coordinated transcriptional programs that are frequently rewired during tumorigenesis. Systematic identification of cancer type-specific gene regulatory networks provides a framework for understanding oncogenic state transitions and for prioritizing candidate therapeutic targets. Here, we present a reproducible network-based workflow for reconstructing and analyzing transcriptional regulatory program…

PubMed2027

Topic-Driven Bibliometrics and Trend Intelligence for Stem Cell and Cancer Research.

The rapid growth of biomedical literature has created an urgent need for computational tools that enable researchers to systematically analyze publication trends, identify emerging research themes, and map the evolution of scientific fields. PubMed Atlas is a command-line and web-enabled workflow for topic-driven bibliometrics and trend intelligence using PubMed E-utilities. The pipeline executes PubMed queries, retrieves matching PMID…

PubMed2026

Nursing Students' Reports of Patient Safety Incidents and Reasons During Clinical Placements: A Secondary Analysis of the International Data.

Clinical placements expose nursing students to patient safety incidents and provide important opportunities for learning about safe care. This study explored how nursing students in four countries recognized and interpreted patient safety incidents encountered or witnessed during clinical practice, including perceived contributing factors. A secondary qualitative content analysis was conducted using narrative data from 1442 undergradua…

PubMed2026

Frequency of Clinically Relevant Drug-Drug Interactions Between Tyrosine Kinase Inhibitors and Proton Pump Inhibitors in Patients With Cancer Using Real-World Data.

BACKGROUND: Proton pump inhibitors (PPIs) raise stomach pH, leading to reduced bioavailability of many tyrosine kinase inhibitors (TKIs), thereby affecting treatment outcomes. To what extent this interaction occurs in clinical practice remains underexplored. OBJECTIVE: To determine the frequency of clinically relevant interactions between TKIs and PPIs in clinical practice and the duration of concomitant prescription. METHODS: A retros…