PubMed دسترسی آزاد

Evaluating Retrieval-Augmented Large Language Models on Anesthesiology Board-Style Questions: Benchmark Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Open-source, mid-scale large language models (LLMs) have emerged as scalable, privacy-preserving alternatives to ultra-large foundation models (eg, GPT-4) in health care systems. Techniques such as retrieval-augmented generation (RAG) enable sub-100-billion-parameter models to address highly specialized medical domains such as anesthesiology. However, studies evaluating RAG architectures on complex medical examinations remain scarce, highlighting the need for rigorous benchmarking to bridge the gap between raw parametric knowledge and clinically relevant application. OBJECTIVE: This study aimed to systematically evaluate RAG pipelines for answering anesthesiology board-style questions, quantify the effects of key design choices including hyperparameter settings, embedding models, source complexity, and chunking strategies, and compare the performance of reasoning-oriented models with that of conventional LLMs. METHODS: We conducted large-scale benchmarking using American Board of Anesthesiology-style multiple-choice questions to compare multiple RAG-enabled configurations with matched standalone LLM baselines. Configurations were first optimized on a 46-item diagnostic set and then validated on a 350-item corpus. Additional experiments on three 100-question subsets derived from the 350-item corpus were used to assess the effects of source selection, source complexity, information density, and chunking strategy on answer accuracy. Models including Llama-3-8B-Instruct, Llama-3.1-8B-Instruct, Llama-3.2-3B-Instruct, Llama-3.3-70B-Instruct, Qwen2.5-7B and Qwen2.5-72B, and Qwen3-8B and Qwen3-32B reasoning models were evaluated under this framework. Self-reflective RAG (self-RAG) with adaptive retrieval techniques was also implemented and evaluated. Cochran Q and McNemar tests were used to assess performance differences across configurations and model pairs. RESULTS: The RAG framework increased the number of correct answers. System stability peaked under highly deterministic sampling configurations (temperature=0.1, top-p [nucleus sampling]=0.1). High-capacity general-text embeddings and applying context-preserving semantic chunking further improved accuracy. Standard RAG provided only modest gains over nonaugmented baselines, improving accuracy from 50.29% to 56.57%, and self-RAG yielded similarly limited gains of up to 4.85 percentage points. Overall, the Qwen family outperformed the Llama series. The 32-billion-parameter reasoning model Qwen-3-32B achieved an 89% correct ratio under complex distractor-heavy retrieval conditions and up to 96% with direct context, significantly outperforming the much larger 72-billion-parameter conventional model Qwen-2.5-72B-Instruct (84%). Smaller reasoning models also showed greater robustness to noise or suboptimal retrieved documents than larger conventional LLMs. Within the Llama family, increasing parameter size to 70 billion did not produce proportional performance gains on this benchmark. CONCLUSIONS: RAG-based LLM systems improved performance on anesthesiology board-style questions, but gains depended strongly on retrieval design. Careful optimization of retrieval settings, embeddings, and chunking strategies improved robustness and answer accuracy. Reasoning-oriented models demonstrated that multistep reasoning can, in some settings, compensate for larger parameter scale. These findings provide a methodological foundation for developing locally deployable LLM systems for anesthesiology education within structured examination settings.

نتیجه فارسی

این مطالعه به ارزیابی عملکرد مدل‌های زبانی بزرگ (LLM) با تکنیک بازیابی (RAG) بر روی سوالات امتحانی تخصصی جراحی پرداخته است. نتایج نشان می‌دهد که RAG می‌تواند دقت پاسخ‌دهی را افزایش دهد، اما بهینه‌سازی تنظیمات بازیابی و مدل‌های تلفیقی کلیدی است. مدل‌های استدلالی کوچک‌تر نسبت به مدل‌های متداول بزرگ‌تر، مقاومت بیشتری در برابر نویز نشان دادند. خانواده مدل Qwen عملکرد بهتری از سری Llama داشت.

  • RAG می‌تواند دقت پاسخ‌دهی به سوالات امتحانی تخصصی جراحی را افزایش دهد.
  • بهینه‌سازی تنظیمات بازیابی، مدل‌های تلفیقی و استراتژی‌های برش برای بهبود عملکرد ضروری است.
  • مدل‌های استدلالی کوچک‌تر نسبت به مدل‌های متداول بزرگ‌تر، مقاومت بیشتری در برابر نویز نشان دادند.
  • خانواده مدل Qwen عملکرد بهتری از سری Llama در این معیار داشت.
  • افزایش اندازه پارامتر در مدل‌های Llama به تنهایی سود عملکرد متناسبی ایجاد نکرد.

ترجمه فارسی چکیده

پس‌زمینه: مدل‌های زبانی بزرگ (LLM) متن‌باز با مقیاس متوسط به عنوان جایگزین‌های مقیاس‌پذیر و حفظ حریم خصوصی برای مدل‌های بنیادی فوق‌بزرگ (مانند GPT-4) در سیستم‌های بهداشتی ظهور کرده‌اند. تکنیک‌هایی مانند تولید تقویت‌شده با بازیابی (RAG) به مدل‌هایی با زیر ۱۰۰ میلیارد پارامتر اجازه می‌دهند تا به حوزه‌های تخصصی پزشکی مانند جراحی دسترسی پیدا کنند. هدف: این مطالعه به منظور ارزیابی سیستماتیک پیکربندی‌های RAG برای پاسخ‌دهی به سوالات امتحانی تخصصی جراحی، کمی‌سازی اثر انتخاب‌های طراحی کلیدی شامل تنظیمات فراپارامتر، مدل‌های تلفیقی و استراتژی‌های برش، و مقایسه عملکرد مدل‌های مبتنی بر استدلال با LLM‌های متداول انجام شد. روش‌ها: ما از سوالات چندگزینه‌ای سبک انجمن جراحی آمریکا برای مقایسه پیکربندی‌های مختلف RAG با خط‌های پایه LLM مستقل استفاده کردیم. پیکربندی‌ها ابتدا روی مجموعه تشخیصی ۴۶ مورد بهینه‌سازی و سپس روی مجموعه ۳۵۰ مورد تأیید شدند. آزمایش‌های اضافی بر روی زیرمجموعه‌های ۱۰۰ سؤالی برای ارزیابی اثر انتخاب منبع، پیچیدگی منبع، چگالی اطلاعات و استراتژی برش بر دقت پاسخ انجام شد. مدل‌هایی شامل Llama-3-8B-Instruct، Qwen2.5-7B و Qwen3-8B و همچنین مدل‌های استدلالی Qwen-3-32B و Qwen-3-8B ارزیابی شدند. همچنین RAG خودبازتابی (self-RAG) با تکنیک‌های بازیابی تطبیقی پیاده‌سازی و ارزیابی شد. از آزمون‌های Cochran Q و McNemar برای ارزیابی تفاوت عملکرد بین پیکربندی‌ها و جفت‌های مدل استفاده شد. نتایج: چارچوب RAG تعداد پاسخ‌های صحیح را افزایش داد. پایداری سیستم در پیکربندی‌های نمونه‌گیری بسیار قطعی (دمای ۰.۱، top-p=۰.۱) اوج گرفت. استفاده از تلفیقی‌های متنی عمومی با ظرفیت بالا و اعمال برش معنایی حفظ‌کننده زمینه دقت را بهبود بخشید. RAG استاندارد تنها سود محدودی نسبت به خط‌های پایه بدون تقویت ایجاد کرد (از ۵۰.۲۹٪ به ۵۶.۵۷٪) و self-RAG نیز سود محدودی تا ۴.۸۵ درصد ایجاد کرد. در مجموع، خانواده Qwen عملکرد بهتری از سری Llama داشت. مدل استدلالی ۳۲ میلیارد پارامتری Qwen-3-32B در شرایط بازیابی با موانع پیچیده ۸۹٪ دقت و تا ۹۶٪ با زمینه مستقیم داشت که عملکرد بهتری نسبت به مدل متداول بسیار بزرگ‌تر Qwen-2.5-72B-Instruct (۸۴٪) نشان داد. مدل‌های استدلالی کوچک‌تر نیز مقاومت بیشتری نسبت به نویز یا مستندات بازیابی نامناسب از مدل‌های LLM متداول بزرگ‌تر نشان دادند. در خانواده Llama، افزایش اندازه پارامتر به ۷۰ میلیارد در این معیار سود عملکرد متناسبی ایجاد نکرد. نتیجه‌گیری: سیستم‌های LLM مبتنی بر RAG عملکرد در سوالات امتحانی تخصصی جراحی را بهبود بخشیدند، اما سود به شدت به طراحی بازیابی وابسته بود. بهینه‌سازی دقیق تنظیمات بازیابی، تلفیقی‌ها و استراتژی‌های برش، پایداری و دقت پاسخ را بهبود داد. مدل‌های مبتنی بر استدلال نشان دادند که استدلال چندمرحله‌ای در برخی تنظیمات می‌تواند مقیاس بزرگ‌تر پارامترها را جبران کند.

روش پژوهش

این مطالعه از سوالات چندگزینه‌ای سبک انجمن جراحی آمریکا برای مقایسه پیکربندی‌های RAG با خط‌های پایه LLM مستقل استفاده کرد. پیکربندی‌ها ابتدا روی مجموعه تشخیصی ۴۶ مورد بهینه‌سازی و سپس روی مجموعه ۳۵۰ مورد تأیید شدند. آزمایش‌های اضافی بر روی زیرمجموعه‌های ۱۰۰ سؤالی برای ارزیابی اثر انتخاب منبع، پیچیدگی منبع، چگالی اطلاعات و استراتژی برش بر دقت پاسخ انجام شد. از آزمون‌های Cochran Q و McNemar برای ارزیابی تفاوت عملکرد بین پیکربندی‌ها و جفت‌های مدل استفاده شد.

محدودیت‌ها

محدودیت‌های این مطالعه شامل تمرکز بر سوالات امتحانی تخصصی جراحی و عدم بررسی سایر حوزه‌های پزشکی است. همچنین، نتایج ممکن است به دلیل ماهیت نمونه‌گیری خاص، به سایر محیط‌ها تعمیم‌پذیر نباشد.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
سوالات امتحانی تخصصی جراحی (سبک انجمن جراحی آمریکا)
مداخله/مواجهه
مدل‌های زبانی بزرگ (LLM) با تکنیک بازیابی (RAG) و مدل‌های استدلالی (مانند Qwen-3-32B و Llama-3-8B-Instruct)
مقایسه
مدل‌های زبانی بزرگ (LLM) متداول (بدون RAG) و خط‌های پایه LLM مستقل
حجم نمونه
مجموعه تشخیصی ۴۶ مورد و مجموعه ۳۵۰ مورد برای تأیید، زیرمجموعه‌های ۱۰۰ سؤالی برای آزمایش‌های اضافی

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

anesthesiology board examinationembedding modellarge language modelreasoning modelretrieval-augmented generation
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and enviro…

PubMed2026

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated respons…

PubMed2026

Liability and Standard of Care in AI-Driven Psychiatric Practice: European Viewpoint.

AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management…

PubMed2026

Routine laboratory panels classify internal medicine ICD-10 code groups: comparison with frontier large language models and laboratory-only specialist assessment.

INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from …