PubMed چکیده/رکورد

"Small" Large Language Models in the Hospital: Evaluation Study on Real-World Data in a Resource-Constrained Setting.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Large language models (LLMs) are increasingly being deployed in health care, but their use and deployment in many real-world hospital environments pose significant challenges and concerns. In particular, state-of-the-art commercial models store or process data externally, which is often in conflict with ensuring patient data protection. At the same time, using LLMs locally is limited by the lack of available computing infrastructure. Small open-source LLMs that do not require substantial computing resources could offer a practical way to resolve these tensions, but their medical utility in real-world local contexts, especially in non-English languages, has not been sufficiently evaluated. OBJECTIVE: This study aimed to evaluate the feasibility of small, locally deployable open-source LLMs for clinically relevant tasks in a resource-constrained hospital setting and to propose a reproducible framework for institution-specific evaluation before deployment. METHODS: We evaluated 6 open-source LLMs ranging from 8B to 24B parameters (from the Mistral, Phi4, Falcon3, Llama3.1, and Meditron3 families) in a zero-shot setting across 7 tasks covering 4 clinical use cases: information extraction, medical text translation, text generation, and clinical decision support. We used deidentified French clinical data from a Swiss tertiary hospital, including discharge letters, clinical notes, and structured electronic health records. Performance was assessed using task-specific metrics, such as precision, recall, F1-score, embedding-based semantic similarity, recall-oriented understudy for gisting evaluation (ROUGE) score, readability indices, and human review by clinicians. RESULTS: Model performance varied substantially between tasks. In the simplest retrieval task (needle-in-the-haystack), several models performed strongly, with Llama3.1 achieving an F1-score of 99.81% and Mistral-small achieving 99.71%. In contrast, performance was poor in more complex tasks. For detecting protected health information, the best-performing LLMs achieved only modest overall macro-F1-scores (0.33-0.34), substantially below a fine-tuned Robustly Optimized BERT Pretraining Approach (RoBERTa) baseline (0.94). In the task of extracting immune-related adverse events from discharge notes, the highest overall macro-F1-score was 0.35 with Phi4. For medical text translation, Phi4 ranked highest in embedding-based evaluation, whereas Meditron3-Phi4 performed the worst, with clinician reviews identifying hallucinations in 55% of its outputs. In the task of summarizing discharge letters, quality was low across all models, with the best penalized ROUGE score reaching only 0.169 with Llama3.1. In the tasks of generating patient-friendly discharge note summaries and clinical decision support, clinician ratings generally ranged from dissatisfied to neutral, and no model achieved consistently satisfactory performance. CONCLUSIONS: Small open-source LLMs appear feasible for simple retrieval-oriented tasks in local hospital deployments but are currently inadequate for more complex applications, such as clinical decision support, deidentification, extraction of adverse events, and medical summarization. These findings highlight the importance of locally grounded evaluation tailored to specific use cases and the need for robust institutional evaluation frameworks to ensure safe and reliable deployment.

نتیجه فارسی

این مطالعه ۶ مدل زبانی بزرگ متن‌باز کوچک را در یک بیمارستان سوئیس ارزیابی کرد. نتایج نشان داد که این مدل‌ها برای وظایف ساده retrieval عملکرد خوبی دارند، اما برای وظایف پیچیده‌تر مانند پشتیبانی از تصمیم بالینی و خلاصه‌سازی، عملکرد مطلوبی ندارند. هیچ مدل‌ای در تمام وظایف مورد بررسی عملکرد consistently satisfactory نداشت.

  • ارزیابی ۶ مدل متن‌باز LLM با پارامترهای ۸ تا ۲۴ میلیارد در یک بیمارستان واقعی.
  • عملکرد مدل‌ها در وظایف retrieval (needle-in-the-haystack) بالا بود (F1 تا 99.81%).
  • عملکرد در وظایف پیچیده‌تر مانند تشخیص اطلاعات محافظت شده و خلاصه‌سازی ضعیف بود.
  • هیچ مدل‌ای در وظایف تولید متن و پشتیبانی از تصمیم بالینی عملکرد consistently satisfactory نداشت.
  • نتیجه‌گیری: مدل‌های کوچک برای کاربردهای ساده مناسب‌اند، اما برای کاربردهای بالینی پیچیده کافی نیستند.

ترجمه فارسی چکیده

مدل‌های زبانی بزرگ (LLM) در حال افزایش در مراقبت‌های بهداشتی استفاده می‌شوند، اما استفاده و پیاده‌سازی آن‌ها در محیط‌های واقعی بیمارستانی چالش‌های و نگرانی‌های جدی ایجاد می‌کند. در حالی که مدل‌های تجاری پیشرفته داده‌ها را خارج از بیمارستان ذخیره یا پردازش می‌کنند که اغلب با تضمین حفاظت از داده‌های بیمار در تضاد است، استفاده از LLM به صورت محلی محدودیت‌های زیرساخت محاسباتی موجود را دارد. مدل‌های کوچک متن‌باز LLM که منابع محاسباتی قابل توجهی نیاز ندارند، می‌توانند راهکار عملی برای حل این تنش‌ها باشند، اما کاربرد پزشکی آن‌ها در محیط‌های محلی واقعی، به‌ویژه در زبان‌های غیرانگلیسی، به‌طور کافی ارزیابی نشده است. هدف این مطالعه ارزیابی قابلیت‌های مدل‌های کوچک متن‌باز LLM قابل پیاده‌سازی محلی برای وظایف clinically relevant در محیط بیمارستانی با منابع محدود و ارائه یک چارچوب قابل تکرار برای ارزیابی خاص institution قبل از پیاده‌سازی بود. ما ۶ مدل متن‌باز LLM با پارامترهای ۸ تا ۲۴ میلیارد (از خانواده‌های Mistral، Phi4، Falcon3، Llama3.1 و Meditron3) را در یک مح_setting zero-shot در ۷ وظیفه پوشش ۴ مورد استفاده بالینی ارزیابی کردیم: استخراج اطلاعات، ترجمه متن پزشکی، تولید متن و پشتیبانی از تصمیم بالینی. ما از داده‌های بالینی فرانسوی بدون شناسایی (deidentified) از یک بیمارستان tertiary سوئیس، شامل نامه‌های ترخیص، یادداشت‌های بالینی و سوابق الکترونیک سلامت ساختاریافته استفاده کردیم. عملکرد با استفاده از معیارهای خاص وظیفه، مانند دقت،recall، امتیاز F1، شباهت معنایی مبتنی بر embedding، امتیاز recall-oriented understudy برای gisting evaluation (ROUGE)، شاخص‌های خوانایی و بازبینی انسانی توسط پزشکان ارزیابی شد. عملکرد مدل‌ها بین وظایف به‌طور قابل توجهی متفاوت بود. در وظیفه ساده retrieval (needle-in-the-haystack)، چندین مدل عملکرد قوی داشتند، با Llama3.1 به دست آوردن امتیاز F1 99.81% و Mistral-small به دست آوردن 99.71%. در مقابل، عملکرد در وظایف پیچیده‌تر ضعیف بود. برای تشخیص اطلاعات بهداشتی محافظت شده، بهترین LLMها فقط امتیازهای macro-F1 کلی محدود (0.33-0.34) به دست آوردند، که به‌طور قابل توجهی زیر پایه‌ای RoBERTa (0.94) بود. در وظیفه استخراج adverse events مرتبط با ایمنی از یادداشت‌های ترخیص، بالاترین امتیاز macro-F1 کلی 0.35 با Phi4 بود. برای ترجمه متن پزشکی، Phi4 در ارزیابی مبتنی بر embedding رتبه اول را داشت، در حالی که Meditron3-Phi4 بدترین عملکرد را داشت، با بازبینی پزشکان که hallucinations را در 55% خروجی‌های آن شناسایی کرد. در وظیفه خلاصه‌سازی نامه‌های ترخیص، کیفیت در تمام مدل‌ها پایین بود، با بهترین امتیاز ROUGE penalized تنها 0.169 با Llama3.1. در وظایف تولید خلاصه‌های نامه‌های ترخیص بیمارپسند و پشتیبانی از تصمیم بالینی، امتیازدهی پزشکان معمولاً از ناراضی تا خنثی ranged بود و هیچ مدل عملکرد مطلوب consistently به دست نیاورد. مدل‌های کوچک متن‌باز LLM برای وظایف ساده retrieval-oriented در پیاده‌سازی‌های محلی بیمارستان به نظر می‌رسد feasible هستند اما در حال حاضر برای کاربردهای پیچیده‌تر، مانند پشتیبانی از تصمیم بالینی، deidentification، استخراج adverse events و خلاصه‌سازی پزشکی، inadequate هستند. این یافته‌ها اهمیت ارزیابی محلی grounded tailored به موارد استفاده خاص و نیاز به چارچوب‌های ارزیابی institution robust برای تضمین پیاده‌سازی safe و reliable را برجسته می‌کنند.

روش پژوهش

این مطالعه یک ارزیابی zero-shot از ۶ مدل متن‌باز LLM (Mistral, Phi4, Falcon3, Llama3.1, Meditron3) در ۷ وظیفه بالینی مختلف انجام داد. داده‌ها از یک بیمارستان tertiary سوئیس (بدون شناسایی) شامل نامه‌های ترخیص و یادداشت‌های بالینی بودند. عملکرد با معیارهای متعدد مانند F1-score، ROUGE و بازبینی پزشکان ارزیابی شد.

محدودیت‌ها

محدودیت‌های متن کامل گزارش نشده‌اند. این مطالعه فقط در یک بیمارستان انجام شد و نتایج ممکن است به سایر محیط‌ها تعمیم‌پذیر نباشد.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
بیماران در یک بیمارستان tertiary سوئیس (داده‌های بالینی فرانسوی بدون شناسایی).
مداخله/مواجهه
استفاده از ۶ مدل زبانی بزرگ متن‌باز کوچک (8B تا 24B پارامتر) برای وظایف بالینی.
مقایسه
هیچ comparator مشخصی گزارش نشده است.
حجم نمونه
تعداد نمونه مشخصی گزارش نشده است.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

French medical textLLMclinical NLPevaluation frameworkhealth carehospital deploymentlarge language modelnatural language processingresource-constrained settingssmall language models
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …