PubMed دسترسی آزاد

Multicenter Evaluation of Large Language Models Versus Hepatologists for Prognostic Prediction in Drug-Induced Liver Injury.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND & AIMS: Drug-induced liver injury (DILI) would progress to chronicity or death. Large language models (LLMs) may enhance clinical decision-making, yet their utility relative to physicians in DILI remains unclear. Therefore, we evaluated their performance in predicting DILI outcomes. METHODS: We enrolled 943 DILI patients from three centers as internal and external cohorts. Based on 12-month follow-up, outcomes were classified as recovery, 6/12-month chronicity, and death. LLMs (Gemini-2.5 Pro, GPT-5.1, DeepSeek-3.2), hepatologists (Junior, middle, senior), and models (Hy's Law, nHy's Law, MELD Score) estimated probabilities of outcomes. LLM-Rules (VOTE, OR, AND) were applied to enhance stability. Model performance was assessed. RESULTS: For 6-month chronicity, the senior achieved highest AUROC (0.61) with an accuracy of 70%. Gemini-2.5 Pro and GPT-5.1 yielded AUROCs of 0.60 and 0.59, respectively, outperforming junior and middle hepatologists. Gemini-2.5 Pro demonstrated strongest agreement with senior (κ = 0.43). LLMs all exhibited lower accuracy and specificity than hepatologists. A similar result was observed in 12-month chronicity. For overall mortality, the senior achieved highest AUROC (0.87) with an accuracy of 83%. Gemini-2.5 Pro and GPT-5.1 achieved AUROCs of 0.86, outperforming junior hepatologist, Hy's Law, and nHy's Law. GPT-5.1 achieved strongest agreement with senior (κ = 0.25). LLM-Rules demonstrated stability for predicting outcomes across cohorts. OR and AND rules improved sensitivity and specificity, respectively. CONCLUSIONS: GPT-5.1 and Gemini-2.5 Pro showed AUROCs approaching senior hepatologists for DILI outcomes with limited accuracy and specificity. LLM-Rules demonstrated stable performance across cohorts with improved sensitivity or specificity, supporting the potential of multi-LLM approaches as clinician-supervised complementary tools.

نتیجه فارسی

این مطالعه عملکرد مدل‌های زبانی بزرگ (LLM) را در پیش‌بینی نتایج آسیب کبد ناشی از دارو (DILI) در مقایسه با متخصصان کبد ارزیابی کرد. LLMها عملکردی نزدیک به متخصصان ارشد در پیش‌بینی مرگ کلی و مزمنی داشتند، اما دقت و ویژگی‌های خاص کمتری نسبت به متخصصان داشتند. استفاده از قوانین ترکیبی LLM-Rules پایداری را در پیش‌بینی نتایج بهبود بخشید.

  • LLMها عملکردی نزدیک به متخصصان ارشد در پیش‌بینی مرگ کلی و مزمنی DILI داشتند.
  • LLMها دقت و ویژگی‌های خاص کمتری نسبت به متخصصان کبد داشتند.
  • قوانین LLM-Rules (OR و AND) پایداری را در پیش‌بینی نتایج بهبود بخشیدند.
  • این مطالعه نشان داد که LLMها می‌توانند ابزارهای تکمیلی تحت نظارت بالینی باشند.

ترجمه فارسی چکیده

پیشرفت آسیب کبد ناشی از دارو (DILI) می‌تواند به مزمنی یا مرگ منجر شود. مدل‌های زبانی بزرگ (LLM) ممکن است تصمیمات بالینی را بهبود بخشند، اما کارایی آن‌ها نسبت به پزشکان در DILI هنوز مشخص نیست. بنابراین، عملکرد آن‌ها در پیش‌بینی نتایج DILI ارزیابی شد. ۹۴۳ بیمار DILI از سه مرکز به عنوان گروه‌های داخلی و خارجی ثبت شدند. بر اساس پیگیری ۱۲ ماهه، نتایج به عنوان بهبودی، مزمنی ۶/۱۲ ماهه و مرگ طبقه‌بندی شدند. LLMها (Gemini-2.5 Pro، GPT-5.1، DeepSeek-3.2)، متخصصان کبد (جونیور، میانی، ارشد) و مدل‌ها (قانون Hy، nHy's Law، امتیاز MELD) احتمالات نتایج را تخمین زدند. LLM-Rules (VOTE، OR، AND) برای افزایش پایداری اعمال شد. عملکرد مدل ارزیابی شد. برای مزمنی ۶ ماهه، ارشد بالاترین AUROC (۰.۶۱) را با دقت ۷۰٪ به دست آورد. Gemini-2.5 Pro و GPT-5.1 به ترتیب AUROCهای ۰.۶۰ و ۰.۵۹ را ارائه دادند که عملکرد بهتری نسبت به متخصصان جونیور و میانی داشتند. Gemini-2.5 Pro بیشترین توافق را با ارشد نشان داد (κ = ۰.۴۳). LLMها همگی دقت و ویژگی‌های خاص کمتری نسبت به متخصصان کبد داشتند. نتیجه مشابهی در مزمنی ۱۲ ماهه مشاهده شد. برای مرگ کلی، ارشد بالاترین AUROC (۰.۸۷) را با دقت ۸۳٪ به دست آورد. Gemini-2.5 Pro و GPT-5.1 AUROCهای ۰.۸۶ را به دست آوردند که عملکرد بهتری نسبت به متخصص جونیور، قانون Hy و nHy's Law داشتند. GPT-5.1 بیشترین توافق را با ارشد نشان داد (κ = ۰.۲۵). LLM-Rules پایداری را برای پیش‌بینی نتایج در طول گروه‌ها نشان داد. قوانین OR و AND به ترتیب حساسیت و ویژگی‌های خاص را بهبود بخشیدند. نتیجه‌گیری: GPT-5.1 و Gemini-2.5 Pro AUROCs نزدیک به متخصصان کبد ارشد را برای نتایج DILI با دقت و ویژگی‌های خاص محدود نشان دادند. LLM-Rules عملکرد پایداری را در طول گروه‌ها نشان داد و حساسیت یا ویژگی‌های خاص را بهبود بخشید، که پتانسیل رویکردهای چند-LLM را به عنوان ابزارهای تکمیلی تحت نظارت بالینی پشتیبانی می‌کند.

روش پژوهش

این مطالعه یک ارزیابی چندمرکزه بود که ۹۴۳ بیمار DILI از سه مرکز را شامل می‌شد. نتایج بر اساس پیگیری ۱۲ ماهه طبقه‌بندی شدند. عملکرد مدل‌ها با استفاده از AUROC و دقت ارزیابی شد.

محدودیت‌ها

محدودیت‌های این مطالعه شامل عدم گزارش دقیق بودجه و منابع بودجه است. همچنین، محدودیت‌های مربوط به دقت و ویژگی‌های خاص LLMها در مقایسه با متخصصان کبد ذکر شده است.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
۹۴۳ بیمار DILI از سه مرکز.
مداخله/مواجهه
مدل‌های زبانی بزرگ (Gemini-2.5 Pro، GPT-5.1، DeepSeek-3.2) و قوانین LLM-Rules (VOTE، OR، AND).
مقایسه
متخصصان کبد (جونیور، میانی، ارشد) و مدل‌های استاندارد (قانون Hy، nHy's Law، امتیاز MELD).
حجم نمونه
۹۴۳ بیمار DILI.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

chemical and drug induced liver injurychronic diseaseclinical outcomelarge language modelsliver failureprognosis
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and enviro…

PubMed2026

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated respons…

PubMed2026

Liability and Standard of Care in AI-Driven Psychiatric Practice: European Viewpoint.

AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management…

PubMed2026

Routine laboratory panels classify internal medicine ICD-10 code groups: comparison with frontier large language models and laboratory-only specialist assessment.

INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from …