PubMed دسترسی آزاد

User Beware: Inaccuracy and Inconsistency of Large Language Models in Providing Precision Dosing Recommendations for Patients With Kidney Impairment-A Case Series.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

INTRODUCTION: Evidence-based dosing guidance for medications in critically ill patients with acute kidney injury (AKI) and receiving continuous kidney replacement therapy (CKRT) is limited. Freely available large language models (LLMs) can generate confident, human-like outputs. The accuracy and reproducibility of LLMs in providing precision drug dosing recommendations in the setting of AKI and CKRT have not been evaluated. OBJECTIVES: We sought to characterize literature concordance and internal consistency of LLM-generated dosing recommendations for cefepime and meropenem in patients with AKI and receiving CKRT. METHODS: Six investigators queried the freely available versions of six LLMs (ChatGPT, Claude, Google Gemini, Microsoft Copilot, OpenEvidence, and Perplexity) from July to September 2025 using three standardized vignettes asking for dosing recommendations and rationale to meet prespecified pharmacodynamic targets: (i) adult with AKI not on dialysis receiving cefepime, (ii) child receiving CKRT and cefepime, and (iii) toddler receiving high-effluent CKRT and meropenem. To assess inter-iteration consistency, each investigator also prompted one LLM three times for each vignette. Responses were parsed for concordance with recommendations in primary literature. LLM "reasoning" was evaluated for use of pharmacokinetic (PK) equations, citation accuracy versus confabulation, acknowledgment of uncertainty, and recommendations for therapeutic drug monitoring (TDM) for efficacy or safety. RESULTS: LLM-generated dosing recommendations varied widely. Mean concordance with literature-based recommendations was 63% (range: 28%-94%). LLMs produced variable responses to the same user upon multiple iterations, with variance in daily maintenance doses recommended ranging from 0 to 1000%. OpenEvidence universally cited relevant sources, whereas other LLMs leveraged sources inconsistently or confabulated them. Each recommended TDM, though only Claude consistently acknowledged its own uncertainty. CONCLUSION: Freely available LLMs produce highly variable and often discordant antibiotic dosing recommendations for patients with AKI and receiving CKRT. Although valuable for hypothesis generation and literature retrieval, LLM outputs should not be used in isolation for drug dosing in critically ill patients with kidney dysfunction.

نتیجه فارسی

مدل‌های زبانی بزرگ توصیه‌های دوز‌دهی آنتی‌بیوتیکی بسیار متغیری برای بیماران با آسیب کلیوی ارائه می‌دهند که اغلب با توصیه‌های علمی همخوانی ندارد. این مدل‌ها در پاسخ‌دهی به تکرار سوالات ناپایدار هستند و اغلب منابع را جعل می‌کنند. نتایج LLM نباید به تنهایی برای دوز‌دهی دارو در بیماران بحرانی استفاده شود.

  • دقت LLMها در توصیه‌های دوز‌دهی سیفپیم و مروپنم در بیماران با CKRT پایین و متغیر بود.
  • پاسخ‌های LLMها در تکرار سوالات متفاوت بود و دوزهای پیشنهادی می‌توانستند تا ۱۰۰٪ تغییر کنند.
  • OpenEvidence منابع را به طور منظم ذکر کرد، اما سایر مدل‌ها منابع را جعل کردند.
  • توصیه به نظارت دارویی (TDM) توسط همه مدل‌ها داده شد، اما فقط Claude به طور منظم عدم قطعیت خود را پذیرفت.

ترجمه فارسی چکیده

محدودیت‌های راهنمای دوز‌دهی مبتنی بر شواهد برای داروهای بیماران بحرانی با آسیب حاد کلیوی (AKI) و دریافت درمان جایگزینی کلیوی مداوم (CKRT) وجود دارد. مدل‌های زبانی بزرگ (LLM) آزاد در دسترس می‌توانند خروجی‌های با اطمینان و انسانی تولید کنند. دقت و تکرارپذیری LLMها در ارائه توصیه‌های دوز‌دهی دقیق دارو در زمینه AKI و CKRT ارزیابی نشده است. هدف این مطالعه بررسی همخوانی با ادبیات و انسجام داخلی توصیه‌های دوز‌دهی LLM برای سیفپیم و مروپنم در بیماران با AKI و دریافت CKRT بود. شش محقق از نسخه‌های رایگان شش مدل LLM (ChatGPT، Claude، Google Gemini، Microsoft Copilot، OpenEvidence و Perplexity) از ژوئیه تا سپتامبر ۲۰۲۵ استفاده کردند. میانگین همخوانی با توصیه‌های مبتنی بر ادبیات ۶۳٪ بود. LLMها پاسخ‌های متغیری تولید کردند و توصیه‌های دوز نگهدارنده روزانه متغیری از ۰ تا ۱۰۰٪ ارائه دادند. OpenEvidence منابع مرتبط را همیشه ذکر کرد، در حالی که سایر LLMها منابع را به طور ناپایدار استفاده یا جعل کردند. هر مدل توصیه به نظارت دارویی (TDM) کرد، هرچند فقط Claude عدم قطعیت خود را به طور منظم پذیرفت.

روش پژوهش

شش مدل LLM از ژوئیه تا سپتامبر ۲۰۲۵ بر اساس سه سناریوی استاندارد مورد بررسی قرار گرفتند. برای ارزیابی انسجام، هر محقق یک مدل را سه بار برای هر سناریو پرسید. پاسخ‌ها برای همخوانی با توصیه‌های ادبیات اولیه تحلیل شدند.

محدودیت‌ها

محدودیت‌های این مطالعه شامل استفاده از نسخه‌های رایگان مدل‌ها و تمرکز بر دو آنتی‌بیوتیک خاص است. همچنین، نتایج ممکن است به دلیل ماهیت متغیر LLMها در زمان‌های مختلف اعمال نشود.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
بیماران با آسیب حاد کلیوی (AKI) و دریافت درمان جایگزینی کلیوی مداوم (CKRT)
مداخله/مواجهه
مدل‌های زبانی بزرگ (LLM) برای توصیه به دوز‌دهی سیفپیم و مروپنم
مقایسه
توصیه‌های مبتنی بر ادبیات اولیه
حجم نمونه
شش مدل LLM (ChatGPT، Claude، Google Gemini، Microsoft Copilot، OpenEvidence، Perplexity)

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

acute kidney injuryartificial intelligencebeta lactamscontinuous kidney replacement therapylarge language modelspharmacokinetics
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Age-Specific Associations of Polygenic Risk Scores With Advanced Fibrosis and Histological Activity in Biopsy-Proven MASLD.

BACKGROUND AND AIMS: The genetic susceptibility to MASLD histological severity remains incompletely defined. We assessed the associations between recently developed genome-wide association studies-derived polygenic risk scores (PRSs) and advanced fibrosis in MASLD. We also evaluated PRSs' associations with histological activity and selected PRSs' predictive ability for advanced fibrosis. METHODS: We analysed 2149 adults and 900 childre…

PubMed2026

Prehospital Births in the Region of Southern Denmark-A Cohort From 2016 to 2024.

BACKGROUND: Unplanned births outside hospitals involve higher risks of complications for mothers and newborns and require special obstetric or pediatric skills. However, ambulance calls for childbirth are rare, making it difficult for ambulance staff to maintain their skills. In Europe, unplanned out-of-hospital deliveries constitute 0.10% to 0.61% of all births. The exact number of prehospital births in the Region of Southern Denmark …

PubMed2026

EACTS Expert Consensus Document on aortic valve repair and valve-sparing aortic root replacement procedures.

Aortic valve repair is a complex and evolving field that requires multidisciplinary teamwork among heart specialists to diagnose and treat various causes and mechanisms of aortic regurgitation and proximal thoracic aortic aneurysms across all age groups. From a surgical standpoint, aortic valve repair includes all procedures aimed at restoring or maintaining native valve function. This European Association for Cardio-Thoracic Surgery (…

PubMed2026

EACTS Expert Consensus Document on Civil Aviation Following Cardiothoracic Surgery and Transcatheter Procedures.

Civil aviation medicine and cardiothoracic surgery intersect at a point where individual clinical outcomes and public safety are tightly coupled. Professional aircrew and other safety-critical aviation personnel must meet legally defined medical standards that are designed around an engineering approach to risk, including a low annual tolerance for sudden incapacitation, whereas passengers and cabin crew are exposed to the physiologica…