PubMed چکیده/رکورد

Performance of large language models on narrow therapeutic index drug monitoring: Implications for clinical pharmacy practice.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Large language models are increasingly investigated as clinical decision support tools, but their reliability for therapeutic drug monitoring interpretation remains poorly explored. Phenytoin and digoxin, two narrow therapeutic index drugs, represent clinically challenging test cases. OBJECTIVES: To evaluate three current-generation LLMs on phenytoin and digoxin TDM interpretation across seven clinical reasoning domains, identify domain-specific strengths and limitations, and characterize failure patterns using a qualitative error taxonomy. METHODS: Thirty structured clinical vignettes were submitted to Claude Sonnet 4.6, ChatGPT 5.5, and Gemini 3.1 Pro. The blinded responses were independently scored by two raters using a seven-domain rubric. Between-model differences were assessed using Friedman and post-hoc Wilcoxon signed-rank tests, with Holm adjustment across domain-level tests and Bonferroni correction for pairwise comparisons. All 90 responses underwent error taxonomy analysis. Second independent responses were subsequently generated and scored using the same procedure to assess agreement across repeated generations. RESULTS: Claude achieved the highest mean performance (94.9% ± 8.3%), significantly outperforming ChatGPT (84.5% ± 12.9%, p<0.001) and Gemini (83.4% ± 11.9%, p=0.001). All models showed high performance on level interpretation and toxicity assessment, but significant between-model differences emerged on pharmacokinetic reasoning, management, monitoring, and uncertainty acknowledgment (all Holm-adjusted p≤0.02). Overconfident reasoning was the most common error category (45.4% of total error appearances). Repeat-generation analysis showed high test-retest agreement across generations (ICC=0.90; 95% CI, 0.82-0.94). CONCLUSION: The evaluated LLMs performed strongly in recognizing TDM problems in structured vignettes but showed variability in reasoning-intensive tasks, with overconfident reasoning representing a potential patient safety concern. Pharmacist oversight remains essential for TDM tasks.

نتیجه فارسی

این مطالعه عملکرد سه مدل زبانی بزرگ (Claude Sonnet 4.6، ChatGPT 5.5 و Gemini 3.1 Pro) را در تفسیر نظارت بر داروهای با پنجره درمانی باریک (فنیوتین و دیگوکسین) بررسی کرد. Claude بالاترین عملکرد را داشت، اما تمام مدل‌ها در استدلال‌های پیچیده متفاوت بودند. خطای اصلی «استدلال بیش از حد اطمینان» بود. تکرار پاسخ‌ها نشان‌دهنده همخوانی بالا بود.

  • Claude Sonnet 4.6 بالاترین عملکرد را داشت.
  • استدلال بیش از حد اطمینان رایج‌ترین خطا بود.
  • نظارت داروساز برای وظایف TDM ضروری است.
  • تکرار پاسخ‌ها همخوانی بالایی نشان داد.
  • مدل‌ها در تفسیر سطح و ارزیابی سمیت عملکرد خوبی داشتند.

ترجمه فارسی چکیده

پس‌زمینه: مدل‌های زبانی بزرگ به عنوان ابزارهای پشتیبانی از تصمیم‌گیری بالینی مورد بررسی قرار می‌گیرند، اما قابلیت اطمینان آن‌ها در تفسیر نظارت بر داروهای با پنجره درمانی باریک هنوز به‌طور کافی بررسی نشده است. فنیوتین و دیگوکسین، دو دارو با پنجره درمانی باریک، نمونه‌های چالش‌برانگیز بالینی هستند. اهداف: ارزیابی سه مدل LLM نسل فعلی بر روی تفسیر TDM فنیوتین و دیگوکسین در هفت حوزه استدلال بالینی، شناسایی نقاط قوت و محدودیت‌های مختص هر حوزه، و توصیف الگوهای شکست با استفاده از یک طبقه‌بندی کیفی خطا. روش‌ها: سیاره سناریوی بالینی ساختاریافته به Claude Sonnet 4.6، ChatGPT 5.5 و Gemini 3.1 Pro ارسال شد. پاسخ‌های کور توسط دو رتبه‌دهنده مستقل با استفاده از یک چک‌لیست هفت‌حوزه‌ای امتیازدهی شدند. تفاوت‌های بین مدل‌ها با استفاده از آزمون‌های فریدمن و Wilcoxon امضا‌شده پس‌ازآزمون ارزیابی شد. نتایج: Claude بالاترین میانگین عملکرد (94.9% ± 8.3%) را داشت و به‌طور معناداری از ChatGPT (84.5% ± 12.9%) و Gemini (83.4% ± 11.9%) برتری داشت. تمام مدل‌ها عملکرد بالایی در تفسیر سطح و ارزیابی سمیت نشان دادند، اما تفاوت‌های معناداری بین مدل‌ها در استدلال فارماکوکینتیک، مدیریت، نظارت و پذیرش عدم قطعیت ظاهر شد. نتیجه‌گیری: مدل‌های LLM ارزیابی شده در شناسایی مشکلات TDM در سناریوهای ساختاریافته عملکرد قوی‌ای داشتند اما در وظایف پرهزینه از نظر استدلال متغیر بودند و استدلال بیش از حد اطمینان یک نگرانی بالقوه ایمنی بیمار را نشان داد. نظارت داروساز همچنان برای وظایف TDM ضروری است.

روش پژوهش

سیاره سناریوی بالینی ساختاریافته به سه مدل LLM ارسال شد. پاسخ‌ها با یک چک‌لیست هفت‌حوزه‌ای توسط دو رتبه‌دهنده کور امتیازدهی شدند. تفاوت‌ها با آزمون‌های آماری ارزیابی شدند و پاسخ‌ها برای خطا تحلیل شدند.

محدودیت‌ها

مطالعه فقط از سناریوهای ساختاریافته استفاده کرد و نتایج ممکن است در سناریوهای واقعی متفاوت باشد. تمرکز بر دو دارو خاص بود.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
فنیوتین و دیگوکسین (داروهای با پنجره درمانی باریک)
مداخله/مواجهه
Claude Sonnet 4.6، ChatGPT 5.5، Gemini 3.1 Pro
مقایسه
همدیگر (برای مقایسه عملکرد)
حجم نمونه
۳۰ سناریوی بالینی (۳۰ پاسخ برای هر مدل)

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

Clinical pharmacyDigoxinLarge language modelsPhenytoinTherapeutic drug monitoringclinical decision support
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Real-world pharmacoclinical implementation of risdiplam under a national SMA protocol: A hospital pharmacy registry-based case series.

BACKGROUND: Spinal muscular atrophy (SMA) is a rare neuromuscular disorder treated with disease-modifying therapies such as risdiplam. In Spain, its use is regulated by a national pharmacoclinical protocol that requires structured monitoring. OBJECTIVES: To evaluate real-world use, protocol adherence and registry completeness of risdiplam in routine clinical practice. METHODS: A retrospective registry-based case series was conducted in…

PubMed2026

Sex Differences in Prescribing Patterns of Anticholinergics and β3-Adrenoceptor Agonists for Overactive Bladder: A Nationwide Study in Japan, FY2017-FY2024.

OBJECTIVES: Drug treatment for overactive bladder (OAB) is changing, with an increasing share of β3-adrenoceptor agonists relative to anticholinergics, but it is not known how far this change has progressed in older people, who are most vulnerable to the risks of anticholinergics. We examined the use of these drugs in Japan by age and sex, and analyzed changes in anticholinergic burden using a large database. METHODS: We analyzed outpa…

PubMed2026

ARMED: an Australian cohort retrospective observational study to understand motivations for switching and clinical outcomes of individuals switching to dolutegravir/lamivudine.

BACKGROUND: Although clinical trials and real-world studies demonstrate the efficacy and tolerability of dolutegravir and lamivudine (DTG + 3TC), Australian real-world evidence remains limited despite differences in access to health care and prescribing context. Here, we present the motivations for switching to DTG/3TC (a fixed-dose, single-tablet regimen) and treatment outcomes. METHODS: We performed a retrospective, observational ana…

PubMed2026

Antibiotic prescribing that aligns with a host-protein test at US acute care settings is associated with fewer subsequent hospitalisations.

BACKGROUND: Differentiating bacterial from viral infections in acute care is challenging, often leading to inappropriate antibiotic use. MeMed BV (MMBV) is a host-protein test integrating TNF-related apoptosis-inducing ligand, induced protein-10 and C-reactive protein to distinguish infection aetiology. We evaluated whether alignment between MMBV results and antibiotic prescribing is associated with downstream clinical outcomes and cos…