PubMed چکیده/رکورد

Large Language Models for World Health Organization-Uppsala Monitoring Centre Drug-Adverse Event Causality Assessment Using Food and Drug Administration Adverse Event Reporting System Cases: Comparative Performance Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Causality assessment is central to pharmacovigilance but remains resource-intensive and subjective. The applicability of large language models (LLMs) to formal World Health Organization-Uppsala Monitoring Centre (WHO-UMC) drug-adverse event causality assessment has not been well established. OBJECTIVE: This study aims to evaluate the performance of LLMs in WHO-UMC causality assessment. METHODS: A curated set of 55 cases derived from the US Food and Drug Administration Adverse Event Reporting System, comprising 337 drug-level assessments, was constructed. Cases involving 2 to 11 suspected drugs were stratified by drug count, and 5 cases were sampled from each stratum. To ensure representation of rare but clinically important categories, 5 additional cases containing at least 1 "Certain" drug-adverse event pair were included. Case data were reorganized into a standardized semistructured format that preserved key elements required for WHO-UMC causality assessment. Domain experts conducted a pilot evaluation to align interpretation criteria prior to independently assessing the final dataset, yielding an interexpert agreement (Fleiss κ) of 0.762 across 337 drug-level assessments. Multiple prompting strategies, including standard prompting, chain-of-thought (CoT), CoT with self-consistency, few-shot, reasoning and acting, and tree-of-thought prompting, were applied across multiple LLMs, including GPT-5.4 and its mini variant and Gemini 2.5 Flash and Pro, via their respective application programming interfaces. Agreement with expert assessments was quantified using Cohen κ, weighted κ, and accuracy metrics. Internal consistency across repeated inferences was evaluated using Fleiss κ. RESULTS: Performance varied across models and prompting strategies. Cohen κ ranged from 0.368 to 0.641, weighted κ ranged from 0.641 to 0.821, accuracy ranged from 0.583 to 0.804, and balanced accuracy ranged from 0.513 to 0.735. Fleiss κ ranged from 0.730 to 0.915, corresponding to substantial to almost perfect agreement. The highest Cohen κ was observed for Gemini 2.5 Flash with CoT prompting (0.641). Gemini 2.5 Flash with CoT-self-consistency prompting showed a Cohen κ of 0.640 and achieved the highest observed point estimates for weighted κ (0.821), accuracy (0.804), and Fleiss κ (0.915), although the gains over other prompting strategies were modest. Category-level performance for this model showed higher performance for "Certain" (F1-score=0.793), "Probable/Likely" (F1-score=0.794), and "Unlikely" (F1-score=0.898), whereas performance for "Possible" remained substantially lower (F1-score=0.293), reflecting the difficulty of intermediate causality assessment. CONCLUSIONS: LLMs demonstrated moderate to substantial agreement in WHO-UMC causality assessment, indicating meaningful but still limited performance relative to expert judgment. Although LLMs are not suitable for independent decision-making, they may serve as supportive tools in pharmacovigilance workflows, particularly for preliminary case triage. Further studies using larger and more diverse datasets and evaluating performance on raw narrative reports are warranted.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

GPTGeminiWHO-UMC causality assessmentWorld Health Organization–Uppsala Monitoring Centre causality assessmentgenerative pretrained transformerlarge language modelsprompt engineering
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Flibanserin safety in real-world use: a decade of US Food and Drug Administration Adverse Event Reporting System pharmacovigilance evidence.

BACKGROUND: Flibanserin, a centrally acting 5-HT1A agonist and 5-HT2A antagonist, is the first approved therapy for acquired, generalized hypoactive sexual desire disorder in premenopausal women. Despite its clinical availability, post-marketing data on real-world safety remain limited. METHODS: A retrospective pharmacovigilance analysis was conducted using the US Food and Drug Administration Adverse Event Reporting System from 2016 to…

PubMed2026

Risk of Major Malformations Following First-Trimester Exposure to Cariprazine: Preliminary Data From the MGH National Pregnancy Registry for Psychiatric Medications.

OBJECTIVE: Systematically collected pregnancy safety data for cariprazine have been lacking, despite growing use of this medication across psychiatric indications. The goal of this analysis was to determine the risk of major malformations among infants of mothers with psychiatric illness who used cariprazine during the first trimester of pregnancy compared to unexposed controls. METHODS: The National Pregnancy Registry for Psychiatric …

PubMed2026

Using Pharmacovigilance Data for Signal Detection of Drug Interactions for Rosuvastatin.

The 3-hydroxy-3-methyl-glutaryl-coenzyme A reductase inhibitor rosuvastatin is a substrate of breast cancer resistance protein (BCRP). BCRP inhibition increases rosuvastatin plasma concentrations and may result in concentration-dependent muscle toxicity, at worst rhabdomyolysis. We investigated if concomitant use of rosuvastatin and drugs identified as in vitro BCRP inhibitors shows an increased number of rhabdomyolysis events based on…

PubMed2026

Assessing the risk of cataracts associated with medications: A pharmacovigilance analysis of the FAERS database.

Cataracts are a leading cause of global blindness and visual impairment, with drug-induced cataracts emerging as a significant yet understudied contributor. This study aimed to comprehensively and systematically investigate medication-related cataract risk signals using data from the FDA Adverse Event Reporting System. We searched the FDA Adverse Event Reporting System database for all reported cases of medication-related cataracts fro…