Natural Language Processing Identification of Nonprescribed Fentanyl Use in Electronic Health Records: Algorithm Development and Validation Study.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
BACKGROUND: Overdose and suicide due to nonprescribed fentanyl use have increased significantly, yet health care systems lack reliable methods to identify patients who use nonprescribed fentanyl. International Classification of Diseases codes are inconsistent and do not specify nonprescribed fentanyl use. OBJECTIVE: This study aimed to develop natural language processing approaches to identifying nonprescribed fentanyl use in electronic health record (EHR) documentation. METHODS: This retrospective study included Veterans Health Administration patients seen between April 5, 2023, and December 23, 2024. A term list was developed to identify fentanyl-related mentions in clinical text, and 250-character snippets surrounding identified mentions were extracted. Veterans (n=3878) were randomly sampled from 5 predefined groups based on the presence of 1 of 4 terms ("fent," "blues," "M30s," and "tranq") in their EHR documentation. Physician annotators classified snippets into "nonprescribed fentanyl use," "prescribed fentanyl use," or "other," with interannotator agreement evaluated using the mean pairwise Cohen κ. Cross-validation folds were constructed at the patient level between training and test sets. Penalized logistic regression, Bio-ClinicalBERT, Llama 3-8B, and Mistral-7B were trained on labeled data and compared. Model performance was evaluated using precision, recall, and F1-scores for each class, with a focus on the nonprescribed fentanyl use class as the primary label of clinical interest using bootstrapped 95% CIs. A fairness analysis and Shapley additive explanations analysis were performed using Bio-ClinicalBERT. External validation was performed using Bio-ClinicalBERT on an independent sample of 200 snippets, each representing a unique patient from January 2025 to June 2026, with precision reported as the primary validation metric. RESULTS: Of 7389 snippets, 9.6% (n=709) were classified as "nonprescribed fentanyl use," 40.3% (n=2981) were classified as "prescribed fentanyl use," and 50% (n=3699) were classified as "other." Interannotator agreement was high (κ=0.822). Llama 3-8B achieved the highest F1-score for nonprescribed fentanyl use (0.87, 95% CI 0.83-0.92), followed by Mistral-7B (0.80, 95% CI 0.75-0.84), Bio-ClinicalBERT (0.80, 95% CI 0.74-0.85), and penalized logistic regression (0.74, 95% CI 0.73-0.75). Performance was consistent across demographic subgroups, with lower performance for the nonprescribed fentanyl use class observed in female and Hispanic subgroups. Shapley additive explanations analysis revealed clinically meaningful discriminating terms for each class, although subword tokens required contextual interpretation. External validation of Bio-ClinicalBERT demonstrated a precision of 0.79 for nonprescribed fentanyl use. CONCLUSIONS: Natural language processing can identify nonprescribed fentanyl use in EHR documentation, although model performance for this class was lower than overall model performance, reflecting the clinical complexity of identifying nonprescribed use and the variable ways in which clinicians document this problem. This approach may support risk prediction and targeting of interventions to patients exposed to nonprescribed fentanyl.
نتیجه فارسی
این مطالعه الگوریتمهای پردازش زبان طبیعی را برای شناسایی استفاده از فنتانیل بدون نسخه در سوابق مراقبتهای الکترونیکی توسعه و اعتبارسنجی کرد. مدلهای یادگیری عمیق مانند Llama 3-8B عملکرد بهتری نسبت به سایر مدلها و رگرسیون لجستیک نشان دادند. دقت مدلها در شناسایی استفاده از فنتانیل بدون نسخه نسبت به سایر کلاسها کمتر بود که نشاندهنده پیچیدگی بالینی این موضوع است. اعتبارسنجی خارجی نشان داد که مدل Bio-ClinicalBERT دقت 0.79 را برای این کلاس نشان داد.
- توسعه الگوریتمهای پردازش زبان طبیعی برای شناسایی استفاده از فنتانیل بدون نسخه.
- Llama 3-8B بالاترین امتیاز F1 (0.87) را برای شناسایی این مصرف داشت.
- دقت مدلها در شناسایی استفاده از فنتانیل بدون نسخه نسبت به سایر کلاسها کمتر بود.
- اعتبارسنجی خارجی Bio-ClinicalBERT دقت 0.79 را نشان داد.
- این روش میتواند در پیشبینی ریسک و هدفمندسازی مداخلات مفید باشد.
ترجمه فارسی چکیده
افزایش مصرف فنتانیل بدون نسخه و مرگومیر ناشی از آن، بهویژه در میان جوانان، نگرانیهای جدی بهوجود آورده است. سیستمهای مراقبت سلامت هنوز روشهای مطمئنی برای شناسایی بیماران مصرفکننده فنتانیل بدون نسخه ندارند. کدهای طبقهبندی بینالمللی بیماریها (ICD) برای این منظور دقیق نیستند. این مطالعه به توسعه رویکردهای پردازش زبان طبیعی برای شناسایی استفاده از فنتانیل بدون نسخه در سوابق مراقبتهای الکترونیکی (EHR) پرداخت. دادهها از بیماران سیستم سلامت وابسته به veterans (VA) جمعآوری شد. الگوریتمهای مختلف شامل رگرسیون لجستیک پاداشدار، Bio-ClinicalBERT، Llama 3-8B و Mistral-7B توسعه یافتند. نتایج نشان داد که Llama 3-8B بالاترین امتیاز F1 را برای شناسایی استفاده از فنتانیل بدون نسخه (0.87) داشت. اعتبارسنجی خارجی با Bio-ClinicalBERT دقت 0.79 را نشان داد. این روش میتواند در پیشبینی ریسک و هدفمندسازی مداخلات مفید باشد.
روش پژوهش
این مطالعه یک مطالعه پسرونده بود که بیماران VA را از 5 آوریل 2023 تا 23 دسامبر 2024 شامل میشد. 7389 نمونه متنی از سوابق بیماران استخراج شد. پزشکان برچسبگذاریها را انجام دادند و مدلها با استفاده از دقت،recall و امتیاز F1 ارزیابی شدند.
محدودیتها
محدودیتهای ذکر نشده در متن موجود است.
نمای PICO و پیامدها
- جمعیت
- بیماران سیستم سلامت وابسته به veterans (VA)
- مداخله/مواجهه
- توسعه الگوریتمهای پردازش زبان طبیعی برای شناسایی استفاده از فنتانیل بدون نسخه
- مقایسه
- رگرسیون لجستیک پاداشدار، Bio-ClinicalBERT، Llama 3-8B، Mistral-7B
- حجم نمونه
- 7389 نمونه متنی، 3878 بیمار
متن کامل اصلی
لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز میشود.
باز کردن متن کامل