Conditional Perplexity Scoring for Large Language Model-Generated Differential Diagnoses in Case Reports: Preliminary Computational Evaluation.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
BACKGROUND: Large language models (LLMs) are increasingly used to generate differential diagnoses from clinical narratives. However, LLM-based diagnostic clinical decision support systems still lack a quantitative measure of how strongly a diagnosis is supported by the available case description. Conditional perplexity score quantifies how predictable a target text is given in a preceding context, with lower scores indicating greater predictability. We hypothesized that this concept can be adapted to diagnostic reasoning by treating the prediagnostic case description as the context and a diagnosis as the target text. OBJECTIVE: This study aims to evaluate whether conditional perplexity scores, computed by an independent LLM and conditioned on case-report narratives, differ between physician-verified correct and incorrect LLM-generated diagnoses. Specifically, we hypothesized that the correct LLM-generated diagnosis verified by physicians would have lower conditional perplexity scores than incorrect LLM-generated differential diagnoses. A secondary outcome was to compare this scoring behavior across differential diagnosis lists generated by different LLMs. METHODS: We performed a preliminary computational analysis of 392 peer-reviewed diagnostic case reports published in the American Journal of Case Reports in 2022. For each case, the prediagnostic clinical description was used as the conditioning context, and the case report-defined final diagnoses were treated as the gold standard. Conditional perplexity scores for differential diagnosis lists previously generated by LLaMA2, Bard, and GPT-4 were computed using an independent longer-context LLM, Qwen2.5-1.5B. We compared case report-defined final diagnoses, correct LLM-generated diagnoses verified by physicians, and incorrect generated diagnoses using nonparametric comparisons and receiver operating characteristic analyses. RESULTS: All 392 cases had complete case descriptions and case report-defined final diagnoses. Across the top-10 differential diagnosis lists generated by LLaMA2, Bard, and GPT-4, 823 correct LLM-generated diagnoses verified by physicians and 10,875 incorrect generated diagnoses were analyzed. Case report-defined final diagnoses had lower conditional perplexity scores than incorrect generated diagnoses (median 39.9, IQR 17.7-119.9 vs median 133.3, IQR 37.5-672.1). Correct LLM-generated diagnoses also had lower conditional perplexity scores than incorrect LLM-generated diagnoses (median 43.3, IQR 16.6-147.5 vs median 133.3, IQR 37.6-672.1). Candidate-level discrimination was moderate overall (area under the receiver operating characteristic curve [AUC] 0.666, 95% CI 0.644-0.689) and was the highest for GPT-4-generated differential diagnosis lists (AUC 0.678, 95% CI 0.652-0.705), followed by LLaMA2 (AUC 0.662, 95% CI 0.625-0.698) and Bard (AUC 0.648, 95% CI 0.617-0.681). In within-case analyses, correct diagnoses had lower conditional perplexity than the mean incorrect diagnosis in 88.1% (237/269) to 91.1% (195/214) of evaluable lists. CONCLUSIONS: Conditional perplexity provided a moderate quantitative signal associated with physician-verified correctness but did not reliably rank the correct diagnosis ahead of the strongest incorrect candidate, limiting its use as a stand-alone reranking method.
نتیجه فارسی
این مطالعه محاسباتی اولیه بررسی کرد که آیا امتیاز تنش مشروط میتواند به عنوان معیاری برای سنجش صحت تشخیصهای تولید شده توسط LLM عمل کند. دادهها از 392 گزارش مورد منتشر شده در سال 2022 استخراج شد. نتایج نشان داد که امتیاز تنش مشروط پایینتری برای تشخیصهای صحیح نسبت به غلطها وجود دارد، اما این معیار قادر به رتبهبندی دقیق تشخیص صحیح در برابر رقیب غلط قوی نبود.
- امتیاز تنش مشروط پایینتری برای تشخیصهای صحیح نسبت به غلطها یافت شد.
- سیگنال کمی متوسطی مرتبط با صحت تأیید شده توسط پزشکان وجود داشت.
- این معیار قادر به رتبهبندی دقیق تشخیص صحیح در برابر رقیب غلط قوی نبود.
- گزارش مورد تعریف شده نهایی دارای امتیاز تنش مشروط پایینتری نسبت به تشخیصهای غلط تولید شده بود.
- گزارش مورد تعریف شده نهایی دارای امتیاز تنش مشروط پایینتری نسبت به تشخیصهای صحیح تولید شده بود.
ترجمه فارسی چکیده
مدلهای زبانی بزرگ (LLM) به طور فزایندهای برای تولید تشخیصهای تفکیکی از روایتهای بالینی استفاده میشوند. با این حال، سیستمهای پشتیبانی تصمیمگیری بالینی مبتنی بر LLM هنوز فاقد معیار کمی برای سنجش میزان پشتیبانی یک تشخیص از توصیف مورد موجود هستند. امتیاز تنش مشروط (Conditional Perplexity) میزان پیشبینیپذیری یک متن هدف را در یک زمینه پیشین کمی میکند، با امتیازات پایینتر نشاندهنده پیشبینیپذیری بیشتر. این مطالعه هدف دارد ارزیابی کند که آیا امتیازات تنش مشروط، محاسبه شده توسط یک LLM مستقل و مشروط بر روایتهای گزارش مورد، بین تشخیصهای صحیح و غلط تولید شده توسط LLM که توسط پزشکان تأیید شدهاند، تفاوت دارند. نتایج نشان داد که تشخیصهای نهایی تعریف شده در گزارش مورد و تشخیصهای صحیح تولید شده توسط LLM امتیازات تنش مشروط پایینتری نسبت به تشخیصهای غلط تولید شده داشتند. امتیاز تنش مشروط سیگنال کمی متوسطی مرتبط با صحت تأیید شده توسط پزشکان ارائه داد، اما قادر به رتبهبندی تشخیص صحیح در برابر رقیب غلط قوی نبود.
روش پژوهش
این مطالعه یک تحلیل محاسباتی اولیه از 392 گزارش مورد منتشر شده در American Journal of Case Reports در سال 2022 انجام داد. برای هر مورد، توصیف بالینی پیشدیagnostیک به عنوان زمینه استفاده شد و تشخیصهای نهایی تعریف شده در گزارش مورد به عنوان استاندارد طلایی در نظر گرفته شد. امتیازات تنش مشروط برای لیستهای تشخیص تفکیکی قبلاً تولید شده توسط LLaMA2، Bard و GPT-4 با استفاده از یک LLM مستقل با زمینه طولانیتر، Qwen2.5-1.5B، محاسبه شد.
محدودیتها
این مطالعه یک ارزیابی محاسباتی اولیه بود و شامل نتایج بالینی مستقیم نبود. نتایج ممکن است به دلیل ماهیت محاسباتی و استفاده از مدلهای زبانی خاص محدودیت داشته باشند.
نمای PICO و پیامدها
- جمعیت
- 392 گزارش مورد منتشر شده در American Journal of Case Reports در سال 2022.
- مداخله/مواجهه
- محاسبه امتیاز تنش مشروط بر اساس توصیف بالینی پیشدیagnostیک و تشخیصهای نهایی تعریف شده در گزارش مورد.
- مقایسه
- تشخیصهای غلط تولید شده توسط LLM.
- حجم نمونه
- 392 مورد (823 تشخیص صحیح و 10,875 تشخیص غلط)
متن کامل اصلی
برای بررسی دسترسی کتابخانهای یا خرید، رکورد اصلی را باز کنید.
رفتن به منبع اصلی