PubMed چکیده/رکورد

Conditional Perplexity Scoring for Large Language Model-Generated Differential Diagnoses in Case Reports: Preliminary Computational Evaluation.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Large language models (LLMs) are increasingly used to generate differential diagnoses from clinical narratives. However, LLM-based diagnostic clinical decision support systems still lack a quantitative measure of how strongly a diagnosis is supported by the available case description. Conditional perplexity score quantifies how predictable a target text is given in a preceding context, with lower scores indicating greater predictability. We hypothesized that this concept can be adapted to diagnostic reasoning by treating the prediagnostic case description as the context and a diagnosis as the target text. OBJECTIVE: This study aims to evaluate whether conditional perplexity scores, computed by an independent LLM and conditioned on case-report narratives, differ between physician-verified correct and incorrect LLM-generated diagnoses. Specifically, we hypothesized that the correct LLM-generated diagnosis verified by physicians would have lower conditional perplexity scores than incorrect LLM-generated differential diagnoses. A secondary outcome was to compare this scoring behavior across differential diagnosis lists generated by different LLMs. METHODS: We performed a preliminary computational analysis of 392 peer-reviewed diagnostic case reports published in the American Journal of Case Reports in 2022. For each case, the prediagnostic clinical description was used as the conditioning context, and the case report-defined final diagnoses were treated as the gold standard. Conditional perplexity scores for differential diagnosis lists previously generated by LLaMA2, Bard, and GPT-4 were computed using an independent longer-context LLM, Qwen2.5-1.5B. We compared case report-defined final diagnoses, correct LLM-generated diagnoses verified by physicians, and incorrect generated diagnoses using nonparametric comparisons and receiver operating characteristic analyses. RESULTS: All 392 cases had complete case descriptions and case report-defined final diagnoses. Across the top-10 differential diagnosis lists generated by LLaMA2, Bard, and GPT-4, 823 correct LLM-generated diagnoses verified by physicians and 10,875 incorrect generated diagnoses were analyzed. Case report-defined final diagnoses had lower conditional perplexity scores than incorrect generated diagnoses (median 39.9, IQR 17.7-119.9 vs median 133.3, IQR 37.5-672.1). Correct LLM-generated diagnoses also had lower conditional perplexity scores than incorrect LLM-generated diagnoses (median 43.3, IQR 16.6-147.5 vs median 133.3, IQR 37.6-672.1). Candidate-level discrimination was moderate overall (area under the receiver operating characteristic curve [AUC] 0.666, 95% CI 0.644-0.689) and was the highest for GPT-4-generated differential diagnosis lists (AUC 0.678, 95% CI 0.652-0.705), followed by LLaMA2 (AUC 0.662, 95% CI 0.625-0.698) and Bard (AUC 0.648, 95% CI 0.617-0.681). In within-case analyses, correct diagnoses had lower conditional perplexity than the mean incorrect diagnosis in 88.1% (237/269) to 91.1% (195/214) of evaluable lists. CONCLUSIONS: Conditional perplexity provided a moderate quantitative signal associated with physician-verified correctness but did not reliably rank the correct diagnosis ahead of the strongest incorrect candidate, limiting its use as a stand-alone reranking method.

نتیجه فارسی

این مطالعه محاسباتی اولیه بررسی کرد که آیا امتیاز تنش مشروط می‌تواند به عنوان معیاری برای سنجش صحت تشخیص‌های تولید شده توسط LLM عمل کند. داده‌ها از 392 گزارش مورد منتشر شده در سال 2022 استخراج شد. نتایج نشان داد که امتیاز تنش مشروط پایین‌تری برای تشخیص‌های صحیح نسبت به غلط‌ها وجود دارد، اما این معیار قادر به رتبه‌بندی دقیق تشخیص صحیح در برابر رقیب غلط قوی نبود.

  • امتیاز تنش مشروط پایین‌تری برای تشخیص‌های صحیح نسبت به غلط‌ها یافت شد.
  • سیگنال کمی متوسطی مرتبط با صحت تأیید شده توسط پزشکان وجود داشت.
  • این معیار قادر به رتبه‌بندی دقیق تشخیص صحیح در برابر رقیب غلط قوی نبود.
  • گزارش مورد تعریف شده نهایی دارای امتیاز تنش مشروط پایین‌تری نسبت به تشخیص‌های غلط تولید شده بود.
  • گزارش مورد تعریف شده نهایی دارای امتیاز تنش مشروط پایین‌تری نسبت به تشخیص‌های صحیح تولید شده بود.

ترجمه فارسی چکیده

مدل‌های زبانی بزرگ (LLM) به طور فزاینده‌ای برای تولید تشخیص‌های تفکیکی از روایت‌های بالینی استفاده می‌شوند. با این حال، سیستم‌های پشتیبانی تصمیم‌گیری بالینی مبتنی بر LLM هنوز فاقد معیار کمی برای سنجش میزان پشتیبانی یک تشخیص از توصیف مورد موجود هستند. امتیاز تنش مشروط (Conditional Perplexity) میزان پیش‌بینی‌پذیری یک متن هدف را در یک زمینه پیشین کمی می‌کند، با امتیازات پایین‌تر نشان‌دهنده پیش‌بینی‌پذیری بیشتر. این مطالعه هدف دارد ارزیابی کند که آیا امتیازات تنش مشروط، محاسبه شده توسط یک LLM مستقل و مشروط بر روایت‌های گزارش مورد، بین تشخیص‌های صحیح و غلط تولید شده توسط LLM که توسط پزشکان تأیید شده‌اند، تفاوت دارند. نتایج نشان داد که تشخیص‌های نهایی تعریف شده در گزارش مورد و تشخیص‌های صحیح تولید شده توسط LLM امتیازات تنش مشروط پایین‌تری نسبت به تشخیص‌های غلط تولید شده داشتند. امتیاز تنش مشروط سیگنال کمی متوسطی مرتبط با صحت تأیید شده توسط پزشکان ارائه داد، اما قادر به رتبه‌بندی تشخیص صحیح در برابر رقیب غلط قوی نبود.

روش پژوهش

این مطالعه یک تحلیل محاسباتی اولیه از 392 گزارش مورد منتشر شده در American Journal of Case Reports در سال 2022 انجام داد. برای هر مورد، توصیف بالینی پیش‌دیagnostیک به عنوان زمینه استفاده شد و تشخیص‌های نهایی تعریف شده در گزارش مورد به عنوان استاندارد طلایی در نظر گرفته شد. امتیازات تنش مشروط برای لیست‌های تشخیص تفکیکی قبلاً تولید شده توسط LLaMA2، Bard و GPT-4 با استفاده از یک LLM مستقل با زمینه طولانی‌تر، Qwen2.5-1.5B، محاسبه شد.

محدودیت‌ها

این مطالعه یک ارزیابی محاسباتی اولیه بود و شامل نتایج بالینی مستقیم نبود. نتایج ممکن است به دلیل ماهیت محاسباتی و استفاده از مدل‌های زبانی خاص محدودیت داشته باشند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
392 گزارش مورد منتشر شده در American Journal of Case Reports در سال 2022.
مداخله/مواجهه
محاسبه امتیاز تنش مشروط بر اساس توصیف بالینی پیش‌دیagnostیک و تشخیص‌های نهایی تعریف شده در گزارش مورد.
مقایسه
تشخیص‌های غلط تولید شده توسط LLM.
حجم نمونه
392 مورد (823 تشخیص صحیح و 10,875 تشخیص غلط)

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

artificial intelligencediagnosisgenerative artificial intelligencelarge language modelnatural language processing
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Computational Network Analysis for Defining Transcriptional Programs.

Cancer cell identity is governed by coordinated transcriptional programs that are frequently rewired during tumorigenesis. Systematic identification of cancer type-specific gene regulatory networks provides a framework for understanding oncogenic state transitions and for prioritizing candidate therapeutic targets. Here, we present a reproducible network-based workflow for reconstructing and analyzing transcriptional regulatory program…

PubMed2027

Topic-Driven Bibliometrics and Trend Intelligence for Stem Cell and Cancer Research.

The rapid growth of biomedical literature has created an urgent need for computational tools that enable researchers to systematically analyze publication trends, identify emerging research themes, and map the evolution of scientific fields. PubMed Atlas is a command-line and web-enabled workflow for topic-driven bibliometrics and trend intelligence using PubMed E-utilities. The pipeline executes PubMed queries, retrieves matching PMID…

PubMed2026

Nursing Students' Reports of Patient Safety Incidents and Reasons During Clinical Placements: A Secondary Analysis of the International Data.

Clinical placements expose nursing students to patient safety incidents and provide important opportunities for learning about safe care. This study explored how nursing students in four countries recognized and interpreted patient safety incidents encountered or witnessed during clinical practice, including perceived contributing factors. A secondary qualitative content analysis was conducted using narrative data from 1442 undergradua…

PubMed2026

Frequency of Clinically Relevant Drug-Drug Interactions Between Tyrosine Kinase Inhibitors and Proton Pump Inhibitors in Patients With Cancer Using Real-World Data.

BACKGROUND: Proton pump inhibitors (PPIs) raise stomach pH, leading to reduced bioavailability of many tyrosine kinase inhibitors (TKIs), thereby affecting treatment outcomes. To what extent this interaction occurs in clinical practice remains underexplored. OBJECTIVE: To determine the frequency of clinically relevant interactions between TKIs and PPIs in clinical practice and the duration of concomitant prescription. METHODS: A retros…