PubMed چکیده/رکورد

Chain-of-verification prompting for NIH stroke scale extraction using small and frontier large language models.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: The National Institutes of Health Stroke Scale (NIHSS) is critical to acute stroke care but is often documented in unstructured notes. Large language models (LLMs) can enable automated extraction, though smaller models often underperform relative to frontier systems. Chain-of-Verification (CoVe) prompting introduces a structured self-verification step that may improve performance. METHODS: We evaluated eight LLMs on 312 discharge summaries. Small models included LLaMA 3.2 3B, Ministral 3B, Gemma 3 4B, and Qwen 3 4B. Frontier models included GPT-5.2, Gemini 3 Pro, Claude Opus 4.5, and Grok 4. Each model was tested under a baseline and CoVe prompt. Outcomes were subscore exact-match accuracy, subscore mean absolute error (MAE), total score exact-match accuracy, and total score MAE. RESULTS: At baseline, small models achieved 53.2 ± 10.0% subscore accuracy and subscore MAE 0.84 ± 0.22, compared with 88.5 ± 10.1% and 0.15 ± 0.16 in frontier models (both p < 0.001). Total exact accuracy was low in both groups (7.7 ± 12.9% vs 35.9 ± 32.4%). CoVe significantly improved small-model performance (subscore accuracy 65.0 ± 10.9%; subscore MAE 0.55 ± 0.21; total MAE 4.84 ± 2.30 vs 7.19 ± 3.54 at baseline; all p < 0.001), although total exact accuracy remained modest (9.6 ± 15.7%). Frontier models showed no significant group-level change with CoVe. CONCLUSION: CoVe prompting substantially improves NIHSS extraction in small LLMs while producing negligible effects in frontier models. Although smaller model performance remains insufficient for standalone clinical deployment, CoVe prompting offers a promising avenue for further exploration.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

Chain-of-verificationClinical information extractionLarge language modelsNIH stroke scaleNatural language processing
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

An Evaluation of AI-Generated Clinical Notes in the OpenNotes Era: A Thematic Analysis of Clinician Discourse.

BACKGROUND: The integration of ambient artificial intelligence (AI) scribes into the OpenNotes environment presents a profound governance crisis in healthcare. While patient access to medical records was designed as a transparency reform, the introduction of machine-generated text introduces novel vulnerabilities regarding record integrity, liability, and patients' trust. OBJECTIVE: This study investigates how clinicians discursively n…

PubMed2026

Factors influencing discussion duration in breast cancer multidisciplinary team meetings: insights for streamlining care.

PURPOSE: Multidisciplinary team meetings (MDTMs) in breast cancer care improve outcomes but are time-consuming and costly. This study investigates using data from the Dutch national cancer registry (NCR) and hospital electronic medical records (EMR) to efficiently calculate MDTM discussion durations, while complying with privacy laws. METHODS: This retrospective study analyzed breast cancer MDTM discussion durations using NCR and EMR d…

PubMed2026

Radiological mass effect and neurological status are associated with mortality after burr-hole drainage for chronic subdural hematoma: a 10-year cohort study.

Chronic subdural hematoma (CSDH) is increasingly common in older adults and in patients receiving antithrombotic therapy. Although burr-hole drainage is generally safe and effective, perioperative mortality remains a concern, and reliable preoperative predictors are incompletely defined. We aimed to identify independent preoperative predictors of in-hospital mortality after burr-hole drainage for CSDH, with explicit characterization of…

PubMed2026

Neurocritical factors associated with mortality and functional recovery in pediatric trauma patients admitted to a tertiary PICU.

PURPOSE: To identify early neuroprognostic factors associated with functional neurological outcome in pediatric trauma patients requiring PICU admission. The primary outcome was Glasgow Outcome Scale (GOS) at discharge and 6 months; secondary outcomes were in-hospital mortality and brain death. METHODS: We conducted a retrospective cohort study of 425 consecutive pediatric trauma admissions to a tertiary PICU from June 2014 to December…