Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
BACKGROUND: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. OBJECTIVE: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. METHODS: The benchmark comprises 3139 objective items from educational resources, 3175 research article-derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article-derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. RESULTS: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07-91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68-70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95-69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial r≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator's per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%-34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). CONCLUSIONS: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.
نتیجه فارسی
این مطالعه AUDIOLOGYBENCH را برای ارزیابی مدلهای زبانی بزرگ در شنواریشناسی توسعه داد. نتایج نشان داد که مدلها در تولید توصیههای بالینی قوی بودند، اما در تفسیر اعداد شنواییسنجی و تشخیص تفکیکی محدودیت داشتند. معیار قابلیتهای مدل را مشخص میکند، نه آمادگی بالینی.
- AUDIOLOGYBENCH شامل ۶۷۱۹ مورد آموزشی و پژوهشی و ۶۷ مطالعه مورد بالینی است.
- مدلها در تولید توصیههای بالینی عملکرد خوبی داشتند (میانگین ۸۹.۷۴).
- تفسیر اعداد شنواییسنجی و تشخیص تفکیکی محدودیتهای قابل توجهی داشتند.
- بررسی انسانی برای یافتههای عددی ضروری است.
- معیار قابلیتهای مدل را مشخص میکند، نه آمادگی بالینی.
ترجمه فارسی چکیده
این مطالعه هدف دارد AUDIOLOGYBENCH را توسعه دهد و ارزیابی کند، که یک معیار سهلایه برای مشخص کردن قابلیتهای مدلهای زبانی بزرگ (LLM) در شنواریشناسی بالینی است. این معیار شامل دانش حوزهای، شواهد مقالات و استدلال بالینی در برابر ورودیهای چندرسانهای است. شامل ۳۱۳۹ مورد آموزشی، ۳۱۷۵ مورد از مقالات پژوهشی و ۶۷ مطالعه مورد بالینی است. هشت مدل ارزیابی شدند. نتایج نشان داد که توصیههای بالینی عملکرد خوبی داشتند، اما تفسیر اعداد شنواییسنجی و تشخیص تفکیکی محدودیتهای قابل توجهی داشتند. نویسندگان توصیه میکنند که برای استفاده از LLM در شنواریشناسی، بررسی انسانی ساختاریافته برای یافتههای عددی ضروری است.
روش پژوهش
این مطالعه یک معیار سهلایه (AUDIOLOGYBENCH) را برای ارزیابی مدلهای زبانی بزرگ در شنواریشناسی توسعه داد. شامل ۳۱۳۹ مورد آموزشی، ۳۱۷۵ مورد از مقالات پژوهشی و ۶۷ مطالعه مورد بالینی است. هشت مدل ارزیابی شدند.
محدودیتها
نویسندگان توصیه میکنند که برای استفاده از LLM در شنواریشناسی، بررسی انسانی ساختاریافته برای یافتههای عددی ضروری است.
نمای PICO و پیامدها
- جمعیت
- مدلهای زبانی بزرگ (LLM) و مطالعه مورد بالینی شنواریشناسی.
- مداخله/مواجهه
- ارزیابی مدلهای زبانی بزرگ (LLM) در شنواریشناسی با استفاده از AUDIOLOGYBENCH.
- مقایسه
- هشت مدل زبانی بزرگ ارزیابی شدند.
- حجم نمونه
- ۶۷۱۹ مورد آموزشی و پژوهشی و ۶۷ مطالعه مورد بالینی.
متن کامل اصلی
لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز میشود.
باز کردن متن کامل