PubMed دسترسی آزاد

Evaluating Large Language Models in Clinical Audiology (AUDIOLOGYBENCH): Benchmark Development and Validation Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Large language models (LLMs) are increasingly being explored for clinical decision support, but their performance in audiology has not been systematically benchmarked using clinically grounded case materials and rubric-based safety evaluations. OBJECTIVE: This study aimed to develop and evaluate AUDIOLOGYBENCH, a 3-tier benchmark for characterizing frontier LLM capability in clinical audiology along (1) curated domain knowledge, (2) literature-derived evidence, and (3) clinical reasoning under multimodal case input, with an explicit human audit of the automated adjudicator on the primary end point. METHODS: The benchmark comprises 3139 objective items from educational resources, 3175 research article-derived items from peer-reviewed articles published between 2015 and 2025, and 67 multimodal clinical case studies graded against a standardized A-F rubric with 6 prespecified critical-error types that cap scores at D or F. Eight models were evaluated on the educational objective items: 4 frontier multimodal models (Gemini 2.5 Pro, Grok 4, OpenAI O3, and Claude Sonnet 4 Thinking) were evaluated on the research article-derived items, and on 804 case study evaluations. Adjudication used Gemini 2.5 Pro (objective and research-derived items) and Claude Opus 4.5 (case studies). The case study adjudicator was independently audited against PhD-level audiologist consensus on blinded subsamples, supplemented by a post-stratified human-calibrated sensitivity analysis. RESULTS: A striking task-type dissociation emerged on case studies: clinical recommendations (Q3) achieved a mean score of 89.74 (SD 13.92, 95% CI 88.07-91.41), a 98.1% (263/268) pass rate, and no dangerous recommendations; audiometric numerical interpretation (Q1) achieved a mean score of 67.89 (SD 18.47, 95% CI 65.68-70.10), with a 35.4% (95/268) critical-error rate; and differential diagnosis (Q2) achieved a mean score of 67.79 (SD 15.33, 95% CI 65.95-69.63). Question type, not model selection, dominated performance (eta-squared_H=0.333 vs 0.001; rank biserial r≥0.679). Interreviewer reliability between audiologists was high (quadratic-weighted κ of 0.78 and 0.85 across the 80-item and 50-item audits, respectively). When 2 audiologists regraded all 80 model Q1 responses with the diagnostic images available, the adjudicator's per-item Q1 labels diverged from human judgment (κ=0.05; overflagging; sensitivity: 19/26, 73%; positive predictive value: 19/53, 36%), yet its reweighted Q1 critical-error rate (36.2%) was broadly consistent with the image-grounded human estimates (28%-34%), suggesting no systematic inflation of the headline rate. The principal Q3>{Q1, Q2} ranking was preserved under post-stratified human calibration. Web-style multiple-choice items showed ceiling effects (>95% accuracy); short-answer prompts remained challenging (best 30%). CONCLUSIONS: Current frontier LLMs show strong recommendation generation but substantial limitations in audiometric numerical interpretation that are shared across models and that an automated adjudicator partially miscalibrated at the per-item level. AUDIOLOGYBENCH characterizes capability boundaries rather than certifying clinical readiness. Deployment of LLM-assisted audiology workflows requires structured human verification of all numerical findings and awareness of fabrication and severity misclassification failure modes documented here.

نتیجه فارسی

این مطالعه AUDIOLOGYBENCH را برای ارزیابی مدل‌های زبانی بزرگ در شنواری‌شناسی توسعه داد. نتایج نشان داد که مدل‌ها در تولید توصیه‌های بالینی قوی بودند، اما در تفسیر اعداد شنوایی‌سنجی و تشخیص تفکیکی محدودیت داشتند. معیار قابلیت‌های مدل را مشخص می‌کند، نه آمادگی بالینی.

  • AUDIOLOGYBENCH شامل ۶۷۱۹ مورد آموزشی و پژوهشی و ۶۷ مطالعه مورد بالینی است.
  • مدل‌ها در تولید توصیه‌های بالینی عملکرد خوبی داشتند (میانگین ۸۹.۷۴).
  • تفسیر اعداد شنوایی‌سنجی و تشخیص تفکیکی محدودیت‌های قابل توجهی داشتند.
  • بررسی انسانی برای یافته‌های عددی ضروری است.
  • معیار قابلیت‌های مدل را مشخص می‌کند، نه آمادگی بالینی.

ترجمه فارسی چکیده

این مطالعه هدف دارد AUDIOLOGYBENCH را توسعه دهد و ارزیابی کند، که یک معیار سه‌لایه برای مشخص کردن قابلیت‌های مدل‌های زبانی بزرگ (LLM) در شنواری‌شناسی بالینی است. این معیار شامل دانش حوزه‌ای، شواهد مقالات و استدلال بالینی در برابر ورودی‌های چندرسانه‌ای است. شامل ۳۱۳۹ مورد آموزشی، ۳۱۷۵ مورد از مقالات پژوهشی و ۶۷ مطالعه مورد بالینی است. هشت مدل ارزیابی شدند. نتایج نشان داد که توصیه‌های بالینی عملکرد خوبی داشتند، اما تفسیر اعداد شنوایی‌سنجی و تشخیص تفکیکی محدودیت‌های قابل توجهی داشتند. نویسندگان توصیه می‌کنند که برای استفاده از LLM در شنواری‌شناسی، بررسی انسانی ساختاریافته برای یافته‌های عددی ضروری است.

روش پژوهش

این مطالعه یک معیار سه‌لایه (AUDIOLOGYBENCH) را برای ارزیابی مدل‌های زبانی بزرگ در شنواری‌شناسی توسعه داد. شامل ۳۱۳۹ مورد آموزشی، ۳۱۷۵ مورد از مقالات پژوهشی و ۶۷ مطالعه مورد بالینی است. هشت مدل ارزیابی شدند.

محدودیت‌ها

نویسندگان توصیه می‌کنند که برای استفاده از LLM در شنواری‌شناسی، بررسی انسانی ساختاریافته برای یافته‌های عددی ضروری است.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
مدل‌های زبانی بزرگ (LLM) و مطالعه مورد بالینی شنواری‌شناسی.
مداخله/مواجهه
ارزیابی مدل‌های زبانی بزرگ (LLM) در شنواری‌شناسی با استفاده از AUDIOLOGYBENCH.
مقایسه
هشت مدل زبانی بزرگ ارزیابی شدند.
حجم نمونه
۶۷۱۹ مورد آموزشی و پژوهشی و ۶۷ مطالعه مورد بالینی.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

AIaudiologybenchmarkingclinical decision supportlarge language models
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Ethical Dilemmas in Audiology: A Comparative Analysis of Perspectives from the Middle East, North Africa, and the United States.

PURPOSE: The current study aimed to identify the ethical sensitivity of audiologists in Arabic-speaking countries of the Middle East and North Africa (MENA) and compare it with audiologists in the United States (U.S.) as reported in the 2002 and 2006 studies by Hawkins and colleagues. METHODS: In a cross-sectional survey, 123 audiologists from different countries in the MENA region were surveyed on their ethical perspectives regarding …

PubMed2026

Clinical characteristics and outcomes of referrals for speech and language delay to the Speech Therapy and Audiology department at a district hospital in Gauteng, South Africa.

BACKGROUND: Speech and language skills are essential for effective communication. Speech and language delays may have far-reaching consequences on a child's development; their early identification and intervention are paramount to a child reaching their full potential. OBJECTIVES: To describe the referral process, clinical characteristics, risk factors and outcomes of paediatric referrals to a Speech Therapy and Audiology (STA) departm…

PubMed2026

Evaluating the efficacy of artificial intelligence in audiology: a head-to-head comparison of ChatGPT and Gemini on hearing aid management.

OBJECTIVE: Hearing aid users frequently require accessible and immediate assistance for daily device management. This study aims to evaluate and compare the performance of two prominent Large Language Models (LLMs)-ChatGPT and Gemini, selected for their widespread public accessibility and market dominance-in providing accurate, comprehensible, and repeatable answers to frequently asked questions regarding hearing aids. METHODS: A compr…

PubMed2026

Current clinical practices in pediatric audiology among audiologists in India.

OBJECTIVE: Evidence-based practice is central to audiology; however, adherence to best-practice guidelines varies, particularly in pediatric settings. This study explored current pediatric audiology practices among audiologists in India. METHODS: A cross-sectional survey was conducted using an adapted version of the American Board of Audiology's Pediatric Audiology Clinical Practice Analysis Survey modified for the Indian context. The …