PubMed دسترسی آزاد

Accuracy, Self-Reported Confidence, and Overconfidence of Large Language Models in Endodontics: An Evaluation Using National Specialty Examination Questions.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND This study aimed to evaluate the accuracy, confidence, and overconfidence behavior of large language models (LLMs) in endodontics using questions derived from a national specialty entrance examination. MATERIAL AND METHODS A total of 123 text-based endodontic questions from the Turkish Dental Specialty Examination (2017-2026) were included after excluding annulled and image-based questions. Three LLMs (ChatGPT, Claude, and Gemini) were assessed. Each model answered all questions using a standardized prompt and provided a confidence score (0%-100%). Accuracy was recorded as correct/incorrect. Overconfidence was defined as incorrect responses with ≥90% confidence. Statistical analyses were performed using Cochran's Q test, Friedman test, and post hoc pairwise comparisons with Bonferroni correction. RESULTS Accuracy rates were 89.4% for ChatGPT, 76.4% for Claude, and 90.2% for Gemini, with significant differences among models (χ²(2)=21.00, P<0.001). Confidence scores differed significantly (c2(2)=213.40, P<0.001), with Gemini demonstrating highest confidence (99.51±1.49), followed by ChatGPT (90.88±9.19) and Claude (80.22±11.10). Overconfidence rates were 9.8% for Gemini and 5.7% for ChatGPT, while no overconfident responses were observed for Claude (χ²(2)=14.53, P=0.001). Despite similar accuracy between ChatGPT and Gemini, confidence patterns differed markedly, demonstrating that comparable accuracy does not necessarily reflect comparable reliability. CONCLUSIONS LLMs demonstrated high accuracy in answering text-based endodontic examination questions; however, significant differences were observed in confidence behavior and overconfidence patterns. The presence of high-confidence incorrect responses suggests that accuracy alone may be insufficient to fully evaluate model reliability. These findings highlight the importance of considering confidence-related behavior alongside accuracy when assessing LLM performance in examination-style endodontic tasks.

نتیجه فارسی

این مطالعه دقت، اعتماد به نفس و بیش‌اعتمادی سه مدل زبان بزرگ (ChatGPT، Claude و Gemini) را در پاسخگویی به سوالات دندان‌پزشکی ریشه‌ای بررسی کرد. نتایج نشان داد که دقت مدل‌ها بالا بود، اما الگوهای اعتماد به نفس و نرخ بیش‌اعتمادی بین آن‌ها متفاوت بود. این یافته‌ها حاکی از آن است که دقت به تنهایی برای ارزیابی کامل عملکرد مدل کافی نیست.

  • سه مدل LLM (ChatGPT، Claude و Gemini) بر روی ۱۲۳ سوال آزمون تخصصی دندان‌پزشکی ارزیابی شدند.
  • دقت ChatGPT ۸۹.۴٪، Claude ۷۶.۴٪ و Gemini ۹۰.۲٪ بود.
  • Gemini بالاترین اعتماد به نفس (۹۹.۵۱٪) و نرخ بیش‌اعتمادی (۹.۸٪) را نشان داد.
  • Claude هیچ پاسخ بیش‌اعتمادی نداشت، هرچند دقت پایین‌تری داشت.
  • دقت قابل مقایسه لزوماً به معنای قابلیت اطمینان قابل مقایسه نیست.

ترجمه فارسی چکیده

این مطالعه هدف داشت دقت، اعتماد به نفس و رفتار بیش‌اعتمادی مدل‌های زبان بزرگ (LLM) را در دندان‌پزشکی ریشه‌ای با استفاده از سوالات استخراج‌شده از آزمون ورودی تخصصی ملی ارزیابی کند. ۱۲۳ سوال متنی دندان‌پزشکی ریشه‌ای از آزمون تخصصی دندان‌پزشکی ترکیه (۲۰۱۷-۲۰۲۶) پس از حذف سوالات لغو‌شده و مبتنی بر تصویر وارد شدند. سه مدل LLM (ChatGPT، Claude و Gemini) ارزیابی شدند. هر مدل با استفاده از یک دستورالعمل استاندارد به تمام سوالات پاسخ داد و امتیاز اعتماد به نفس (۰ تا ۱۰۰٪) ارائه کرد. دقت به عنوان درست/غلط ثبت شد. بیش‌اعتمادی به پاسخ‌های غلط با اعتماد به نفس ≥۹۰٪ تعریف شد. تحلیل‌های آماری با استفاده از آزمون کوکران Q، آزمون فریدمن و مقایسه‌های جفت‌به‌جفت پس از آزمون با اصلاح Bonferroni انجام شد. نرخ‌های دقت برای ChatGPT ۸۹.۴٪، برای Claude ۷۶.۴٪ و برای Gemini ۹۰.۲٪ بود که تفاوت‌های معناداری بین مدل‌ها مشاهده شد (χ²(2)=۲۱.۰۰، P<0.001). امتیازات اعتماد به نفس تفاوت معناداری داشتند (c2(2)=۲۱۳.۴۰، P<0.001)، با نشان دادن بالاترین اعتماد به نفس توسط Gemini (۹۹.۵۱±۱.۴۹)، به دنبال آن ChatGPT (۹۰.۸۸±۹.۱۹) و Claude (۸۰.۲۲±۱۱.۱۰). نرخ‌های بیش‌اعتمادی برای Gemini ۹.۸٪ و برای ChatGPT ۵.۷٪ بود، در حالی که برای Claude هیچ پاسخ بیش‌اعتمادی مشاهده نشد (χ²(2)=۱۴.۵۳، P=0.001). با وجود دقت مشابه بین ChatGPT و Gemini، الگوهای اعتماد به نفس به طور قابل توجهی متفاوت بودند که نشان می‌دهد دقت قابل مقایسه لزوماً به معنای قابلیت اطمینان قابل مقایسه نیست. مدل‌های LLM دقت بالایی در پاسخگویی به سوالات آزمون دندان‌پزشکی ریشه‌ای متنی نشان دادند؛ با این حال، تفاوت‌های معناداری در رفتار اعتماد به نفس و الگوهای بیش‌اعتمادی مشاهده شد. وجود پاسخ‌های غلط با اعتماد به نفس بالا نشان می‌دهد که دقت به تنهایی ممکن است برای ارزیابی کامل قابلیت اطمینان مدل کافی نباشد. این یافته‌ها اهمیت در نظر گرفتن رفتار مرتبط با اعتماد به نفس در کنار دقت را هنگام ارزیابی عملکرد LLM در وظایف دندان‌پزشکی ریشه‌ای با سبک آزمون برجسته می‌کنند.

روش پژوهش

این مطالعه یک مطالعه ارزیابی‌کننده با استفاده از سوالات استاندارد شده از آزمون تخصصی ملی بود. تحلیل‌های آماری شامل آزمون کوکران Q، آزمون فریدمن و مقایسه‌های جفت‌به‌جوست.

محدودیت‌ها

محدودیت‌های گزارش‌شده شامل حذف سوالات لغو‌شده و مبتنی بر تصویر است. همچنین، استفاده از سوالات یک آزمون خاص (ترکیه) ممکن است تعمیم‌پذیری را محدود کند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
سوالات متنی دندان‌پزشکی ریشه‌ای از آزمون تخصصی دندان‌پزشکی ترکیه (۲۰۱۷-۲۰۲۶).
مداخله/مواجهه
ارزیابی سه مدل زبان بزرگ (ChatGPT، Claude و Gemini) با استفاده از یک دستورالعمل استاندارد.
مقایسه
هیچ مقایسه‌کننده‌ای در متن گزارش نشده است.
حجم نمونه
۱۲۳ سوال.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Different irrigation solutions in calcium hydroxide removal: Do high-frequency ultrasonic and laser activation make the difference?

This study aimed to comparatively evaluate the effectiveness of different irrigation solutions and activation techniques in removing calcium hydroxide (CH) from standardized artificial grooves located in the apical third of root canals. Ninety extracted single-rooted human maxillary central incisors were instrumented using a standardized preparation protocol. Artificial grooves were created in the apical third of split roots and filled…

PubMed2026

Indications for endodontic surgery in private practice and temporal trends before and after adoption of laser-activated irrigation: a retrospective practice-based study.

OBJECTIVES: To characterize indications for endodontic microsurgery (EMS) and compare the relative proportion of these indications across two calendar periods before and after the practice's adoption of laser-activated irrigation (LAI). MATERIALS AND METHODS: A retrospective time-period comparison of consecutive EMS cases (n = 1,672) treated by six endodontists (April 2019-August 2025) was analyzed. Two periods were prespecified based …

PubMed2026

Short-term radiographic changes of extruded NeoSealer Flo in a simulated enlarged-apex model: influence of moisture and periapical support.

OBJECTIVES: To radiographically evaluate short-term changes in NeoSealer Flo extruded beyond enlarged apical preparations under different moisture and apical-support conditions. MATERIALS AND METHODS: Thirty-two extracted single-rooted human teeth were randomly allocated to four experimental conditions (n = 8 per group). Teeth were prepared with reciprocating NiTi instruments to size 50/0.05 and filled with a matched gutta-percha cone …

PubMed2026

Effect of Root Canal Dressings on Extraradicular pH in a Sealed Apex Model: An In Vitro Study.

OBJECTIVES: Endodontic treatment is mandatory after severe dental trauma such as avulsion or intrusion. Initial intracanal dressing with calcium hydroxide should be avoided according to international guidelines due to potential periodontal ligament damage by pH elevation. However, whether an intracanal dressing is able to alter the extraradicular site remains unclear. This in vitro study quantified extraradicular pH changes through den…