PubMed دسترسی آزاد

Accuracy and Error Patterns of References Generated by Large Language Models in Endodontics: The Role of Prompt Design and Model Selection.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND Large language models (LLMs) are increasingly used in healthcare; concerns persist regarding the accuracy of generated bibliographic references. The effect of prompt design on reference reliability has not been clearly established. This comparative experimental study evaluated the impact of prompt specificity on LLM-generated reference accuracy in endodontics and compared model performance. MATERIAL AND METHODS We used ChatGPT 5 and Claude Sonnet 4.6. Ten predefined endodontic queries were combined with 3 prompt types of increasing specificity. Each model generated 5 references per query-prompt combination (total: 300 references). References were verified using PubMed, Google Scholar, and CrossRef. Accuracy was classified as fabricated (0), partially accurate (1; existing references containing ≥1 bibliographic inaccuracy), or fully accurate (2). Digital object identifier (DOI) accuracy was assessed separately. Statistical analyses were performed using mixed-effects models and Pearson's chi-square test or Fisher's exact test. RESULTS Accuracy scores tended to increase with greater prompt specificity (P=0.249). Claude demonstrated significantly higher accuracy than ChatGPT (mean score: 1.79 vs 1.25; P<0.001). DOI accuracy did not differ among prompt groups (P=0.338); it was significantly higher for Claude than for ChatGPT (90.0% vs 35.3%; P<0.001). ChatGPT produced significantly more title, journal, and DOI errors (P<0.001); author and year errors were similar between models. CONCLUSIONS Prompt specificity had limited effects on reference accuracy; model selection played a greater role. DOI accuracy was strongly model-dependent and largely unaffected by prompt design under the test conditions, highlighting the need for external verification of LLM-generated references.

نتیجه فارسی

این مطالعه اثر طراحی پرامپت و انتخاب مدل بر دقت ارجاعات تولید شده توسط LLM در دندان‌پزشکی درمانی را بررسی کرد. Claude دقت کلی و دقت DOI بالاتری نسبت به ChatGPT داشت. ویژگی خاص پرامپت اثر محدودی بر دقت داشت.

  • Claude دقت ارجاعات (۱.۷۹) و دقت DOI (۹۰٪) را به‌طور معناداری بالاتر از ChatGPT (۱.۲۵ و ۳۵.۳٪) نشان داد.
  • دقت ارجاعات با افزایش ویژگی خاص پرامپت افزایش یافت، اما این اثر معنادار نبود (P=0.249).
  • ChatGPT خطاهای بیشتری در عنوان، مجله و DOI تولید کرد.
  • دقت DOI تحت شرایط آزمایش به‌طور کلی تحت تأثیر طراحی پرامپت قرار نگرفت.

ترجمه فارسی چکیده

مدل‌های زبانی بزرگ (LLM) در حال افزایش در مراقبت‌های بهداشتی استفاده می‌شوند، اما نگرانی‌هایی در مورد دقت ارجاعات کتابشناختی تولید شده وجود دارد. اثر طراحی پرامپت بر قابلیت اطمینان ارجاعات به طور واضح مشخص نشده است. این مطالعه تجربی مقایسه‌ای اثر ویژگی‌های خاص پرامپت بر دقت ارجاعات LLM در دندان‌پزشکی درمانی را ارزیابی کرد و عملکرد مدل‌ها را مقایسه نمود. از ChatGPT 5 و Claude Sonnet 4.6 استفاده شد. ده پرسش دندان‌پزشکی پیش‌تعریف شده با ۳ نوع پرامپت با افزایش ویژگی خاص ترکیب شدند. هر مدل ۵ ارجاع برای هر ترکیب پرسش-پرامپت تولید کرد (مجموع: ۳۰۰ ارجاع). ارجاعات با استفاده از PubMed، Google Scholar و CrossRef تأیید شدند. دقت به عنوان جعلی (۰)، به‌طور جزئی دقیق (۱؛ ارجاعات موجود شامل ≥۱ خطای کتابشناختی) یا کاملاً دقیق (۲) طبقه‌بندی شد. شناسه شیء دیجیتال (DOI) به طور جداگانه ارزیابی شد. تحلیل‌های آماری با استفاده از مدل‌های اثرات مخلوط و آزمون کای-مربع پیرسون یا آزمون دقیق فیشر انجام شد. امتیازات دقت با افزایش ویژگی خاص پرامپت افزایش یافت (P=0.249). Claude دقت به‌طور معناداری بالاتری نسبت به ChatGPT نشان داد (امتیاز میانگین: ۱.۷۹ در برابر ۱.۲۵؛ P<0.001). دقت DOI بین گروه‌های پرامپت تفاوتی نداشت (P=0.338)؛ اما برای Claude به‌طور معناداری بالاتر از ChatGPT بود (۹۰.۰٪ در برابر ۳۵.۳٪؛ P<0.001). ChatGPT خطاهای بیشتری در عنوان، مجله و DOI تولید کرد (P<0.001)؛ خطاهای نویسنده و سال بین مدل‌ها مشابه بود. نتیجه‌گیری: ویژگی خاص پرامپت اثر محدودی بر دقت ارجاعات داشت؛ انتخاب مدل نقش بیشتری ایفا کرد. دقت DOI به شدت به مدل وابسته بود و تحت شرایط آزمایش به‌طور کلی تحت تأثیر طراحی پرامپت قرار نگرفت، که نیاز به تأیید خارجی ارجاعات LLM را برجسته می‌کند.

روش پژوهش

این مطالعه تجربی مقایسه‌ای با استفاده از ChatGPT 5 و Claude Sonnet 4.6 انجام شد. ده پرسش دندان‌پزشکی با ۳ نوع پرامپت ترکیب شدند و هر مدل ۵ ارجاع تولید کرد (مجموع ۳۰۰ ارجاع). ارجاعات با PubMed، Google Scholar و CrossRef تأیید شدند.

محدودیت‌ها

محدودیت‌های متن کامل در دسترس نیستند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
ارجاعات کتابشناختی مرتبط با دندان‌پزشکی درمانی.
مداخله/مواجهه
تولید ارجاعات توسط LLM با استفاده از پرامپت‌های مختلف.
مقایسه
مقایسه عملکرد ChatGPT 5 و Claude Sonnet 4.6.
حجم نمونه
۳۰۰ ارجاع تولید شده.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Different irrigation solutions in calcium hydroxide removal: Do high-frequency ultrasonic and laser activation make the difference?

This study aimed to comparatively evaluate the effectiveness of different irrigation solutions and activation techniques in removing calcium hydroxide (CH) from standardized artificial grooves located in the apical third of root canals. Ninety extracted single-rooted human maxillary central incisors were instrumented using a standardized preparation protocol. Artificial grooves were created in the apical third of split roots and filled…

PubMed2026

Indications for endodontic surgery in private practice and temporal trends before and after adoption of laser-activated irrigation: a retrospective practice-based study.

OBJECTIVES: To characterize indications for endodontic microsurgery (EMS) and compare the relative proportion of these indications across two calendar periods before and after the practice's adoption of laser-activated irrigation (LAI). MATERIALS AND METHODS: A retrospective time-period comparison of consecutive EMS cases (n = 1,672) treated by six endodontists (April 2019-August 2025) was analyzed. Two periods were prespecified based …

PubMed2026

Short-term radiographic changes of extruded NeoSealer Flo in a simulated enlarged-apex model: influence of moisture and periapical support.

OBJECTIVES: To radiographically evaluate short-term changes in NeoSealer Flo extruded beyond enlarged apical preparations under different moisture and apical-support conditions. MATERIALS AND METHODS: Thirty-two extracted single-rooted human teeth were randomly allocated to four experimental conditions (n = 8 per group). Teeth were prepared with reciprocating NiTi instruments to size 50/0.05 and filled with a matched gutta-percha cone …

PubMed2026

Effect of Root Canal Dressings on Extraradicular pH in a Sealed Apex Model: An In Vitro Study.

OBJECTIVES: Endodontic treatment is mandatory after severe dental trauma such as avulsion or intrusion. Initial intracanal dressing with calcium hydroxide should be avoided according to international guidelines due to potential periodontal ligament damage by pH elevation. However, whether an intracanal dressing is able to alter the extraradicular site remains unclear. This in vitro study quantified extraradicular pH changes through den…