Large Language Models for Traumatic Dental Injuries Across Web-Based and Mobile-Based Interfaces: Assessing Accuracy, Quality, and Temporal Consistency.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
OBJECTIVES: Traumatic dental injuries (TDIs) are frequent in clinical practice and require rapid, guideline-based decisions, yet accessing accurate and reliable information may be challenging. Large language models (LLMs) such as ChatGPT, Gemini, DeepSeek, and Qwen are increasingly used as quick online information tools; however, evidence regarding their accuracy, consistency, and the influence of different user interfaces is limited. This study aimed to evaluate the performance of several LLMs in answering TDI-related questions through both web-based interfaces and mobile phone applications. MATERIAL AND METHODS: Twenty questions were prepared according to the 2020 International Association of Dental Traumatology (IADT) guidelines, including 10 open-ended and 10 yes-no items. Four LLMs (ChatGPT-4o, DeepSeek-V3, Gemini 2.0 Flash, Qwen2.5-Max) were queried simultaneously via web and mobile interfaces over five consecutive days, generating 800 responses. Open-ended answers were assessed using the Global Quality Score (GQS) and modified DISCERN (mDISCERN), while yes-no responses were compared with a predetermined answer key. Statistical analyses were performed using IBM SPSS v23.0, with significance set at p < 0.05. RESULTS: Qwen2.5-Max demonstrated comparatively higher GQS and mDISCERN scores across both interfaces. Accuracy for yes-no questions ranged from 86% to 91% without significant differences among models. Interface comparisons showed that ChatGPT-4o generated comparatively higher-quality responses on the web, whereas Qwen2.5-Max performed better on mobile. Over the 5-day period, Qwen2.5-Max showed relatively higher temporal consistency, while DeepSeek-V3 exhibited notable day-to-day variation. CONCLUSIONS: LLMs may serve as useful supplementary tools for providing guideline-based information on TDIs, especially for straightforward, closed-ended clinical questions. However, their performance varies by model, interface, and question type. Qwen2.5-Max demonstrated comparatively higher performance across several evaluated measures. Despite these results, LLM-generated information should be interpreted cautiously and verified by dental professionals before being used in clinical decision-making.
متن کامل اصلی
لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز میشود.
باز کردن متن کامل