Periodontitis Risk Assessment and Prevention Planning: Comparative Study of Multimodal Large Language Models and Periodontist Evaluations.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
BACKGROUND: Periodontitis is one of the most prevalent yet preventable oral diseases, as indicated by multiple clinical and radiographic factors. As these factors are recorded in electronic health records (EHRs), their reuse offers opportunities for personalized risk assessment and targeted prevention. Predictive AI and traditional machine learning models support fragmented detection tasks but lack the integration of textual and imaging predictors. Emerging multimodal large language models (M-LLMs) show promise in combining these data sources for clinical assessment. Evaluating the capabilities of M-LLMs and comparing them against the current clinical standard are therefore essential to determine their potential as digital assistants. OBJECTIVE: This study aimed to evaluate the ability of M-LLMs to assess periodontitis risk and suggest prevention strategies, based on EHR data and radiographic findings. Each M-LLM was individually evaluated by periodontal experts, benchmarked against other models, and compared with a periodontist as a reference. METHODS: A vignette study was conducted following TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines for the evaluation of LLMs. Ten periodontal vignettes were created, each including a panoramic radiograph and textual EHR data. Three LLMs capable of reasoning and handling multimodal data were compared to a periodontist who generated outputs manually, based on the same prompts and input data. Periodontal experts rated all outputs across 6 predefined criteria on a 5-point Likert scale. Statistical analyses evaluated overall performance per model and tested whether performance varied per model, scenario complexity, or rater. RESULTS: GPT o1 Pro and Claude Sonnet 4 showed strong performance, with 86.7% and 85.6% of ratings deemed acceptable-comparable to the periodontist's output (87.8%). Gemini 2.5 Pro was rated significantly lower than both the periodontist and the other models (59.4% acceptable; P<.002). Radiographic interpretation consistently received lower scores than other abilities across all models and the periodontist, with Gemini rated below the acceptable threshold. The time required for completion ranged from approximately 10 seconds for Claude to 37 seconds for Gemini; 3 minutes, 22 seconds, for GPT; and 5 minutes, 57 seconds, for the periodontist. CONCLUSIONS: M-LLMs demonstrated strong reasoning abilities in periodontal assessment. Across all models, unacceptable elements were consistently related to errors in radiographic interpretation, though refined prompting or newer model versions may improve this. Notably, even when radiographic findings were incorrect and plaque-retentive factors were absent, outputs were still rated well, indicating that EHR data alone provide a substantial basis. For clinical applicability, M-LLMs must at least perform comparably to a periodontist and meet the quality standards set by periodontal experts-a bar that GPT and Claude appear to approach.
نتیجه فارسی
این مطالعه مقایسهای توانایی مدلهای زبانی بزرگ چندوجهی (M-LLM) را در ارزیابی ریسک پریودنتیت و پیشنهاد استراتژیهای پیشگیری بر اساس دادههای سوابق الکترونیک سلامت و رادیوگرافی بررسی کرد. سه مدل M-LLM با یک پریودنتیست مقایسه شدند. نتایج نشان داد GPT o1 Pro و Claude Sonnet 4 عملکردی مشابه پریودنتیست داشتند، در حالی که Gemini 2.5 Pro عملکرد ضعیفتری نشان داد. خطاهای اصلی در تفسیر رادیوگرافیک بود.
- M-LLMها توانایی استدلال قوی در ارزیابی پریودنتیت نشان دادند.
- GPT o1 Pro و Claude Sonnet 4 عملکردی قابل مقایسه با پریودنتیست داشتند.
- Gemini 2.5 Pro عملکرد پایینتری نسبت به سایر مدلها و پریودنتیست داشت.
- خطاهای اصلی در تفسیر رادیوگرافیک بود.
- دادههای EHR به تنهایی میتوانند پایهای قابل توجه برای ارزیابی باشند.
ترجمه فارسی چکیده
پریودنتیت یکی از شایعترین اما قابل پیشگیری بیماریهای دهانی است. مدلهای هوش مصنوعی پیشبین و یادگیری ماشین کلاسیک وظایف تشخیصی پراکنده را پشتیبانی میکنند اما در ادغام دادههای متنی و تصویری ناکارآمدند. مدلهای زبانی بزرگ چندوجهی (M-LLM) در ترکیب این منابع داده برای ارزیابی بالینی امیدوارکننده هستند. این مطالعه هدف دارد توانایی M-LLMها در ارزیابی ریسک پریودنتیت و پیشنهاد استراتژیهای پیشگیری را بر اساس دادههای سوابق الکترونیک سلامت (EHR) و یافتههای رادیوگرافیک بررسی کند. هر مدل M-LLM توسط متخصصان پریودنتیت ارزیابی شد و با یک پریودنتیست به عنوان مرجع مقایسه گردید. این مطالعه شامل ۱۰ سناریوی پریودنتیت شامل رادیوگرافی پانوراما و دادههای متنی EHR بود. سه مدل M-LLM قادر به استدلال و پردازش دادههای چندوجهی با یک پریودنتیست که خروجیها را دستی تولید میکرد، مقایسه شدند. متخصصان پریودنتیت خروجیها را بر اساس ۶ معیار تعریف شده در مقیاس لیکرت ۵ نمرهای ارزیابی کردند. نتایج نشان داد GPT o1 Pro و Claude Sonnet 4 عملکرد قوی داشتند (۸۶.۷٪ و ۸۵.۶٪ نمرات قابل قبول) که مشابه خروجی پریودنتیست (۸۷.۸٪) بود. Gemini 2.5 Pro نمرات پایینتری دریافت کرد (۵۹.۴٪ قابل قبول؛ P<.002). تفسیر رادیوگرافیک در تمام مدلها و پریودنتیست نمرات پایینتری نسبت به سایر تواناییها داشت. زمان تکمیل کار بین حدود ۱۰ ثانیه برای Claude تا ۳۷ ثانیه برای Gemini متغیر بود.
روش پژوهش
این مطالعه یک مطالعه سناریویی (vignette study) با رعایت دستورالعملهای TRIPOD برای ارزیابی مدلهای زبانی بزرگ انجام شد. ده سناریوی پریودنتیت شامل رادیوگرافی پانوراما و دادههای متنی EHR ایجاد شد. سه مدل M-LLM قادر به استدلال و پردازش دادههای چندوجهی با یک پریودنتیست مقایسه شدند. متخصصان پریودنتیت خروجیها را بر اساس ۶ معیار تعریف شده در مقیاس لیکرت ۵ نمرهای ارزیابی کردند.
محدودیتها
محدودیتهای گزارش شده شامل عدم وجود دادههای واقعی بیماران و استفاده از سناریوهای مصنوعی است. همچنین، مقایسه با یک پریودنتیست به عنوان مرجع ممکن است محدودیتهایی داشته باشد.
نمای PICO و پیامدها
- جمعیت
- سناریوهای پریودنتیت شامل رادیوگرافی پانوراما و دادههای متنی EHR.
- مداخله/مواجهه
- ارزیابی ریسک پریودنتیت و پیشنهاد استراتژیهای پیشگیری توسط مدلهای زبانی بزرگ چندوجهی (M-LLM).
- مقایسه
- ارزیابی توسط پریودنتیست و سایر مدلهای M-LLM.
- حجم نمونه
- ده سناریوی پریودنتیت.
متن کامل اصلی
لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز میشود.
باز کردن متن کامل