PubMed دسترسی آزاد

Periodontitis Risk Assessment and Prevention Planning: Comparative Study of Multimodal Large Language Models and Periodontist Evaluations.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Periodontitis is one of the most prevalent yet preventable oral diseases, as indicated by multiple clinical and radiographic factors. As these factors are recorded in electronic health records (EHRs), their reuse offers opportunities for personalized risk assessment and targeted prevention. Predictive AI and traditional machine learning models support fragmented detection tasks but lack the integration of textual and imaging predictors. Emerging multimodal large language models (M-LLMs) show promise in combining these data sources for clinical assessment. Evaluating the capabilities of M-LLMs and comparing them against the current clinical standard are therefore essential to determine their potential as digital assistants. OBJECTIVE: This study aimed to evaluate the ability of M-LLMs to assess periodontitis risk and suggest prevention strategies, based on EHR data and radiographic findings. Each M-LLM was individually evaluated by periodontal experts, benchmarked against other models, and compared with a periodontist as a reference. METHODS: A vignette study was conducted following TRIPOD (Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis) guidelines for the evaluation of LLMs. Ten periodontal vignettes were created, each including a panoramic radiograph and textual EHR data. Three LLMs capable of reasoning and handling multimodal data were compared to a periodontist who generated outputs manually, based on the same prompts and input data. Periodontal experts rated all outputs across 6 predefined criteria on a 5-point Likert scale. Statistical analyses evaluated overall performance per model and tested whether performance varied per model, scenario complexity, or rater. RESULTS: GPT o1 Pro and Claude Sonnet 4 showed strong performance, with 86.7% and 85.6% of ratings deemed acceptable-comparable to the periodontist's output (87.8%). Gemini 2.5 Pro was rated significantly lower than both the periodontist and the other models (59.4% acceptable; P<.002). Radiographic interpretation consistently received lower scores than other abilities across all models and the periodontist, with Gemini rated below the acceptable threshold. The time required for completion ranged from approximately 10 seconds for Claude to 37 seconds for Gemini; 3 minutes, 22 seconds, for GPT; and 5 minutes, 57 seconds, for the periodontist. CONCLUSIONS: M-LLMs demonstrated strong reasoning abilities in periodontal assessment. Across all models, unacceptable elements were consistently related to errors in radiographic interpretation, though refined prompting or newer model versions may improve this. Notably, even when radiographic findings were incorrect and plaque-retentive factors were absent, outputs were still rated well, indicating that EHR data alone provide a substantial basis. For clinical applicability, M-LLMs must at least perform comparably to a periodontist and meet the quality standards set by periodontal experts-a bar that GPT and Claude appear to approach.

نتیجه فارسی

این مطالعه مقایسه‌ای توانایی مدل‌های زبانی بزرگ چندوجهی (M-LLM) را در ارزیابی ریسک پریودنتیت و پیشنهاد استراتژی‌های پیشگیری بر اساس داده‌های سوابق الکترونیک سلامت و رادیوگرافی بررسی کرد. سه مدل M-LLM با یک پریودنتیست مقایسه شدند. نتایج نشان داد GPT o1 Pro و Claude Sonnet 4 عملکردی مشابه پریودنتیست داشتند، در حالی که Gemini 2.5 Pro عملکرد ضعیف‌تری نشان داد. خطاهای اصلی در تفسیر رادیوگرافیک بود.

  • M-LLMها توانایی استدلال قوی در ارزیابی پریودنتیت نشان دادند.
  • GPT o1 Pro و Claude Sonnet 4 عملکردی قابل مقایسه با پریودنتیست داشتند.
  • Gemini 2.5 Pro عملکرد پایین‌تری نسبت به سایر مدل‌ها و پریودنتیست داشت.
  • خطاهای اصلی در تفسیر رادیوگرافیک بود.
  • داده‌های EHR به تنهایی می‌توانند پایه‌ای قابل توجه برای ارزیابی باشند.

ترجمه فارسی چکیده

پریودنتیت یکی از شایع‌ترین اما قابل پیشگیری بیماری‌های دهانی است. مدل‌های هوش مصنوعی پیش‌بین و یادگیری ماشین کلاسیک وظایف تشخیصی پراکنده را پشتیبانی می‌کنند اما در ادغام داده‌های متنی و تصویری ناکارآمدند. مدل‌های زبانی بزرگ چندوجهی (M-LLM) در ترکیب این منابع داده برای ارزیابی بالینی امیدوارکننده هستند. این مطالعه هدف دارد توانایی M-LLMها در ارزیابی ریسک پریودنتیت و پیشنهاد استراتژی‌های پیشگیری را بر اساس داده‌های سوابق الکترونیک سلامت (EHR) و یافته‌های رادیوگرافیک بررسی کند. هر مدل M-LLM توسط متخصصان پریودنتیت ارزیابی شد و با یک پریودنتیست به عنوان مرجع مقایسه گردید. این مطالعه شامل ۱۰ سناریوی پریودنتیت شامل رادیوگرافی پانوراما و داده‌های متنی EHR بود. سه مدل M-LLM قادر به استدلال و پردازش داده‌های چندوجهی با یک پریودنتیست که خروجی‌ها را دستی تولید می‌کرد، مقایسه شدند. متخصصان پریودنتیت خروجی‌ها را بر اساس ۶ معیار تعریف شده در مقیاس لیکرت ۵ نمره‌ای ارزیابی کردند. نتایج نشان داد GPT o1 Pro و Claude Sonnet 4 عملکرد قوی داشتند (۸۶.۷٪ و ۸۵.۶٪ نمرات قابل قبول) که مشابه خروجی پریودنتیست (۸۷.۸٪) بود. Gemini 2.5 Pro نمرات پایین‌تری دریافت کرد (۵۹.۴٪ قابل قبول؛ P<.002). تفسیر رادیوگرافیک در تمام مدل‌ها و پریودنتیست نمرات پایین‌تری نسبت به سایر توانایی‌ها داشت. زمان تکمیل کار بین حدود ۱۰ ثانیه برای Claude تا ۳۷ ثانیه برای Gemini متغیر بود.

روش پژوهش

این مطالعه یک مطالعه سناریویی (vignette study) با رعایت دستورالعمل‌های TRIPOD برای ارزیابی مدل‌های زبانی بزرگ انجام شد. ده سناریوی پریودنتیت شامل رادیوگرافی پانوراما و داده‌های متنی EHR ایجاد شد. سه مدل M-LLM قادر به استدلال و پردازش داده‌های چندوجهی با یک پریودنتیست مقایسه شدند. متخصصان پریودنتیت خروجی‌ها را بر اساس ۶ معیار تعریف شده در مقیاس لیکرت ۵ نمره‌ای ارزیابی کردند.

محدودیت‌ها

محدودیت‌های گزارش شده شامل عدم وجود داده‌های واقعی بیماران و استفاده از سناریوهای مصنوعی است. همچنین، مقایسه با یک پریودنتیست به عنوان مرجع ممکن است محدودیت‌هایی داشته باشد.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
سناریوهای پریودنتیت شامل رادیوگرافی پانوراما و داده‌های متنی EHR.
مداخله/مواجهه
ارزیابی ریسک پریودنتیت و پیشنهاد استراتژی‌های پیشگیری توسط مدل‌های زبانی بزرگ چندوجهی (M-LLM).
مقایسه
ارزیابی توسط پریودنتیست و سایر مدل‌های M-LLM.
حجم نمونه
ده سناریوی پریودنتیت.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

benchmarkingclinical decision supportclinical specialistelectronic health recordsgenerative AIlarge language modelmultimodal AIperiodontitisradiographs
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …