PubMed دسترسی آزاد

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response characteristics were analysed using category-based coding and a six-dimensional LLM-as-a-judge assessment, with LLM-based ratings compared with rheumatologist ratings in random subsets.Between September 2025 and January 2026, 6291 questions were recorded. Thirteen question categories were identified, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%). The chatbots were unable to answer in 263 interactions (4.2%). Of 2671 responses rated by users, 2481 (92.9%) received a positive rating. Insufficient detail was the most common reason for negative ratings (125/190, 65.8%). Among 602 questionnaire respondents, 84.6% reported that the chatbot was easy to use, 84.1% that answers were easy to understand, and 80.2% that it was a useful addition to patient education. In the LLM-based evaluation, 95.3% of answers were rated as completely safe and 79.1% as completely correct. Guideline adherence was assessed separately, with 45.0% rated as fully adherent; agreement with physician assessment was weak.Guideline-grounded chatbots received predominantly positive user feedback in real-world use, while LLM-based evaluation suggested that most responses were safe and correct. User questions and feedback may help guide iterative improvements to source content and patient education materials. Further studies are needed to evaluate educational effectiveness and independently validate response quality and clinical safety.

نتیجه فارسی

ده چت‌بات اختصاصی مبتنی بر راهنما در ۱۳ مرکز روماتولوژی توسعه یافتند. در یک دوره ۵ ماهه، ۶۲۹۱ سوال ثبت شد و اکثر پاسخ‌ها (۹۲.۹٪) توسط کاربران مثبت ارزیابی شدند. ارزیابی‌های LLM نشان دادند که اکثر پاسخ‌ها (۹۵.۳٪) ایمن و (۷۹.۱٪) صحیح بودند، اما همسویی با نظر پزشکان ضعیف بود.

  • توسعه ده چت‌بات اختصاصی مبتنی بر راهنماهای آلمانی در ۱۳ مرکز روماتولوژی.
  • ۶۲۹۱ سوال توسط کاربران ثبت شد و اکثر پاسخ‌ها (۹۲.۹٪) امتیاز مثبت دریافت کردند.
  • ارزیابی‌های LLM نشان دادند که ۹۵.۳٪ پاسخ‌ها ایمن و ۷۹.۱٪ صحیح بودند.
  • همسویی با نظر پزشکان ضعیف بود.
  • نیاز به مطالعات بیشتر برای تأیید مستقل ایمنی بالینی و اثربخشی آموزشی.

ترجمه فارسی چکیده

بیماران مبتلا به بیماری‌های روماتولوژیک نیازهای اطلاعاتی پایداری دارند که در مراقبت‌های استاندارد به‌طور کامل برطرف نمی‌شوند. ما مدل‌های زبانی بزرگ (LLM) مبتنی بر راهنما را برای پشتیبانی از خودمدیریت و آموزش بیماران توسعه دادیم و ارزیابی کردیم. ده چت‌بات اختصاصی بر اساس راهنماهای آلمانی در ۱۳ مرکز روماتولوژی و شش سازمان بیمار توسعه و پیاده‌سازی شدند. کاربران چت‌بات‌ها پاسخ‌ها را ارزیابی کردند. سوالات، بازخوردها و ویژگی‌های پاسخ با کدگذاری مبتنی بر دسته و ارزیابی شش‌بعدی LLM-as-a-judge تحلیل شدند. در بازه سپتامبر ۲۰۲۵ تا ژانویه ۲۰۲۶، ۶۲۹۱ سوال ثبت شد. ۱۳ دسته سوال شناسایی شد، شامل سوالات اختصاصی بیماری (۵۰.۲٪)، دارو و نظارت (۳۷.۳٪) و تشخیص (۲۸.۲٪). چت‌بات‌ها در ۲۶۳ تعامل (۴.۲٪) نتوانستند پاسخ دهند. از ۲۶۷۱ پاسخ ارزیابی شده توسط کاربران، ۲۴۸۱ (۹۲.۹٪) امتیاز مثبت دریافت کردند. کمبود جزئیات رایج‌ترین دلیل امتیاز منفی بود. از ۶۰۲ پاسخ‌دهنده پرسشنامه، ۸۴.۶٪ گزارش کردند چت‌بات آسان است، ۸۴.۱٪ پاسخ‌ها قابل فهم هستند و ۸۰.۲٪ آن را افزودنی مفید برای آموزش بیماران می‌دانند. در ارزیابی مبتنی بر LLM، ۹۵.۳٪ پاسخ‌ها کاملاً ایمن و ۷۹.۱٪ کاملاً صحیح ارزیابی شدند. اطاعت از راهنما ۴۵.۰٪ به عنوان کاملاً مطابق ارزیابی شد. همسویی با ارزیابی پزشکان ضعیف بود. چت‌بات‌های مبتنی بر راهنما بازخورد مثبت دریافت کردند، در حالی که ارزیابی LLM نشان داد که اکثر پاسخ‌ها ایمن و صحیح هستند. سوالات و بازخوردهای کاربران می‌توانند به بهبود تدریجی محتوای منبع و مواد آموزشی کمک کنند. مطالعات بیشتری برای ارزیابی اثربخشی آموزشی و تأیید مستقل کیفیت پاسخ و ایمنی بالینی مورد نیاز است.

روش پژوهش

ده چت‌بات اختصاصی بر اساس راهنماهای آلمانی در ۱۳ مرکز روماتولوژی و شش سازمان بیمار توسعه و پیاده‌سازی شدند. سوالات و بازخوردها با کدگذاری مبتنی بر دسته و ارزیابی شش‌بعدی LLM-as-a-judge تحلیل شدند.

محدودیت‌ها

ارزیابی همسویی با نظر پزشکان ضعیف بود. مطالعات بیشتری برای تأیید مستقل کیفیت پاسخ و ایمنی بالینی مورد نیاز است.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
بیماران مبتلا به بیماری‌های روماتولوژیک
مداخله/مواجهه
چت‌بات‌های زبانی بزرگ (LLM) مبتنی بر راهنما
مقایسه
ارزیابی با نظر پزشکان و ارزیابی LLM-as-a-judge
حجم نمونه
۶۲۹۱ سوال ثبت شده، ۲۶۷۱ پاسخ ارزیابی شده توسط کاربران، ۶۰۲ پاسخ‌دهنده پرسشنامه

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

Artificial intelligenceChatbotsLarge language modelsPatient educationRetrieval-augmented generationRheumatology
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Community Stakeholder Perspectives on Children With Cerebral Palsy and Their Caregivers: A Qualitative Study From Rural Malawi.

BACKGROUND: Cerebral palsy (CP) disproportionately affects children in low- and middle-income countries, where stigma and discrimination lead to marginalization and reduced participation. To design culturally appropriate community-based health promotion strategies that can address these challenges, the role of community stakeholders is crucial, yet unstudied in rural sub-Saharan Africa. This study aimed to explore stakeholders' percept…

PubMed2026

Implementation of a Multidisciplinary Non-Pharmacological Program to Improve Urinary Incontinence in an Intermediate Care Hospital.

INTRODUCTION: Urinary incontinence (UI) is a prevalent geriatric syndrome that significantly affects the physical, psychological, and social well-being of older adults. Non-pharmacological interventions are recommended as the first-line approach, especially in frail older adults. DESIGN: This was a prospective pre-post observational study. METHODS: The study included sixty-one patients with rehabilitable UI who were admitted to an inte…

PubMed2026

Describing an integrated partnership approach to improving sexual health literacy and service navigation for international students in Sydney, Australia.

BACKGROUND: The population of international students (IS) in New South Wales (NSW) continues to grow post-COVID-19, with IS a key priority within NSW Health HIV and sexually transmissible infection (STI) strategies and efforts. Evidence shows gaps in sexual and reproductive health knowledge (SRH), and barriers to accessing services, including stigma, unfamiliarity with the Australian healthcare system and service cost concerns. Althoug…

PubMed2026

Moving beyond sexual and reproductive health knowledge: an evaluation of a co-designed digital engagement tool to promote sexual and reproductive health among international students in New South Wales, Australia.

BACKGROUND: This study evaluates the perceived impact and needed enhancement of a co-designed digital sexual and reproductive health (SRH) tool ('the Hub') developed to address SRH literacy gaps among international students (IS) in New South Wales, Australia. Despite Australia's large IS population and recognition of SRH as a fundamental human right, many report limited exposure to SRH information and face barriers to navigating health…