PubMed چکیده/رکورد

Evaluating the Reliability, Accuracy, and Readability of ChatGPT-3.5 and GPT-4 in Providing Patient Education Related to Glenohumeral Joint Osteoarthritis.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

OBJECTIVES: Chatbots have been increasingly recognized as modern tools that provide patients with reliable health-related information. This study aimed to evaluate and compare ChatGPT-3.5 and GPT-4's ability to answer glenohumeral osteoarthritis-related questions. METHODS: Fifteen questions were derived from the 2020 AAOS Clinical Practice Guidelines for the Surgical Management of Glenohumeral Joint Osteoarthritis. Questions were categorized into three groups: risk factors, implant/intraoperative considerations, and pain/functional outcomes. ChatGPT-3.5 and GPT-4 were prompted with these questions, and responses were evaluated by four fellowship-trained shoulder and elbow surgeons. Each response was rated on a scale (scores:1-5) based on relevance, accuracy, clarity, completeness, and evidence-based support. Data was analyzed descriptively and statistically to compare the scores between ChatGPT-3.5 and GPT-4. RESULTS: Average score for ChatGPT-3.5 was 19.7/25, with "Risk Factor" prompts achieving the highest mean score. GPT-4 averaged 18.7/25, with "Functional Outcomes" prompts scoring highest. However, there were no statistically significant differences between different prompt themes for GPT-3.5 and GPT-4. "Clarity" category received the highest score for GPT-3.5, while "Relevance" was highest for GPT-4. Both models scored lowest on "Evidence-based" prompts. On the Flesch- Kincaid scale, GPT-3.5 responses had a significantly higher score of 18.3 compared to GPT-4's 15.4, indicating a more difficult reading level in GPT-3.5's responses. CONCLUSION: Both ChatGPT-3.5 and GPT-4 performed adequately in providing well-informed medical responses to patient queries about glenohumeral osteoarthritis. Future chatbot versions should focus on providing evidence-based content through systematic and reliable reviews of literature, in an accessible readable manner.

نتیجه فارسی

این مطالعه توانایی دو مدل هوش مصنوعی ChatGPT-3.5 و GPT-4 در پاسخگویی به سوالات پزشکی بیماران درباره آرتروز مفصل شانه را بررسی کرد. هر دو مدل عملکرد قابل قبولی داشتند، هرچند GPT-3.5 پاسخ‌های خواناتری ارائه داد. تفاوت معناداری بین عملکرد دو مدل در پاسخ‌های مبتنی بر شواهد مشاهده نشد.

  • پرسش‌ها از دستورالعمل‌های AAOS ۲۰۲۰ استخراج شدند.
  • ChatGPT-3.5 نمره میانگین ۱۹.۷/۲۵ و GPT-4 نمره میانگین ۱۸.۷/۲۵ داشتند.
  • GPT-3.5 در دسته‌بندی «وضوح» نمره بالاتری داشت.
  • هر دو مدل در پاسخ‌های مبتنی بر شواهد نمرات پایینی گرفتند.
  • GPT-3.5 پاسخ‌های خواناتری (Flesch-Kincaid ۱۸.۳) نسبت به GPT-4 (۱۵.۴) ارائه داد.

ترجمه فارسی چکیده

هدف این مطالعه ارزیابی و مقایسه توانایی چت‌بات‌های ChatGPT-3.5 و GPT-4 در پاسخگویی به سوالات مربوط به آرتروز مفصل شانه بود. پرسش‌ها از دستورالعمل‌های بالینی ۲۰۲۰ AAOS برای مدیریت جراحی این بیماری استخراج شدند. پاسخ‌ها توسط چهار جراح متخصص شانه و مچ دست ارزیابی شدند. میانگین نمره ChatGPT-3.5 ۱۹.۷ از ۲۵ بود و نمره بالاتر در دسته‌بندی عوامل خطر بود. GPT-4 میانگین ۱۸.۷ از ۲۵ داشت و نمره بالاتر در دسته‌بندی نتایج عملکردی بود. تفاوت‌های معنادار آماری بین تم‌های مختلف پرسش‌ها مشاهده نشد. ChatGPT-3.5 در دسته‌بندی «وضوح» نمره بالاتری داشت و GPT-4 در دسته‌بندی «مرتبط بودن». هر دو مدل در پاسخ‌های مبتنی بر شواهد نمرات پایینی گرفتند. بر اساس مقیاس Flesch-Kincaid، پاسخ‌های ChatGPT-3.5 سطح خواندن سخت‌تری داشتند.

روش پژوهش

پرسش‌ها از دستورالعمل‌های بالینی ۲۰۲۰ AAOS استخراج شدند. پاسخ‌ها توسط چهار جراح متخصص ارزیابی شدند. نمرات بر اساس مرتبط بودن، دقت، وضوح، کامل بودن و پشتیبانی مبتنی بر شواهد تعیین شد.

محدودیت‌ها

محدودیت‌های گزارش نشده‌اند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
بیماران با آرتروز مفصل شانه (گلنوهومرال)
مداخله/مواجهه
پرسش‌های مربوط به آرتروز مفصل شانه (گلنوهومرال)
مقایسه
هیچ مقایسه‌کننده‌ای گزارش نشده است.
حجم نمونه
تعداد نمونه‌ای گزارش نشده است.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

Flesch-Kincaidartificial intelligencechatbothealth informationpatients
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Community Stakeholder Perspectives on Children With Cerebral Palsy and Their Caregivers: A Qualitative Study From Rural Malawi.

BACKGROUND: Cerebral palsy (CP) disproportionately affects children in low- and middle-income countries, where stigma and discrimination lead to marginalization and reduced participation. To design culturally appropriate community-based health promotion strategies that can address these challenges, the role of community stakeholders is crucial, yet unstudied in rural sub-Saharan Africa. This study aimed to explore stakeholders' percept…

PubMed2026

Implementation of a Multidisciplinary Non-Pharmacological Program to Improve Urinary Incontinence in an Intermediate Care Hospital.

INTRODUCTION: Urinary incontinence (UI) is a prevalent geriatric syndrome that significantly affects the physical, psychological, and social well-being of older adults. Non-pharmacological interventions are recommended as the first-line approach, especially in frail older adults. DESIGN: This was a prospective pre-post observational study. METHODS: The study included sixty-one patients with rehabilitable UI who were admitted to an inte…

PubMed2026

Describing an integrated partnership approach to improving sexual health literacy and service navigation for international students in Sydney, Australia.

BACKGROUND: The population of international students (IS) in New South Wales (NSW) continues to grow post-COVID-19, with IS a key priority within NSW Health HIV and sexually transmissible infection (STI) strategies and efforts. Evidence shows gaps in sexual and reproductive health knowledge (SRH), and barriers to accessing services, including stigma, unfamiliarity with the Australian healthcare system and service cost concerns. Althoug…

PubMed2026

Moving beyond sexual and reproductive health knowledge: an evaluation of a co-designed digital engagement tool to promote sexual and reproductive health among international students in New South Wales, Australia.

BACKGROUND: This study evaluates the perceived impact and needed enhancement of a co-designed digital sexual and reproductive health (SRH) tool ('the Hub') developed to address SRH literacy gaps among international students (IS) in New South Wales, Australia. Despite Australia's large IS population and recognition of SRH as a fundamental human right, many report limited exposure to SRH information and face barriers to navigating health…