OsteoCHAT: real-world patient evaluation and benchmarking of a guideline-grounded osteoporosis chatbot.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
UNLABELLED: Patients with osteoporosis have information needs that routine care inconsistently addresses. In this multicenter real-world evaluation, the LLM-based, guideline-grounded chatbot OsteoCHAT showed high participant acceptability and achieved expert-rated answer quality comparable to general-purpose LLMs. Despite challenges in complex treatment and safety-related topics, guideline-grounded chatbots may represent useful educational adjuncts. PURPOSE: Patients with osteoporosis have information needs that routine care inconsistently addresses. Guideline-grounded chatbots may offer scalable educational support METHODS: We developed OsteoCHAT, a retrieval-augmented conversational system based on the current German osteoporosis guideline, and conducted a multicenter evaluation at seven centers in Germany. Participants interacted with OsteoCHAT, provided real-time feedback on individual responses, and completed a standardized questionnaire assessing usability, comprehensibility, usefulness, trust, perceived quality, and preference over conventional internet search. In parallel, OsteoCHAT was benchmarked against three general-purpose large language models (LLMs) using ten osteoporosis frequently asked questions derived from the BfO patient guideline as the gold standard. Five clinical experts independently rated these responses across five domains, each scored 0-3. RESULTS: Overall, 1,075 question-answer interactions were recorded. Of 389 individually rated responses, 379 (97.4%) received positive feedback. Among 217 complete questionnaire responses, 84.3% rated answers as easily understandable, 79.8% found OsteoCHAT easy to use, 77.2% perceived time savings, and 78.7% considered it a useful addition to existing educational materials. Trust was comparatively lower (64.7% agreement). Blinded expert benchmarking showed acceptable-to-high-quality responses across all LLMs (median total scores 14/15 for Gemini, ChatGPT, and OsteoCHAT and 13/15 for Meta AI), with expert-identified weaknesses mainly concerning pharmacological treatment indication and communication of adverse and rare safety-critical events. CONCLUSION: OsteoCHAT demonstrated high user acceptability and usability in a real-world multicenter evaluation, was rated as a valuable addition to existing educational materials, and achieved expert-rated response quality comparable to general-purpose LLMs, although limitations in safety communication were identified across all examined systems.
نتیجه فارسی
OsteoCHAT یک ربات گفتگوی مبتنی بر راهنماي استخوانشکني آلمانی است که در ارزیابی واقعی چندمرکزه، پذيرش بالايي و کاربردپذيري را نشان داد. پاسخهاي آن کیفیت مشابهی با مدلهای زباني عمومی داشت، اما در ارتباطات ایمنی محدودیتهايي وجود داشت.
- OsteoCHAT بر اساس راهنماي آلمانی استخوانشکني ساخته شد.
- پذيرش بالايي (97.4%) و کاربردپذيري در شرکتکنندگان گزارش شد.
- کیفیت پاسخها با مدلهای زباني عمومی مقایسهپذير بود.
- در ارتباطات ایمنی محدودیتهايي شناسایی شد.
- مبتلايان به استخوانشکني به اطلاعات بیشتری نیاز دارند که مراقبتهای عادی فراهم نمیکند.
ترجمه فارسی چکیده
مبتلايان به استخوانشکنی نيازهاي اطلاعاتي دارند که مراقبتهای عادی به طور ناگهانی آنها را برطرف نمیکند. در این ارزیابی چندمرکزه واقعی، ربات گفتگوی مبتنی بر مدل زبانی بزرگ (LLM) و راهنماي OsteoCHAT، پذيرش بالايي از سوي شرکتکنندگان را نشان داد و کیفیت پاسخهاي ارزیابی شده توسط متخصصان را با مدلهای زباني عمومی مقایسهپذير ساخت. با وجود چالشها در موضوعات درماني پیچیده و ایمنی، رباتهاي مبتنی بر راهنما ممکن است مکملهای آموزشي مفیدی باشند.
روش پژوهش
OsteoCHAT یک سیستم گفتگوي مبتنی بر بازیابی است که بر اساس راهنماي آلمانی استخوانشکني ساخته شد. ارزیابی در هفت مرکز آلمان انجام شد و شرکتکنندگان با ربات تعامل کردند و بازخورد دادند. همچنین ربات با سه مدل زباني عمومی مقایسه شد.
محدودیتها
در ارتباطات ایمنی محدودیتهايي در تمام سیستمهاي بررسی شده وجود داشت. اعتماد شرکتکنندگان نسبتاً پایینتر بود. پاسخهاي ربات در موضوعات پیچیده درمان و ایمنی چالشبرانگیز بودند.
نمای PICO و پیامدها
- جمعیت
- مبتلايان به استخوانشکني
- مداخله/مواجهه
- OsteoCHAT (ربات گفتگوي مبتنی بر راهنما)
- مقایسه
- جستجوی اینترنتی متداول و مدلهای زباني عمومی (Gemini, ChatGPT, Meta AI)
- حجم نمونه
- 1,075 تعامل سوال و پاسخ; 389 پاسخ ارزیابی شده; 217 پرسشنامه کامل
متن کامل اصلی
برای بررسی دسترسی کتابخانهای یا خرید، رکورد اصلی را باز کنید.
رفتن به منبع اصلی