PubMed چکیده/رکورد

Generative artificial intelligence to augment ethical problem solving in ophthalmology: GPT-5.1 versus a human ethicist.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

PURPOSE: While many large language models (LLMs) have been extensively investigated for their clinical decision-making capabilities, few studies have characterized their abilities to reason through complex, open-ended ethics cases. This study compared GPT-5.1 and expert human ethicist responses to real-world ethical scenarios specific to ophthalmology. METHODS: Ten ethical scenarios from the American Academy of Ophthalmology's Ask the Ethicist website were randomly selected and presented to GPT-5.1 using ChatGPT. AI-generated responses were subsequently compared to those of the expert human ethicist using conventional readability metrics. A panel of 10 physicians independently rated all responses via 6-point and 5-point Likert scales for both outcomes of likelihood-of-use in their own careers and perceived patient impact, respectively. RESULTS: GPT-5.1 versus human ethicist responses differed significantly on Flesch Reading Ease (10.9 ± 9.6 vs. 25.4 ± 10.9, p = 0.002) and Gunning Fog Index (21.3 ± 1.9 vs. 19.5 ± 2.9, p = 0.037). For likelihood-of-use, median ratings were 4.1 [3.8-4.3] for human ethicist versus 5.0 [4.7-5.2] for GPT-5.1 responses (p = 0.013). Median ratings for perceived patient impact of human ethicist versus GPT-5.1 responses were 3.6 [3.3-3.7] versus 3.9 [3.8-4.4], p = 0.008. Inter-rater reliability was moderate for both outcomes and response sources (ICC [2, 10]: 0.56-0.69). CONCLUSIONS: Compared to human ethicist responses, GPT-5.1 responses demonstrated lower readability; however, GPT-5.1 responses received significantly higher ratings for both outcomes of likelihood-of-use and perceived patient impact, respectively. These results advocate for further explorations of GPT-5.1 and other LLMs as potentially useful tools for evaluating common bioethical challenges in ophthalmology.

نتیجه فارسی

این مطالعه GPT-5.1 را در حل مسائل اخلاقی چشم‌پزشکی با یک متخصص انسانی مقایسه کرد. اگرچه GPT-5.1 پاسخ‌های کمتر خوانا تولید کرد، اما رتبه‌های بالاتری برای کاربرد بالینی و تأثیر بر بیمار دریافت کرد. نتایج نشان می‌دهد که LLMs می‌توانند ابزارهای مفیدی برای این کار باشند.

  • GPT-5.1 در برابر یک متخصص انسانی در حل مسائل اخلاقی چشم‌پزشکی مقایسه شد.
  • GPT-5.1 پاسخ‌های کمتر خوانا تولید کرد اما رتبه‌های بالاتری برای کاربرد و تأثیر بر بیمار دریافت کرد.
  • پاسخ‌های GPT-5.1 برای احتمال استفاده در حرفه پزشکان و تأثیر درک شده بر بیمار معنی‌داراً بالاتر بودند.
  • روایی بین نمره‌دهندگان برای هر دو خروجی متوسط بود.
  • این مطالعه پیشنهاد می‌کند که LLMs می‌توانند ابزارهای مفیدی برای ارزیابی چالش‌های اخلاقی باشند.

ترجمه فارسی چکیده

هدف: در حالی که بسیاری از مدل‌های زبانی بزرگ (LLM) برای توانایی‌های تصمیم‌گیری بالینی مورد بررسی قرار گرفته‌اند، تعداد کمی از مطالعات توانایی آن‌ها برای استدلال در مورد موارد اخلاقی پیچیده و باز را توصیف کرده‌اند. این مطالعه پاسخ‌های GPT-5.1 و یک متخصص اخلاق‌دان انسانی را در برابر سناریوهای اخلاقی واقعی مرتبط با چشم‌پزشکی مقایسه کرد. روش‌ها: ده سناریوی اخلاقی از وب‌سایت Ask the Ethicist آکادمی چشم‌پزشکی آمریکا به صورت تصادفی انتخاب شد و به GPT-5.1 در چت‌جی‌پی‌تی ارائه شد. پاسخ‌های تولید شده توسط هوش مصنوعی با پاسخ‌های متخصص انسانی با استفاده از شاخص‌های خوانایی استاندارد مقایسه شدند. یک پنل از ۱۰ پزشک پاسخ‌ها را به طور مستقل با مقیاس‌های لیکرت ۶ و ۵ نمره‌گذاری کردند: یکی برای احتمال استفاده در حرفه خود و دیگری برای تأثیر درک شده بر بیمار. نتایج: پاسخ‌های GPT-5.1 در برابر متخصص انسانی تفاوت معنی‌داری در شاخص خوانایی Flesch (10.9 ± 9.6 در برابر 25.4 ± 10.9، p = 0.002) و شاخص مهندسی Gunning (21.3 ± 1.9 در برابر 19.5 ± 2.9، p = 0.037) داشتند. برای احتمال استفاده، میانگین رتبه‌ها 4.1 [3.8-4.3] برای متخصص انسانی و 5.0 [4.7-5.2] برای پاسخ‌های GPT-5.1 بود (p = 0.013). میانگین رتبه‌ها برای تأثیر درک شده بر بیمار متخصص انسانی در برابر GPT-5.1 3.6 [3.3-3.7] در برابر 3.9 [3.8-4.4] بود، p = 0.008. روایی بین نمره‌دهندگان برای هر دو خروجی و منبع پاسخ متوسط بود (ICC [2, 10]: 0.56-0.69). نتیجه‌گیری: در برابر پاسخ‌های متخصص انسانی، پاسخ‌های GPT-5.1 خوانایی کمتری داشتند؛ با این حال، پاسخ‌های GPT-5.1 برای هر دو خروجی احتمال استفاده و تأثیر درک شده بر بیمار رتبه‌های معنی‌داری بالاتری دریافت کردند. این نتایج پیشنهاد می‌کنند که GPT-5.1 و سایر LLMs به عنوان ابزارهای بالقوه مفید برای ارزیابی چالش‌های بیواخلاقی رایج در چشم‌پزشکی مورد بررسی قرار گیرند.

روش پژوهش

ده سناریوی اخلاقی از Ask the Ethicist آکادمی چشم‌پزشکی آمریکا به صورت تصادفی انتخاب شدند. پاسخ‌های GPT-5.1 با پاسخ‌های متخصص انسانی با استفاده از شاخص‌های خوانایی Flesch و Gunning Fog مقایسه شدند. یک پنل از ۱۰ پزشک پاسخ‌ها را با مقیاس‌های لیکرت نمره‌گذاری کردند.

محدودیت‌ها

محدودیت‌های گزارش نشده در متن موجود است.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
پزشکان (پنل ۱۰ نفره) و متخصص اخلاق‌دان انسانی.
مداخله/مواجهه
GPT-5.1 (در چت‌جی‌پی‌تی) برای حل سناریوهای اخلاقی.
مقایسه
متخصص اخلاق‌دان انسانی.
حجم نمونه
۱۰ سناریوی اخلاقی، ۱۰ متخصص انسانی، ۱۰ پزشک نمره‌دهنده.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

Artificial intelligence (AI)Ethical reasoningEthics in ophthalmologyGPT-5.1Large language models (LLMs)
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and enviro…

PubMed2026

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated respons…

PubMed2026

Liability and Standard of Care in AI-Driven Psychiatric Practice: European Viewpoint.

AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management…

PubMed2026

Routine laboratory panels classify internal medicine ICD-10 code groups: comparison with frontier large language models and laboratory-only specialist assessment.

INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from …