PubMed دسترسی آزاد

Performance of GPT-4o and Claude in Medical Ethics Scenarios: Comparative Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: The emergence of AI technology has sparked curiosity regarding the capabilities of large language models (LLMs) in the field of medicine. Minimal research exists regarding the proficiency of various AI models in ethics scenarios, specifically in specialty-based scenarios. OBJECTIVE: This study aimed to compare the performance of GPT-4o and Claude Sonnet 4 on ethics questions with that of medical students and orthopedic residents. METHODS: A total of 200 ethical or legal scenario questions were randomly selected from question banks targeted for third- and fourth-year medical students (UWorld, AMBOSS) and orthopedic residents (OrthoBullets). Questions at the medical student level were exclusively text-based, while resident-level questions included text-based questions accompanied by images. Each question was entered identically into each AI model 3 separate times. If answers varied between trials, the answer provided most frequently by the model was used as the selected answer. RESULTS: GPT-4o correctly answered 140 (70%) of 200 questions, which was similar to the average human test taker score of 71% (~142/200 questions). Claude correctly answered 180 (89%) questions, a score greater than that of human test takers and significantly better than GPT-4o (P<.001). Claude scored significantly higher than GPT-4o in almost all question categories. GPT-4o provided different responses to identically worded trials for 27 (21%) of 130 general questions and 3 (4%) of 70 orthopedic questions (P=.002), while Claude did not have a significant difference in variability between these 2 groups (general: 16/130, 12% vs orthopedic: 3/70, 4%; P=.06). GPT-4o selected the incorrect response for 60 (30%) total questions and chose the incorrect response most commonly selected by humans significantly more frequently on UWorld interpersonal-specific questions (30/40, 75%) than on UWorld all social sciences (27/40, 68%; P=.03). Claude showed no significant difference in the rate of most common incorrect response selection between question categories. CONCLUSIONS: These results suggest that GPT-4o can potentially answer both general and specialty-specific ethical questions with similar proficiency to sample groups of both medical students and orthopedic residents, while Claude AI performs significantly better than both humans and GPT-4o. Variables such as AI model framework and training data may drive the observed difference in performance, but the exact cause cannot be definitively isolated without intentional testing. Therefore, further research is needed to ensure safety by minimizing output variability before integrating AI as a patient-facing resource.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

ChatGPTClaudeartificial intelligenceethicshealthlegalmedical educationmedical studentorthopedicsresident
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Beyond Primum Non Nocere: Therapeutic Injury, Comparative Harm Management, and the Ethics of Modern Medicine.

RATIONALE, AIMS AND OBJECTIVE: The maxim primum non nocere ("first, do no harm") remains emblematic of medical ethics, yet much of modern therapeutics accepts foreseeable and sometimes near-certain injury in pursuit of benefit: surgery injures tissue, chemotherapy produces systemic toxicity, and ablation and embolization achieve their effect through controlled destruction. This paper asks how such therapeutic injury is to be ethically …

PubMed2026

[Ambiguities of medicine in the face of forms of death].

For thirty years, several healthcare systems have been offering medical assistance in dying (MAID) in the form of euthanasia or assisted suicide. This development is transforming medicine's relationship to death. From a legal perspective, MAID does not constitute a right to die, but rather, as a form of conditional decriminalization, an exception to the general prohibition on killing, under certain strict and controlled conditions. Nev…

PubMed2026

Albert Schweitzer: A Reverence for Life, Empathy, and the Medical Vocation.

Albert Schweitzer's life integrated theology, philosophy, music, and medicine into a unified ethical commitment to humanity. Awarded the 1952 Nobel Peace Prize, he articulated a moral vision grounded in deep empathy and respect for all living beings. This article examines the developmental, relational, and cultural influences that shaped his life trajectory, including his upbringing in Alsace, formative experiences of privilege and mor…

PubMed2026

Underdiagnosis of Health: A Key Driver of Overdiagnosis of Disease.

BACKGROUND: In reducing medical overactivity and healthcare waste, there is much focus on reducing overdiagnosis of disease. Surprisingly, there is little attention on the underdiagnosis of health. OBJECTIVE: This article investigates how avoiding underdiagnosis of health can reduce medical overactivity. METHODS: Conceptual analysis drawing philosophy of medicine and medical ethics is used to address five specific research questions: 1…