PubMed دسترسی آزاد

Artificial intelligence advancements for orthopaedic clinical reasoning: longitudinal assessment of newer models (ChatGPT-5, Grok-3, Gemini 2.5 Flash) compared to clinicians.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

INTRODUCTION: This descriptive study aimed to longitudinally evaluate the performance of contemporary large language models - ChatGPT-5, Gemini 2.5 Flash, and Grok-3 - on orthopaedic clinical multiple-choice tasks, benchmarked against pooled clinician consensus. A secondary aim was to assess whether recent advances in generative AI translated into improved alignment with clinician consensus compared with previous AI models. MATERIALS AND METHODS: A total of 97 multiple-choice clinical cases spanning eight orthopaedic subspecialties were sourced from OrthoBullets and previously benchmarked against aggregated responses from thousands of practising clinicians. Using identical methodology to our 2023 study of ChatGPT-3.5, ChatGPT-4, and Bard, each model was prompted with standardised case stems and response options. The primary outcome was the proportion of AI responses matching the most popular clinician response; secondary analyses assessed agreement within 10% and 20% of clinician consensus, performance on 'controversial' (< 25% margin) questions, and inter-model concordance using Cohen's kappa coefficients. RESULTS: Gemini 2.5 Flash achieved the highest alignment with clinician consensus (69.1%), followed by Grok-3 (66.0%) and ChatGPT-5 (58.8%). None of the LLMs refused to respond to any prompts, representing a reduction from 7.2% from our 2023 study. Subspecialty analysis demonstrated that Gemini 2.5 Flash performed best in Hand and Paediatric domains, while Grok-3 excelled in Reconstruction, Trauma, and 'controversial' cases. Inter-model agreement was highest between Grok-3 and Gemini 2.5 Flash (κ = 0.678), indicating improved consistency compared with prior-generation systems. CONCLUSIONS: Contemporary LLMs can be promising adjuncts for orthopaedic education by simulating peer reasoning and offering structured explanations in non-critical settings. Despite incremental gains in reasoning capability compared to previous AI models, contemporary LLMs remain unsuitable for independent clinical use. Future research should develop hybrid clinician-AI workflows and longitudinal benchmarks to distinguish true reasoning improvements from memorisation.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

BenchmarkingClinician consensusGenerative AILarge language modelsOrthopaedic decision-making
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Sustainable metallic biomaterials for orthopaedic implants: a comprehensive review of biodegradable and conventional metals.

The selection of biomaterial is crucial for the long-term success of implants. Materials that perform an adequate function and reduce negative biological responses should be taken. Due to their good mechanical strength, stainless steel, titanium, and Co-based alloys have been utilized for implant purposes; however, their permanent nature and very low corrosion rates may lead to long-term clinical complications. Researchers are looking …

PubMed2026

Artificial intelligence meets pediatric orthopedics: A comparative analysis of ChatGPT-4o, Gemini 2.0, and Claude 3.5 in detecting supracondylar humeral fractures.

BACKGROUND: Supracondylar humeral fractures constitute 10-16% of pediatric skeletal injuries, requiring timely diagnosis to prevent neurovascular complications. Developmental variations in pediatric bone structures pose diagnostic challenges for clinicians. This study evaluated three next-generation large language models (LLMs) (ChatGPT-4o, Gemini 2.0, Claude 3.5) for detecting pediatric supracondylar humeral fractures and their classi…

PubMed2026

Leaving orthopaedic surgical training: the LOST surgeons - a qualitative study exploring why UK trauma and orthopaedic registrars discontinue surgical training.

OBJECTIVES: The study aimed to explore why trauma and orthopaedic registrars decide to discontinue surgical training. Understanding the factors that influence a decision to leave may help to inform changes which enhance the experiences of surgeons and retention of the future workforce. DESIGN: Qualitative study using semi-structured interviews. SETTING: Between October 2022 and March 2026, interviews were conducted with participants wh…

PubMed2026

e-Learning, Distance Education, and Virtual and Augmented Reality in Orthopedic Training: European Cross-Sectional Survey of Trainee Acceptance Guided by the Technology Acceptance Model and Unified Theory of Acceptance and Use of Technology.

BACKGROUND: Digital technologies increasingly shape postgraduate medical education, yet orthopedic and trauma training face unique challenges because of the tactile, procedurally focused skills involved. Digital tools partially address these needs, but gaps remain, particularly across diverse European contexts. OBJECTIVE: Our primary aim was to quantitatively assess predictors of digital learning technology acceptance (e-learning, dist…