PubMed دسترسی آزاد

Comparative performance of large language models for appraising bias in real-world evidence studies.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Real-world evidence (RWE) is increasingly used to inform regulatory and payer policy decisions and health technology assessment, yet appraising the methodological credibility of RWE studies remains time-intensive and requires specialized expertise. The appraisal task could involve using an appraisal tool that provides a structured approach for evaluating bias in observational studies of comparative effectiveness and safety. Large language models (LLMs) may offer a scalable means to support this appraisal process, but their performance on structured bias assessment tasks has not been fully characterized. OBJECTIVE: To compare the performance of LLMs from 6 major artificial intelligence (AI) technology providers against human expert assessments in appraising bias in published RWE studies using the Appraisal of Potential Bias in Real-World Evidence Studies framework. METHODS: We conducted a comparative diagnostic accuracy study evaluating 40 LLMs from OpenAI, Anthropic, Google, xAI, Meta, and DeepSeek. Ten published RWE studies representing diverse pharmacoepidemiological designs and data sources were appraised by each LLM using a structured chain-of-thought prompt with conditional rubric injection based on the Appraisal of Potential Bias in Real-World Evidence Studies framework. Two independent human reviewers with pharmacoepidemiology training evaluated each study, with a third adjudicator resolving disagreements to establish the reference standard. LLM performance was assessed using overall accuracy and macro-averaged precision, recall, and F1 scores. Assessment time was compared between models and benchmarked against human reviewers. Bootstrap method was used to construct 95% CI for performance measures. RESULTS: Across 280 item-level assessments per model (10 studies × 28 items), overall accuracy ranged from 12.9% to 66.1%. The highest-performing model was Claude-Sonnet-4.6 (66.1%), followed by o3 (65.4%) and Gemini-3.1-pro-preview (65.0%). Macro-averaged F1 scores ranged from 30.9% to 66.9%; o3 achieved the highest F1 score (66.9%), followed by GROK-4 (65.5%) and Gemini-3.1-pro-preview (65.4%). Human reviewers required an average of 61.05 minutes per study; all LLMs completed assessments substantially faster, with average time per study ranging from 0.80 to 17.22 minutes relative to humans. CONCLUSIONS: LLMs hold considerable promise for automating methodological appraisal of RWE studies; however, their performance is variable and model dependent. Their greatest value may lie in enhancing efficiency and supporting human-led appraisal as decision-support tools rather than replacing expert review. Future research should assess performance across larger, more diverse RWE study collections and evaluate output reproducibility across repeated runs.

نتیجه فارسی

این مطالعه عملکرد ۴۰ مدل زبانی بزرگ را در ارزیابی سوگیری در مطالعات شواهد واقعی با متخصصان انسانی مقایسه کرد. دقت کلی مدل‌ها بین ۱۲.۹٪ تا ۶۶.۱٪ متغیر بود و مدل‌ها به طور قابل توجهی سریع‌تر از انسان‌ها عمل کردند. LLM‌ها به عنوان ابزارهای پشتیبانی انسانی ارزش بالایی دارند، اما جایگزین ارزیابی متخصصان نیستند.

  • ۴۰ مدل LLM از ۶ شرکت هوش مصنوعی در ارزیابی ۱۰ مطالعه RWE مقایسه شدند.
  • دقت کلی مدل‌ها بین ۱۲.۹٪ تا ۶۶.۱٪ متغیر بود.
  • مدل Claude-Sonnet-4.6 بالاترین دقت کلی (۶۶.۱٪) را داشت.
  • مدل o3 بالاترین امتیاز F1 (۶۶.۹٪) را داشت.
  • مدل‌ها زمان ارزیابی را به طور قابل توجهی کاهش دادند.

ترجمه فارسی چکیده

شواهد واقعی (RWE) به طور فزاینده‌ای برای تصمیم‌گیری‌های سیاستی نظارتی و پرداخت و ارزیابی فناوری سلامت استفاده می‌شود، اما ارزیابی اعتبار روش‌شناختی مطالعات RWE همچنان زمان‌بر و نیازمند تخصص ویژه است. مدل‌های زبانی بزرگ (LLM) ممکن است ابزاری مقیاس‌پذیر برای پشتیبانی از این فرآیند ارزیابی باشند، اما عملکرد آن‌ها در تسک‌های ساختاریافته ارزیابی سوگیری هنوز به طور کامل مشخص نشده است. هدف این مطالعه مقایسه عملکرد LLM‌های ۶ ارائه‌دهنده اصلی هوش مصنوعی در برابر ارزیابی‌های متخصصان انسانی در ارزیابی سوگیری در مطالعات RWE منتشرشده با استفاده از چارچوب ارزیابی پتانسیل سوگیری در مطالعات شواهد واقعی است. این مطالعه یک مطالعه تشخیصی مقایسه‌ای بود که ۴۰ مدل LLM از OpenAI، Anthropic، Google، xAI، Meta و DeepSeek را ارزیابی کرد. هر مدل ۱۰ مطالعه RWE منتشرشده را با استفاده از یک فرآیند زنجیره‌ای تفکر ساختاریافته و با استفاده از یک چک‌لیست شرطی بر اساس چارچوب ارزیابی پتانسیل سوگیری ارزیابی کرد. دو نویسنده انسانی مستقل با تخصص اپیدمیولوژی دارویی هر مطالعه را ارزیابی کردند و یک قاضی سوم اختلافات را حل کرد تا استاندارد مرجع ایجاد شود. عملکرد LLM با دقت کلی و میانگین ماکروی دقت،recall و امتیاز F1 ارزیابی شد. زمان ارزیابی بین مدل‌ها مقایسه و با نویسندگان انسانی مقایسه شد. روش بوت‌استرپ برای ساخت ۹۵٪ CI برای معیارهای عملکرد استفاده شد.

روش پژوهش

این مطالعه یک مطالعه تشخیصی مقایسه‌ای بود که ۴۰ مدل LLM را ارزیابی کرد. هر مدل ۱۰ مطالعه RWE را با استفاده از یک فرآیند زنجیره‌ای تفکر ساختاریافته و چک‌لیست شرطی ارزیابی کرد. دو نویسنده انسانی مستقل با تخصص اپیدمیولوژی دارویی ارزیابی کردند و یک قاضی سوم اختلافات را حل کرد.

محدودیت‌ها

محدودیت‌ها در متن گزارش نشده‌اند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
مطالعات شواهد واقعی (RWE) منتشرشده در زمینه اپیدمیولوژی دارویی.
مداخله/مواجهه
ارزیابی سوگیری با استفاده از چارچوب ارزیابی پتانسیل سوگیری در مطالعات شواهد واقعی.
مقایسه
ارزیابی متخصصان انسانی (دو نویسنده مستقل و یک قاضی).
حجم نمونه
۴۰ مدل LLM، ۱۰ مطالعه RWE، ۲۸ مورد ارزیابی در هر مدل (۲۸۰ مورد در مجموع).

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and enviro…

PubMed2026

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated respons…

PubMed2026

Liability and Standard of Care in AI-Driven Psychiatric Practice: European Viewpoint.

AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management…

PubMed2026

Routine laboratory panels classify internal medicine ICD-10 code groups: comparison with frontier large language models and laboratory-only specialist assessment.

INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from …