PubMed چکیده/رکورد

Mitigating medical bias in large language models by prompt engineering: an empirical study of effectiveness and trade-offs.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

Large language models (LLMs) demonstrate expert-level performance in various medical scenarios, yet their outputs can exhibit bias against groups or individuals with specific sensitive attributes, posing risks to patient safety and undermining trust in LLMs for healthcare. Recent research suggests that prompt engineering offers a convenient way to adjust model outputs, with the potential to mitigate such biases. However, there is a lack of empirical studies that systematically examine the effectiveness of prompt engineering and its trade-offs among fairness, accuracy and inference overhead. To fill this gap, we empirically evaluate five widely used prompting strategies across five influential LLMs in the latest medical bias benchmark. Results reveal substantial heterogeneity in both effectiveness and overhead across models, with no strategy proving universally effective and some even exacerbating bias. Chain-of-thought prompting yields the largest reduction, lowering the average gap across all scenarios by 2.4 percentage points, where the largest reduction is 6.2 percentage points, obtained in the DeepSeek-V3.1-sex case. Furthermore, the results of the McNemar test also show that it achieves the largest number of significant bias reduction cases (8/15), primarily by improving performance on unprivileged groups. These findings provide practical guidance for the fair deployment of LLMs in healthcare and highlight that mitigating medical bias remains a challenging problem requiring sustained efforts from both the artificial intelligence (AI) and medical communities. To support future research on fair AI in healthcare, we shall release all results and source code. This article is part of the theme issue 'Safe, secure and robust AI for safety-critical systems'.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

LLMsLLMs for healthbias mitigationfair healthcareprompt engineering
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Factors Influencing Medication Safety Competence of Nurses According to Clinical Career Stage: A Cross-Sectional Survey.

Medication errors pose a serious threat to patient safety, causing preventable harm and increasing healthcare costs. Nurses' competence in medication safety is critical and develops with clinical experience. This study aimed to identify factors influencing medication safety competence across different clinical career stages, focusing on nursing organizational culture, communication competence, and patient safety culture. A survey was c…

PubMed2026

Key Strategies for Improving Psychiatric Home-Based Care in Spain: A Modified Delphi Study.

AIMS: This study aimed to identify and prioritize key strategies for improving a psychiatric home-based care programme, the Crisis Resolution and Home Treatment (CRHT) intervention in Catalonia, Spain. The objective was to incorporate the perspectives of service users, family caregivers and healthcare professionals to guide quality improvement efforts. METHODS: A modified Delphi method was used to reach consensus among stakeholders pre…

PubMed2026

'HOPtimise' Logic Model: An Organisational Innovation to Improve Nursing Practices in Paediatric Oncology in Québec.

INTRODUCTION: HOPtimise is an organisational innovation in paediatric oncology in the province of Québec (Canada) destined to improve and reinforce best nursing practices through digital training using a serious game and create dashboards focused on indicators sensitive to the quality of care. This complex intervention is led by a quality improvement committee, a joint committee of user partners (former patients and family members) and…

PubMed2026

Cross-Cultural Adaptation and Psychometric Validation of the Hospital Survey on Patient Safety Culture Version 2 (HSOPSv2) in Spanish Hospitals.

INTRODUCTION: Patient safety culture is a key component of healthcare quality. The Hospital Survey on Patient Safety Culture (HSOPS), developed by AHRQ in 2004, has been widely used internationally to assess healthcare professionals' perceptions of safety. In 2019, version 2 (HSOPSv2) was released, with improvements in structure and item clarity. Although a Spanish version for North America exists, it has not yet been validated in the …