PubMed چکیده/رکورد

Bilingual Performance of Large Language Models in Answering Consumer Health Questions in English and Chinese: Comparative Benchmark Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Large language models (LLMs) are increasingly used as health information intermediaries. Whether they provide comparable accuracy and communication quality across languages has direct implications for health information equity; however, systematic bilingual evaluations remain limited. OBJECTIVE: This study aimed to provide a preliminary bilingual benchmark evaluating whether 11 LLMs deliver comparable accuracy and communication quality when answering identical consumer health questions in English and Chinese. METHODS: We conducted a controlled evaluation of 11 LLMs (GPT-4.5, Claude Sonnet 4, Gemini 2.5 Flash, Grok 3, DeepSeek R1, Qwen 3, Doubao, Kimi k1.5, Hunyuan T1, ERNIE X1 Turbo, and ChatGLM 4) using 150 binary consumer health questions from the Text Retrieval Conference Health Misinformation Track (2019, 2021, and 2022). All models were accessed through official public-facing web interfaces during May 2025. Models were assessed under 2 full-benchmark prompting conditions (no-context and expert), evaluating accuracy, comprehensiveness, precision, and understandability. Four post hoc error-correction strategies (chain-of-thought [CoT], retrieval-augmented generation [RAG], CoT+RAG, and error attribution) were applied to baseline-incorrect responses. Composite ranking used the technique for order of preference by similarity to ideal solution (TOPSIS), with sensitivity analysis across 3 weighting schemes. Generalized estimating equations and linear mixed models with Benjamini-Hochberg false discovery rate (FDR) correction were applied using a full 3-way interaction specification (model×language×prompt). RESULTS: English and Chinese inputs showed comparable overall accuracy under no-context conditions (1572/1650, 95.27% vs 1548/1650, 93.82%), with no significant language main effect (β=0.00; P=.99). No language main effects for any individual model remained significant after FDR correction. TOPSIS analysis identified ChatGPT and Qwen as the most consistently top-ranked models (tier 1 in 12/12 condition×weight-scheme combinations). A model-specific language interaction emerged for communication quality: DeepSeek showed a significant English-language decrement in understandability (β=-0.73; FDR=-0.016), while its decrements in precision and comprehensiveness were not significant after correction. One 3-way interaction survived: Grok showed a disproportionate accuracy reduction when English input and expert prompting were combined (β=-1.88; FDR=-0.022). Among post hoc correction strategies, error attribution achieved the highest correction rate (Δ55.56%), although this condition provided models with privileged information. CONCLUSIONS: Contemporary LLMs achieved high binary accuracy on consumer health questions in both English and Chinese, with no significant aggregate language effect. The only robust model-specific language interaction was DeepSeek's English understandability decrement, independently confirmed by TOPSIS tier analysis. These findings suggested that cross-linguistic communication quality concerns were model-specific rather than universal and warrant targeted monitoring.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

AI in medicinebilingual benchmarkconsumer health questionslarge language modelsprompt engineering
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Community Stakeholder Perspectives on Children With Cerebral Palsy and Their Caregivers: A Qualitative Study From Rural Malawi.

BACKGROUND: Cerebral palsy (CP) disproportionately affects children in low- and middle-income countries, where stigma and discrimination lead to marginalization and reduced participation. To design culturally appropriate community-based health promotion strategies that can address these challenges, the role of community stakeholders is crucial, yet unstudied in rural sub-Saharan Africa. This study aimed to explore stakeholders' percept…

PubMed2026

Implementation of a Multidisciplinary Non-Pharmacological Program to Improve Urinary Incontinence in an Intermediate Care Hospital.

INTRODUCTION: Urinary incontinence (UI) is a prevalent geriatric syndrome that significantly affects the physical, psychological, and social well-being of older adults. Non-pharmacological interventions are recommended as the first-line approach, especially in frail older adults. DESIGN: This was a prospective pre-post observational study. METHODS: The study included sixty-one patients with rehabilitable UI who were admitted to an inte…

PubMed2026

Describing an integrated partnership approach to improving sexual health literacy and service navigation for international students in Sydney, Australia.

BACKGROUND: The population of international students (IS) in New South Wales (NSW) continues to grow post-COVID-19, with IS a key priority within NSW Health HIV and sexually transmissible infection (STI) strategies and efforts. Evidence shows gaps in sexual and reproductive health knowledge (SRH), and barriers to accessing services, including stigma, unfamiliarity with the Australian healthcare system and service cost concerns. Althoug…

PubMed2026

Moving beyond sexual and reproductive health knowledge: an evaluation of a co-designed digital engagement tool to promote sexual and reproductive health among international students in New South Wales, Australia.

BACKGROUND: This study evaluates the perceived impact and needed enhancement of a co-designed digital sexual and reproductive health (SRH) tool ('the Hub') developed to address SRH literacy gaps among international students (IS) in New South Wales, Australia. Despite Australia's large IS population and recognition of SRH as a fundamental human right, many report limited exposure to SRH information and face barriers to navigating health…