PubMed دسترسی آزاد

Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. OBJECTIVE: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists' item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. METHODS: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. RESULTS: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816-0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted P<.001 and adjusted P=.002, respectively), whereas GPT-4o did not (adjusted P=.15). Clinician adjudication attributed 97 of 248 (39.1%) GPT-4o mismatches, 88 of 301 (29.2%) Claude 3.5 mismatches, and 118 of 355 (33.2%) Gemini 2.5 mismatches to intrinsic ambiguity in patient speech. Among definite errors, misapplication of severity thresholds was the predominant mechanism across models. Exploratory mixed-effects models showed that higher expansion and lower mapping scores were associated with larger absolute errors. However, the association with expansion may reflect case difficulty or ambiguity rather than causation. CONCLUSIONS: In this exploratory clinician-benchmarked evaluation of authentic interviews from a Korean psycho-oncology sample comprising predominantly women and patients with breast cancer, LLMs showed high aggregate concordance with psycho-oncologists' item-level ratings while differing in their error profiles. These findings support further evaluation of LLM-based item-level symptom-rating approaches in psycho-oncology. Validation in larger, more diverse, and independent cohorts is needed.

نتیجه فارسی

این مطالعه ارزیابی کرد که مدل‌های زبانی بزرگ (LLM) چگونه رتبه‌بندی‌های علائم را در مصاحبه‌های واقعی روان‌درمانی-سرطان‌شناسی کره‌ای بازتولید می‌کنند. نتایج نشان داد که LLMها همخوانی خوبی با پزشکان دارند، اما الگوهای خطا متفاوتی دارند. GPT-4o بالاترین همخوانی را نشان داد، در حالی که Claude 3.5 و Gemini 2.5 بار کل علائم بالاتری را گزارش کردند. اکثر خطاها به ابهام در گفتار بیمار نسبت داده شد.

  • LLMها همخوانی خوبی (ICC: 0.816-0.872) با رتبه‌بندی‌های پزشکان در ۳۹۳۱ رتبه‌بندی آیتم نشان دادند.
  • GPT-4o بالاترین همخوانی را داشت، در حالی که Claude 3.5 و Gemini 2.5 بار کل علائم بالاتری را گزارش کردند.
  • بیشتر خطاها به ابهام در گفتار بیمار نسبت داده شد.
  • خطاهای قطعی عمدتاً ناشی از کاربرد نادرست آستانه‌های شدت بودند.
  • ارزیابی‌های متا نشان دادند که توضیحات گسترده‌تر با خطای مطلق بزرگتری مرتبط است.

ترجمه فارسی چکیده

پایگاه و هدف: درد روانی در بیماران سرطان شایع است، اما غربالگری مبتنی بر مصاحبه سیستماتیک هنوز در مراقبت‌های بالینی روتین دشوار است. مدل‌های زبانی بزرگ (LLM) پتانسیل ابزارهای مقیاس‌پذیر برای ارزیابی سلامت روان را نشان داده‌اند، اما اکثر شواهد موجود از سوابق نوشته شده توسط پزشکان، متون ترجمه شده یا داده‌های واسط است. ویژگی‌های عملکردی، الگوهای خطا و رفتارهای توضیحی LLMهای امروزی در مصاحبه‌های روان‌پزشکی غیرانگلیسی باقی مانده‌اند. این مطالعه آزمایشی ارزیابی کرد که LLMهای امروازی چگونه رتبه‌بندی‌های علائم سطح آیتم از مصاحبه‌های واقعی روان‌درمانی-سرطان‌شناسی کره‌ای را بازتولید می‌کنند. روش‌ها: بین آوریل ۲۰۲۴ و مه ۲۰۲۵، ۱۰۱ بالغ در کره جنوبی مصاحبه‌های نیمه‌ساخت‌یافته دریافت کردند. رتبه‌بندی‌های ۳۹ آیتم توسط پزشکان متخصص روان‌درمانی-سرطان‌شناسی در زمان واقعی ارائه شد. GPT-4o، Claude 3.5 Sonnet و Gemini 2.5 Flash امتیازات و دلایل کوتاه را با استفاده از همان معیارهای صفر-نمونه کره‌ای تولید کردند. همخوانی بین رتبه‌بندی‌های پزشکان با استفاده از معیارهای غربالگری ترتیبی و باینری ارزیابی شد. آزمون جفت‌شده ویلکاکسون برای بار کل علائم بیمار استفاده شد. خطاهای باینری توسط پزشکان طبقه‌بندی شدند. دلایل مدل‌ها با استفاده از Claude Sonnet 4.6 و Gemini Pro 3.1 در ۴ بعد ارزیابی شدند. نتایج: از ۱۰۱ شرکت‌کننده، ۸۸ نفر (۸۷.۱٪) عمدتاً زن و سرطان پستان به عنوان نوع اصلی سرطان داشتند. در میان ۳۹۳۱ رتبه‌بندی سطح آیتم، تمام مدل‌ها همخوانی خوبی با پزشکان نشان دادند (ضریب همبستگی درون‌کلاسی: ۰.۸۱۶-۰.۸۷۲). Claude 3.5 و Gemini ۲.۵ بار کل علائم بالاتری داشتند، در حالی که GPT-4o این را نشان نداد. ۹۷ مورد از خطاهای GPT-4o به ابهام در گفتار بیمار نسبت داده شد. نتیجه‌گیری: در این ارزیابی آزمایشی، LLMها همخوانی کلیدی با رتبه‌بندی‌های پزشکان را نشان دادند.

روش پژوهش

این مطالعه آزمایشی شامل ۱۰۱ شرکت‌کننده در کره جنوبی بود که مصاحبه‌های نیمه‌ساخت‌یافته دریافت کردند. پزشکان رتبه‌بندی‌های ۳۹ آیتم را در زمان واقعی ارائه دادند. GPT-4o، Claude 3.5 Sonnet و Gemini 2.5 Flash امتیازات را با استفاده از معیارهای صفر-نمونه کره‌ای تولید کردند. همخوانی با استفاده از معیارهای ترتیبی و باینری ارزیابی شد.

محدودیت‌ها

مطالعه شامل اکثریت زنان و بیماران با سرطان پستان بود. نتایج به طور آزمایشی است و نیاز به تأیید در گروه‌های بزرگتر و متنوع‌تر دارد. همبستگی با گسترش توضیحات ممکن است ناشی از دشواری مورد باشد.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
۱۰۱ بالغ در کره جنوبی که مراقبت سرطانی دریافت می‌کردند. ۸۸ نفر (۸۷.۱٪) عمدتاً زن و ۷۰ نفر (۶۹.۳٪) سرطان پستان داشتند.
مداخله/مواجهه
GPT-4o، Claude 3.5 Sonnet و Gemini 2.5 Flash برای تولید امتیازات و دلایل کوتاه با استفاده از معیارهای صفر-نمونه کره‌ای استفاده شدند.
مقایسه
رتبه‌بندی‌های ۳۹ آیتم توسط پزشکان متخصص روان‌درمانی-سرطان‌شناسی در زمان واقعی.
حجم نمونه
۱۰۱ شرکت‌کننده. ۳۹۳۱ رتبه‌بندی آیتم.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

distresslarge language modelmental healthpatients with cancerpsycho-oncology
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Global Genomic Surveillance.

Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and enviro…

PubMed2026

Development and Nationwide Multicentre Evaluation of Guideline-Grounded Large Language Model Chatbots to Support Patient Self-Management and Education in Rheumatology.

Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated respons…

PubMed2026

Liability and Standard of Care in AI-Driven Psychiatric Practice: European Viewpoint.

AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management…

PubMed2026

Routine laboratory panels classify internal medicine ICD-10 code groups: comparison with frontier large language models and laboratory-only specialist assessment.

INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from …