Large Language Models for Distress Rating in Korean Psycho-Oncology Interviews: Exploratory Clinician-Benchmarked Evaluation Study.
پخش حرفهای فارسی و انگلیسی
در حال بررسی نسخههای صوتی ذخیرهشده…
تنظیم صدای طبیعی و سرعت
صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده میشود معمولاً طبیعیترند. انتخاب صدا به صداهای نصبشده در ویندوز و مرورگر شما بستگی دارد.
چکیده اصلی
BACKGROUND: Psychiatric distress is common among patients with cancer; yet, systematic interview-based screening remains difficult to scale in routine clinical care. Large language models (LLMs) have shown promise as scalable tools for mental health assessment, but most existing evidence is derived from clinician-authored records, translated text, or proxy data. The performance characteristics, error patterns, and explanatory behaviors of contemporary LLMs when applied to authentic, non-English psychiatric interviews remain insufficiently characterized. OBJECTIVE: This exploratory study evaluated how contemporary LLMs reproduced psycho-oncologists' item-level symptom ratings from real-world Korean psycho-oncology interviews, focusing on concordance, directional bias, and clinician-adjudicated error characteristics. METHODS: Between April 2024 and May 2025, 101 adults receiving oncologic care in South Korea underwent semistructured interviews. Board-certified psycho-oncologists provided real-time, time-stamped ratings of 39 items. Generative pretrained transformer 4o (GPT-4o), Claude 3.5 Sonnet, and Gemini 2.5 Flash generated item scores and brief rationales using identical Korean zero-shot rubrics. Concordance between clinician ratings was evaluated using ordinal and binary screening metrics. A paired Wilcoxon test compared patient-level total symptom burden. Binary mismatches were clinically adjudicated as ambiguous or definite overestimation or underestimation with an 8-etiology taxonomy. Model-generated rationales were further meta-evaluated using GPT-5.4, Claude Sonnet 4.6, and Gemini Pro 3.1 across 4 dimensions: citation (use of quoted supporting statements), structure (logical organization of the rationale), mapping (consistency between the rationale and the assigned item rating), and expansion (degree of interpretive elaboration beyond the explicit transcript content). The associations between meta-evaluation results and absolute error were examined using cross-classified mixed-effects models. RESULTS: Out of 101 participants, 88 (87.1%) were predominantly female and had breast cancer as the primary cancer type (n=70, 69.3%). Across 3931 item-level ratings, all models showed good agreement with clinicians (intraclass correlation coefficient: 0.816-0.872), with GPT-4o showing the highest agreement. Claude 3.5 and Gemini 2.5 yielded significantly higher patient-level symptom burden (adjusted P<.001 and adjusted P=.002, respectively), whereas GPT-4o did not (adjusted P=.15). Clinician adjudication attributed 97 of 248 (39.1%) GPT-4o mismatches, 88 of 301 (29.2%) Claude 3.5 mismatches, and 118 of 355 (33.2%) Gemini 2.5 mismatches to intrinsic ambiguity in patient speech. Among definite errors, misapplication of severity thresholds was the predominant mechanism across models. Exploratory mixed-effects models showed that higher expansion and lower mapping scores were associated with larger absolute errors. However, the association with expansion may reflect case difficulty or ambiguity rather than causation. CONCLUSIONS: In this exploratory clinician-benchmarked evaluation of authentic interviews from a Korean psycho-oncology sample comprising predominantly women and patients with breast cancer, LLMs showed high aggregate concordance with psycho-oncologists' item-level ratings while differing in their error profiles. These findings support further evaluation of LLM-based item-level symptom-rating approaches in psycho-oncology. Validation in larger, more diverse, and independent cohorts is needed.
نتیجه فارسی
این مطالعه ارزیابی کرد که مدلهای زبانی بزرگ (LLM) چگونه رتبهبندیهای علائم را در مصاحبههای واقعی رواندرمانی-سرطانشناسی کرهای بازتولید میکنند. نتایج نشان داد که LLMها همخوانی خوبی با پزشکان دارند، اما الگوهای خطا متفاوتی دارند. GPT-4o بالاترین همخوانی را نشان داد، در حالی که Claude 3.5 و Gemini 2.5 بار کل علائم بالاتری را گزارش کردند. اکثر خطاها به ابهام در گفتار بیمار نسبت داده شد.
- LLMها همخوانی خوبی (ICC: 0.816-0.872) با رتبهبندیهای پزشکان در ۳۹۳۱ رتبهبندی آیتم نشان دادند.
- GPT-4o بالاترین همخوانی را داشت، در حالی که Claude 3.5 و Gemini 2.5 بار کل علائم بالاتری را گزارش کردند.
- بیشتر خطاها به ابهام در گفتار بیمار نسبت داده شد.
- خطاهای قطعی عمدتاً ناشی از کاربرد نادرست آستانههای شدت بودند.
- ارزیابیهای متا نشان دادند که توضیحات گستردهتر با خطای مطلق بزرگتری مرتبط است.
ترجمه فارسی چکیده
پایگاه و هدف: درد روانی در بیماران سرطان شایع است، اما غربالگری مبتنی بر مصاحبه سیستماتیک هنوز در مراقبتهای بالینی روتین دشوار است. مدلهای زبانی بزرگ (LLM) پتانسیل ابزارهای مقیاسپذیر برای ارزیابی سلامت روان را نشان دادهاند، اما اکثر شواهد موجود از سوابق نوشته شده توسط پزشکان، متون ترجمه شده یا دادههای واسط است. ویژگیهای عملکردی، الگوهای خطا و رفتارهای توضیحی LLMهای امروزی در مصاحبههای روانپزشکی غیرانگلیسی باقی ماندهاند. این مطالعه آزمایشی ارزیابی کرد که LLMهای امروازی چگونه رتبهبندیهای علائم سطح آیتم از مصاحبههای واقعی رواندرمانی-سرطانشناسی کرهای را بازتولید میکنند. روشها: بین آوریل ۲۰۲۴ و مه ۲۰۲۵، ۱۰۱ بالغ در کره جنوبی مصاحبههای نیمهساختیافته دریافت کردند. رتبهبندیهای ۳۹ آیتم توسط پزشکان متخصص رواندرمانی-سرطانشناسی در زمان واقعی ارائه شد. GPT-4o، Claude 3.5 Sonnet و Gemini 2.5 Flash امتیازات و دلایل کوتاه را با استفاده از همان معیارهای صفر-نمونه کرهای تولید کردند. همخوانی بین رتبهبندیهای پزشکان با استفاده از معیارهای غربالگری ترتیبی و باینری ارزیابی شد. آزمون جفتشده ویلکاکسون برای بار کل علائم بیمار استفاده شد. خطاهای باینری توسط پزشکان طبقهبندی شدند. دلایل مدلها با استفاده از Claude Sonnet 4.6 و Gemini Pro 3.1 در ۴ بعد ارزیابی شدند. نتایج: از ۱۰۱ شرکتکننده، ۸۸ نفر (۸۷.۱٪) عمدتاً زن و سرطان پستان به عنوان نوع اصلی سرطان داشتند. در میان ۳۹۳۱ رتبهبندی سطح آیتم، تمام مدلها همخوانی خوبی با پزشکان نشان دادند (ضریب همبستگی درونکلاسی: ۰.۸۱۶-۰.۸۷۲). Claude 3.5 و Gemini ۲.۵ بار کل علائم بالاتری داشتند، در حالی که GPT-4o این را نشان نداد. ۹۷ مورد از خطاهای GPT-4o به ابهام در گفتار بیمار نسبت داده شد. نتیجهگیری: در این ارزیابی آزمایشی، LLMها همخوانی کلیدی با رتبهبندیهای پزشکان را نشان دادند.
روش پژوهش
این مطالعه آزمایشی شامل ۱۰۱ شرکتکننده در کره جنوبی بود که مصاحبههای نیمهساختیافته دریافت کردند. پزشکان رتبهبندیهای ۳۹ آیتم را در زمان واقعی ارائه دادند. GPT-4o، Claude 3.5 Sonnet و Gemini 2.5 Flash امتیازات را با استفاده از معیارهای صفر-نمونه کرهای تولید کردند. همخوانی با استفاده از معیارهای ترتیبی و باینری ارزیابی شد.
محدودیتها
مطالعه شامل اکثریت زنان و بیماران با سرطان پستان بود. نتایج به طور آزمایشی است و نیاز به تأیید در گروههای بزرگتر و متنوعتر دارد. همبستگی با گسترش توضیحات ممکن است ناشی از دشواری مورد باشد.
نمای PICO و پیامدها
- جمعیت
- ۱۰۱ بالغ در کره جنوبی که مراقبت سرطانی دریافت میکردند. ۸۸ نفر (۸۷.۱٪) عمدتاً زن و ۷۰ نفر (۶۹.۳٪) سرطان پستان داشتند.
- مداخله/مواجهه
- GPT-4o، Claude 3.5 Sonnet و Gemini 2.5 Flash برای تولید امتیازات و دلایل کوتاه با استفاده از معیارهای صفر-نمونه کرهای استفاده شدند.
- مقایسه
- رتبهبندیهای ۳۹ آیتم توسط پزشکان متخصص رواندرمانی-سرطانشناسی در زمان واقعی.
- حجم نمونه
- ۱۰۱ شرکتکننده. ۳۹۳۱ رتبهبندی آیتم.
متن کامل اصلی
لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز میشود.
باز کردن متن کامل