PubMed چکیده/رکورد

Embedding-Based Evaluation and Benchmarking Framework for Optimizing LLM Post-Processing of Medical Transcriptions from a Local Whisper System.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

INTRODUCTION: Locally deployed speech-to-text systems such as Whisper enable privacy-preserving transcription of medical encounters. However, the resulting transcripts are often lengthy, noisy, and insufficiently structured for direct integration into clinical documentation workflows. Large language models (LLMs) can transform such transcripts into concise clinical summaries, yet evaluating their quality and reliability remains challenging, particularly in the absence of gold-standard reference summaries. MATERIALS AND METHODS: We propose a reference-free evaluation framework for benchmarking LLM-based post-processing of clinical transcripts generated by a locally deployed Whisper system. The framework combines embedding-based transcript-summary semantic similarity, stability analysis across repeated generations, and structured human evaluation. It was applied to 26 German physician-patient conversations. Four LLMs generated four summaries per transcript using standardized prompts and decoding parameters, resulting in 416 summaries. A subset of summaries was additionally assessed by three human raters using six predefined quality criteria. RESULTS: GPT-OSS-120B and MedGemma-27B achieved the highest transcript-summary similarity scores across most embedding models. Although absolute similarity values varied across embedding models, relative model rankings remained largely consistent, indicating robustness of the evaluation framework. Stability analysis showed high consistency across repeated runs, with cosine similarities typically exceeding 0.90, while higher sampling temperatures reduced semantic similarity. Human evaluation showed partial agreement between embedding-based similarity signals and human judgments of summary quality. CONCLUSION: The proposed framework enables scalable, reference-free evaluation of LLM-based clinical summarization in privacy-sensitive settings. By combining semantic similarity, stability analysis, and human evaluation, it supports systematic model benchmarking, relative model comparison, and optimization without requiring reference summaries. These findings suggest that embedding-based metrics can provide useful signals for selecting and optimizing LLM-based post-processing models in local speech-to-text pipelines.

نتیجه فارسی

این مطالعه چارچوبی بدون مرجع برای ارزیابی مدل‌های زبانی بزرگ (LLM) در تولید خلاصه‌های بالینی از ترانسکریپت‌های Whisper محلی ارائه می‌دهد. این چارچوب از شباهت معنایی مبتنی بر تلفیق، تحلیل پایداری و ارزیابی انسانی استفاده می‌کند. نتایج نشان می‌دهد که GPT-OSS-120B و MedGemma-27B عملکرد بهتری دارند و این چارچوب قابلیت اعتماد بالایی برای مقایسه نسبی مدل‌ها در محیط‌های حساس به حریم خصوصی دارد.

  • ارائه چارچوب ارزیابی بدون مرجع برای خلاصه‌سازی بالینی LLM
  • استفاده از شباهت معنایی مبتنی بر تلفیق، پایداری و ارزیابی انسانی
  • ارزیابی بر روی ۲۶ مکالمه پزشک-بیمار آلمانی
  • GPT-OSS-120B و MedGemma-27B بالاترین امتیازات را کسب کردند
  • ارزیابی انسانی همسانی جزئی با معیارهای مبتنی بر تلفیق نشان داد

ترجمه فارسی چکیده

سیستم‌های تبدیل گفتار به متن محلی مانند Whisper امکان ترانسکریپت‌سازی ملاقات‌های پزشکی با حفظ حریم خصوصی را فراهم می‌کنند. با این حال، ترانسکریپت‌های حاصل اغلب طولانی، پر از نویز و ساختارمند نیستند تا بتوانند مستقیماً در جریان‌های کاری مستندات بالینی ادغام شوند. مدل‌های زبانی بزرگ (LLM) می‌توانند چنین ترانسکریپت‌هایی را به خلاصه‌های بالینی مختصر تبدیل کنند، اما ارزیابی کیفیت و قابلیت اعتماد آن‌ها همچنان چالش‌برانگیز است، به‌ویژه در نبود خلاصه‌های مرجع استاندارد. ما چارچوب ارزیابی بدون مرجع را برای بنچمارک پس‌پردازش LLM بر اساس ترانسکریپت‌های بالینی تولید شده توسط یک سیستم Whisper محلی پیشنهاد می‌کنیم. این چارچوب ترکیبی از شباهت معنایی ترانسکریپت-خلاصه مبتنی بر تلفیق، تحلیل پایداری در تولیدهای تکراری و ارزیابی ساختاریافته انسانی است. این رویکرد بر ۲۶ مکالمه پزشک-بیمار آلمانی اعمال شد. چهار مدل LLM چهار خلاصه برای هر ترانسکریپت با استفاده از پرامپت‌های استاندارد و پارامترهای کدگذاری تولید کردند که منجر به ۴۱۶ خلاصه شد. یک زیرمجموعه از خلاصه‌ها همچنین توسط سه ارزیاب انسانی با شش معیار کیفیت پیش‌فرض ارزیابی شد. نتایج نشان داد که GPT-OSS-120B و MedGemma-27B بالاترین امتیازات شباهت ترانسکریپت-خلاصه را در اکثر مدل‌های تلفیقی کسب کردند. تحلیل پایداری نشان‌دهنده همسانی بالا در اجراهای تکراری بود، در حالی که دما (temperature) بالاتر نمونه‌برداری کاهش معناداری در شباهت معنایی ایجاد کرد. ارزیابی انسانی نشان داد همسانی جزئی بین سیگنال‌های شباهت مبتنی بر تلفیق و قضاوت‌های انسانی در کیفیت خلاصه وجود دارد.

روش پژوهش

چارچوب پیشنهادی شامل ترکیب شباهت معنایی مبتنی بر تلفیق، تحلیل پایداری و ارزیابی انسانی است. این رویکرد بر ۲۶ مکالمه پزشک-بیمار آلمانی اعمال شد و چهار مدل LLM چهار خلاصه برای هر ترانسکریپت تولید کردند.

محدودیت‌ها

محدودیت‌های گزارش شده شامل عدم وجود خلاصه‌های مرجع (gold-standard) و همسانی جزئی بین ارزیابی‌های مبتنی بر تلفیق و ارزیابی انسانی است.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
مکالمات پزشک-بیمار آلمانی (۲۶ مورد)
مداخله/مواجهه
پس‌پردازش LLM برای تولید خلاصه‌های بالینی
مقایسه
هیچ مقایسه‌کننده‌ای به طور خاص ذکر نشده، اما مدل‌های مختلف LLM مقایسه شدند.
حجم نمونه
۲۶ مکالمه پزشک-بیمار (منجر به ۴۱۶ خلاصه)

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

Artificial IntelligenceBenchmarkingLarge Language ModelsMedical RecordsNatural Language ProcessingSpeech Recognition Software
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …