PubMed دسترسی آزاد

Anonymization of Portuguese Clinical Notes Using Large Language Models and Quantum-Enhanced Hybrid Architectures: Comparative Evaluation Study.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: The widespread adoption of electronic health records (EHRs) has generated large-scale repositories of highly sensitive clinical information, emphasizing the need for robust anonymization strategies to enable secondary use for research while safeguarding patient privacy. Conventional rule-based and machine learning approaches for deidentifying medical text face limitations with the linguistic complexity, variability, and context dependence inherent to clinical documentation. Recent advances in large language models (LLMs), combined with emerging quantum computing paradigms, present novel opportunities to enhance the accuracy, scalability, and resilience of health care data anonymization. OBJECTIVE: This study aims to evaluate the efficacy of LLM-based and quantum-enhanced hybrid architectures for medical text anonymization, assessing the effectiveness and computational efficiency across multiple entity types in Portuguese clinical notes. METHODS: We constructed a gold-standard corpus of 1000 Portuguese outpatient clinical notes, manually annotated by 5 trained researchers for 5 protected-entity categories: patient names, dates, identifiers, organizations, and geographic locations. Four anonymization strategies were evaluated: 2 stand-alone LLMs (Llama-3.1-8B-instruct and Llama-3.3-70B-instruct) and 2 quantum-enhanced hybrid models (Dynex-QML with 8B and 70B base models) incorporating quantum optimization via Quadratic Unconstrained Binary Optimization (QUBO) formulations. The quantum-enhanced approach transforms the final attention layer of the LLM into a global constraint satisfaction problem solved via neuromorphic quantum annealing. Model performance was measured on a held-out test set of 500 notes using precision, recall, and F1-score metrics. Computational efficiency was quantified through end-to-end processing time. RESULTS: The quantum-enhanced Dynex-QML-70B model achieved the highest overall performance with a macro-F1-score of 0.855 (95% CI 0.823-0.880), outperforming the stand-alone Llama-3.3-70B (0.726, 95% CI 0.704-0.747), Dynex-QML-8B (0.733, 95% CI 0.709-0.756), and Llama-3.1-8B (0.602, 95% CI 0.588-0.615). Compared with Llama 3.3 70B, Dynex-QML (Llama 70B) improved macro-F1-score by 0.128 (95% CI 0.091-0.163; empirical 2-sided bootstrap P<.001). Most entity-level within-size comparisons favored the Dynex-QML models and were statistically significant, although the ORGANIZATION comparison between Dynex-QML (Llama 8B) and Llama 3.1 8B was not significant. For total elapsed time, Dynex-QML (Llama 70B) was faster than stand-alone Llama 3.3 70B (7.97, 95% CI 7.72-8.21 seconds per note vs 8.52, 95% CI 8.23-8.81 seconds per note). In a paired note-level bootstrap comparison, this corresponded to a mean reduction of 0.55 (SD 2.36; 95% CI -0.75 to -0.34 seconds per note; P<.001). CONCLUSIONS: Quantum-enhanced hybrid architectures provide substantial improvements in medical text anonymization accuracy compared to stand-alone LLMs, particularly reductions in false-positive rates while preserving high sensitivity. The Dynex-QML-70B model achieved the best balance between performance and efficiency, suggesting that quantum-enhanced optimization offers a strategy for high-fidelity, scalable, and regulation-compliant deidentification of clinical text. These findings highlight the potential of emerging quantum-AI paradigms to advance the secondary secure use of health care data.

نتیجه فارسی

این مطالعه به مقایسه چهار استراتژی استانجام‌نامه‌سازی در یادداشت‌های بالینی پرتغالی پرداخته است. دو مدل زبانی بزرگ مستقل (Llama) و دو مدل هیبریدی تقویت‌شده با کوانتوم (Dynex-QML) ارزیابی شدند. مدل Dynex-QML-70B بالاترین عملکرد کلی را با نمره F1-مجموعی 0.855 نشان داد و نسبت به مدل‌های Llama-3.3-70B و Llama-3.1-8B عملکرد بهتری داشت. همچنین، مدل Dynex-QML-70B زمان پردازش را نسبت به مدل Llama-3.3-70B کاهش داد.

  • ارزیابی چهار مدل استانجام‌نامه‌سازی در یادداشت‌های بالینی پرتغالی.
  • مدل Dynex-QML-70B بالاترین نمره F1-مجموعی (0.855) را کسب کرد.
  • مدل Dynex-QML-70B نسبت به مدل Llama-3.3-70B زمان پردازش را کاهش داد.
  • معماری‌های تقویت‌شده با کوانتوم بهبود دقت در استانجام‌نامه‌سازی متن پزشکی را نشان دادند.

ترجمه فارسی چکیده

پذیرش گسترده سوابق الکترونیک سلامت (EHR) مخازن بزرگی از اطلاعات بالینی حساس ایجاد کرده است که نیاز به استراتژی‌های قوی برای استانجام‌نامه‌سازی را برای استفاده ثانویه تحقیقاتی، در حالی که حفظ حریم خصوصی بیمار، برجسته می‌کند. رویکردهای سنتی مبتنی بر قوانین و یادگیری ماشین برای حذف هویت متن پزشکی با پیچیدگی‌های زبانی، متغیر بودن و وابستگی به زمینه در مستندات بالینی محدودیت‌هایی را تجربه می‌کنند. پیشرفت‌های اخیر در مدل‌های زبانی بزرگ (LLM)، همراه با پارادایم‌های در حال ظهور محاسبات کوانتومی، فرصت‌های جدیدی برای ارتقای دقت، مقیاس‌پذیری و تاب‌آوری استانجام‌نامه‌سازی داده‌های سلامت ارائه می‌دهند. این مطالعه به ارزیابی اثربخشی معماری‌های مبتنی بر LLM و تقویت‌شده با کوانتوم برای استانجام‌نامه‌سازی متن پزشکی، در دسته‌های مختلف موجودیت‌ها در یادداشت‌های بالینی پرتغالی می‌پردازد.

روش پژوهش

مطالعه شامل ۱۰۰۰ یادداشت بالینی پرتغالی (۵۰۰ یادداشت تست) بود. مدل‌ها با استفاده از معیارهای دقت،recall و نمره F1 ارزیابی شدند. زمان پردازش به‌صورت end-to-end اندازه‌گیری شد.

محدودیت‌ها

محدودیت‌های متن در این متن گزارش نشده است.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
یادداشت‌های بالینی پرتغالی (۱۰۰۰ یادداشت آموزشی و ۵۰۰ یادداشت تست)
مداخله/مواجهه
استراتژی‌های استانجام‌نامه‌سازی شامل دو مدل Llama-3.1-8B-instruct، Llama-3.3-70B-instruct و دو مدل Dynex-QML با پایه 8B و 70B
مقایسه
مدل‌های Llama-3.1-8B-instruct و Llama-3.3-70B-instruct
حجم نمونه
۱۰۰۰ یادداشت آموزشی و ۵۰۰ یادداشت تست

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

GDPR complianceHIPAA complianceLGPD complianceLLMsNLPQMLdata privacyelectronic health recordlarge language modelsmedical text anonymizationnatural language processingquantum machine learning
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …