PubMed دسترسی آزاد

Quantifying the Impact of Anonymization-Induced Clinical Data Quality Loss: Methodological Quantitative Case Study Using Primary Diagnosis Codes and Hospital Length of Stay.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: The secondary use of electronic health record data requires robust privacy protection. k-Anonymity is widely used to enable data sharing by ensuring that each quasi-identifier combination occurs in at least k records; yet, its analytical impact on clinically meaningful structures remains insufficiently characterized, particularly for the combination of record suppression and microaggregation that arises when a numeric attribute lacks a natural generalization hierarchy. A further gap is that anonymization tools report internal information-loss values but do not signal the downstream distributional and inferential distortions these transformations introduce. OBJECTIVE: This study evaluated the analytical footprint of k-anonymity at k=5, 10, and 15 on 2 core data elements in retrospective hospital research: primary International Classification of Diseases, 10th Revision, German Modification (ICD-10-GM) diagnosis codes, and hospital length of stay (LOS). It aimed to determine and quantify whether anonymization introduces meaningful distortions not captured by the anonymization tool itself, and whether diagnosis-specific LOS patterns remain reproducible after anonymization. METHODS: We analyzed 719,387 inpatient encounters from University Hospital Mannheim from 2010 to 2024. Anonymization was performed with the ARX tool. It used record suppression and microaggregation. Distributional distortion was assessed with the Kolmogorov-Smirnov D statistic, quantile shifts, IQR changes, and tail changes. Categorical fidelity was assessed with the Jaccard coefficient and Cramer V. Inferential reproducibility was assessed with a 3-level linear mixed model. The model included random intercepts for the three-character diagnosis codes from the International Classification of Diseases (ICD-3) and patients. We compared the intraclass correlation coefficient and diagnosis-level effect concordance. Concordance was quantified using Spearman ρ and Lin concordance correlation coefficient, both with 95% CIs. A composite traffic-light verdict summarized the results. RESULTS: ARX masked quasi-identifier cells rather than deleting rows; the proportion of encounters with a masked cell rose from 0.77% (k=5) to 2.62% (k=15), distributed almost uniformly across admission years. Kolmogorov-Smirnov D was stable at 0.147. Median LOS shifted by 1 day, and the SD declined by about 6.5 days, while the diagnosis-level mean changed considerably (absolute mean shift of -1.81 days). Jaccard overlap fell from 0.624 to 0.421. The diagnosis intraclass correlation coefficient rose from 0.294 to 0.837, reflecting variance compression rather than improved signal. Best linear unbiased prediction rank concordance (Spearman ρ 0.964-0.970) and aggregate magnitude agreement (Lin concordance correlation coefficient 0.959-0.966) were high; yet, about 7% of low-signal diagnoses showed sign reversals. Most distributional change occurred at k=5. CONCLUSIONS: k-Anonymity preserved the ranking of diagnosis-specific LOS effects but altered distributional shape, individual effect magnitudes, and diagnostic vocabulary, none of which was flagged by the internal loss metric. Anonymized data of this type may support ordinal analyses but can mislead analyses requiring faithful variance structure, accurate absolute effects, or complete rare-diagnosis representation. A reporting checklist is provided to document these effects.

نتیجه فارسی

این مطالعه نشان داد که k-ناشناس‌سازی بر رتبه‌بندی اثرات مدت اقامت خاص به تشخیص تأثیر مثبت گذاشت اما شکل توزیع و واریانس را تغییر داد. ابزار ناشناس‌سازی این تغییرات را ثبت نکرد. داده‌های ناشناس‌سازی شده ممکن است برای تحلیل‌های ترتیبی مناسب باشند اما برای تحلیل‌های دقیق مطلق یا واریانس نادر ممکن است گمراه‌کننده باشند.

  • k-ناشناس‌سازی رتبه‌بندی اثرات LOS را حفظ کرد اما شکل توزیع و واریانس را تغییر داد.
  • ابزار ناشناس‌سازی (ARX) تغییرات را ثبت نکرد.
  • تغییر علامت در حدود 7% تشخیص‌های سیگنال ضعیف رخ داد.
  • این مطالعه بر داده‌های بستری از بیمارستان دانشگاه ماینز (2010-2024) انجام شد.
  • تغییرات توزیعی بیشتر در k=5 رخ داد.

ترجمه فارسی چکیده

پیش‌نیاز استفاده مجدد از داده‌های سوابق سلامت الکترونیکی، حفاظت قوی از حریم خصوصی است. k-ناشناس‌سازی برای اشتراک‌گذاری داده‌ها استفاده می‌شود، اما اثر تحلیلی آن بر ساختارهای بالینی مهم هنوز به‌طور کافی مشخص نشده است. این مطالعه اثر تحلیلی k-ناشناس‌سازی را در k=5، 10 و 15 بر روی دو عنصر داده اصلی در تحقیقات بیمارستانی بازگشتی ارزیابی کرد: کدهای تشخیص اولیه ICD-10-GM و مدت اقامت بیمارستانی (LOS). هدف بررسی و کمی‌سازی این بود که آیا ناشناس‌سازی تغییرات معناداری ایجاد می‌کند که ابزار ناشناس‌سازی خود ثبت نمی‌کند و الگوهای LOS خاص به تشخیص آیا پس از ناشناس‌سازی قابل تکرار باقی می‌مانند. ما 719،387 رویداد بستری را از بیمارستان دانشگاه ماینز از 2010 تا 2024 تحلیل کردیم. ناشناس‌سازی با ابزار ARX با حذف رکورد و میکروآگgregatio انجام شد. تغییرات توزیعی با آماره D کولموگوروف-اسمیرنوف، تغییرات کوانتیل، تغییرات IQR و دم ارزیابی شد. وفاداری دسته‌بندی با ضریب Jaccard و Cramer V ارزیابی شد. تکرار استنباطی با مدل خطی آمیخته سه سطحی ارزیابی شد. ما ضریب همبستگی درون‌گروهی و همسویی اثرات سطح تشخیص را مقایسه کردیم. همسویی با ضریب اسپیرمن ρ و ضریب همبستگی همسویی لین، هر دو با 95% CI، کمی‌سازی شد. یک خلاصه تصمیم ترافیک‌نمایی نتایج را خلاصه کرد. نتایج نشان داد که ARX سلول‌های شناسایی نیمه‌شخصی را ماسک کرد تا از حذف ردیف جلوگیری شود. ضریب Jaccard از 0.624 به 0.421 کاهش یافت. ضریب همبستگی درون‌گروهی تشخیصی از 0.294 به 0.837 افزایش یافت که نشان‌دهنده فشردگی واریانس است. همسویی رتبه‌بندی پیش‌بینی خطی بهینه (Spearman ρ 0.964-0.970) و توافق اندازه کلان (Lin concordance correlation coefficient 0.959-0.966) بالا بود، اما حدود 7% تشخیص‌های سیگنال ضعیف تغییر علامت نشان دادند. k-ناشناس‌سازی رتبه‌بندی اثرات LOS خاص به تشخیص را حفظ کرد اما شکل توزیع، اندازه اثرات فردی و واژگان تشخیصی را تغییر داد که هیچ‌کدام توسط شاخص داخلی ضرر ثبت نشدند. داده‌های ناشناس‌سازی شده از این نوع ممکن است تحلیل‌های ترتیبی را پشتیبانی کند اما می‌تواند تحلیل‌هایی که نیاز به ساختار واریانس وفادار، اثرات مطلق دقیق یا نمای کامل تشخیص‌های نادر دارند، گمراه کند.

روش پژوهش

تحلیل 719،387 رویداد بستری با استفاده از ابزار ARX برای ناشناس‌سازی. ارزیابی تغییرات توزیعی با D کولموگوروف-اسمیرنوف و تغییرات کوانتیل. ارزیابی وفاداری دسته‌بندی با ضریب Jaccard و Cramer V. ارزیابی تکرار استنباطی با مدل خطی آمیخته سه سطحی.

محدودیت‌ها

محدودیت‌های گزارش نشده در متن موجود.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
بیماران بستری در بیمارستان دانشگاه ماینز (2010 تا 2024).
مداخله/مواجهه
k-ناشناس‌سازی با k=5، 10 و 15 با استفاده از ابزار ARX (شامل حذف رکورد و میکروآگgregatio).
مقایسه
داده‌های اصلی قبل از ناشناس‌سازی (برای مقایسه تغییرات).
حجم نمونه
719،387 رویداد بستری.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

ADAQI reporting checklistARX Data Anonymization ToolAnonymization and Data Analysis Quality and Impact ReportingICD codesInternational Classification of Diseasesclinical data qualitydata anonymizationdistributional fidelityinferential reproducibilityk-anonymitylength of staylinear mixed modelmedical informatics
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …