PubMed دسترسی آزاد

Bridging the data latency gap: automated extraction of genomic biomarkers from unstructured clinical documents to support real-world oncology data.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

PURPOSE: Real-world oncology data are essential for clinical research and precision cancer care. However, genomic biomarkers are often embedded in scanned, unstructured clinical documents requiring manual abstraction before becoming available in cancer registries, delaying real-world evidence generation. This study evaluated and compared three open-source optical character recognition (OCR) approaches, Tesseract, EasyOCR, and a hybrid implementation, to determine which best enables automated extraction of Oncotype DX recurrence scores and improves the timeliness and quality of real-world oncology data. METHODS: We evaluated the feasibility of automated genomic data extraction using 675 Oncotype DX reports from a Midwestern U.S. health system. EasyOCR, Tesseract, and a hybrid OCR approach were used to extract recurrence scores from scanned reports. OCR-derived values were compared with manually abstracted scores and local cancer registry data. Performance was assessed using agreement, precision, recall, F1 score, and processing time. Multivariable logistic regression was performed to identify factors associated with discordance between registry-reported and manually abstracted scores. RESULTS: The hybrid OCR approach demonstrated the highest performance, achieving 97% agreement with manual abstraction, precision of 0.997, recall of 0.972, and an F1 score of 0.984. Registry abstraction demonstrated comparable performance but required greater manual effort. Automated extraction substantially reduced processing time while maintaining high accuracy. Logistic regression showed registry discordance was largely independent of patient and tumor characteristics, with unknown progesterone receptor (PR) status as the only significant predictor. CONCLUSION: Automated extraction of genomic biomarkers represents a scalable approach to reducing delays in cancer data availability. Earlier capture of genomic information may support cancer registry modernization and improve real-world evidence generation in precision oncology.

نتیجه فارسی

این مطالعه رویکردی هیبریدی OCR را برای استخراج خودکار امتیاز تکرار Oncotype DX از اسناد اسکن‌شده ارزیابی کرد. این روش عملکردی بسیار بالا (۹۷٪ توافق) با زمان پردازش کمتر نسبت به انتزاع دستی نشان داد. ناهماهنگی در داده‌های ثبت سرطان عمدتاً مستقل از ویژگی‌های بیمار بود.

  • ارزیابی سه روش OCR متن‌باز برای استخراج خودکار امتیاز تکرار Oncotype DX.
  • رویکرد هیبریدی OCR بالاترین دقت و توافق را با انتزاع دستی نشان داد.
  • استخراج خودکار زمان پردازش را به طور قابل توجهی کاهش داد.
  • ناهماهنگی در ثبت سرطان عمدتاً مستقل از ویژگی‌های بیمار و تومور بود.
  • استخراج خودکار نشانگرهای ژنومیک می‌تواند به مدرن‌سازی ثبت سرطان کمک کند.

ترجمه فارسی چکیده

هدف: داده‌های واقعی در سرطان‌شناسی برای تحقیقات بالینی و مراقبت دقیق از سرطان ضروری هستند. با این حال، نشانگرهای ژنومیک اغلب در اسناد اسکن‌شده و غیرسازمان‌یافته بالینی گنجانده شده‌اند و نیاز به انتزاع دستی دارند تا در ثبت‌های سرطان در دسترس قرار گیرند، که منجر به تأخیر در تولید شواهد واقعی می‌شود. این مطالعه سه رویکرد استخراج متن با نوری (OCR) متن‌باز، تست‌رکت، ایزی‌OCR و یک پیاده‌سازی هیبریدی را ارزیابی و مقایسه کرد تا مشخص شود کدام بهترین امکان استخراج خودکار امتیاز تکرار Oncotype DX را فراهم می‌کند و کیفیت داده‌های واقعی در سرطان‌شناسی را بهبود می‌بخشد. روش‌ها: ما قابلیت استخراج خودکار داده‌های ژنومیک را با استفاده از ۶۷۵ گزارش Oncotype DX از یک سیستم بهداشتی در ایالات متحده میانی ارزیابی کردیم. برای استخراج امتیازات تکرار از گزارش‌های اسکن‌شده، از ایزی‌OCR، تست‌رکت و یک رویکرد هیبریدی OCR استفاده شد. مقادیر استخراج‌شده توسط OCR با امتیازات انتزاع دستی و داده‌های ثبت سرطان محلی مقایسه شدند. عملکرد با استفاده از توافق، دقت، یادآوری، امتیاز F1 و زمان پردازش ارزیابی شد. رگرسیون لجستیک چندمتغیره برای شناسایی عوامل مرتبط با ناهماهنگی بین امتیازات گزارش‌شده در ثبت و امتیازات انتزاع دستی انجام شد. نتایج: رویکرد هیبریدی OCR بالاترین عملکرد را نشان داد، با ۹۷٪ توافق با انتزاع دستی، دقت ۰.۹۹۷، یادآوری ۰.۹۷۲ و امتیاز F1 ۰.۹۸۴. انتزاع ثبت سرطان عملکرد قابل مقایسه‌ای نشان داد اما نیاز به تلاش دستی بیشتری داشت. استخراج خودکار زمان پردازش را به طور قابل توجهی کاهش داد در حالی که دقت بالا را حفظ کرد. رگرسیون لجستیک نشان داد که ناهماهنگی ثبت عمدتاً مستقل از ویژگی‌های بیمار و تومور بود و وضعیت رقبای پروژسترون (PR) ناشناخته تنها پیش‌بینی‌کننده معنی‌دار بود. نتیجه‌گیری: استخراج خودکار نشانگرهای ژنومیک یک رویکرد مقیاس‌پذیر برای کاهش تأخیرهای دسترسی به داده‌های سرطان است. به‌دست آوردن زودهنگام اطلاعات ژنومیک ممکن است به مدرن‌سازی ثبت سرطان و بهبود تولید شواهد واقعی در سرطان‌شناسی دقیق کمک کند.

روش پژوهش

ارزیابی سه روش OCR (تست‌رکت، ایزی‌OCR، هیبریدی) با استفاده از ۶۷۵ گزارش Oncotype DX. مقایسه مقادیر استخراج‌شده با انتزاع دستی و داده‌های ثبت سرطان محلی. ارزیابی عملکرد با استفاده از توافق، دقت، یادآوری و امتیاز F1.

محدودیت‌ها

مطالعه فقط در یک سیستم بهداشتی در ایالات متحده میانی انجام شد. نتایج ممکن است به طور مستقیم به سایر سیستم‌های بهداشتی یا روش‌های OCR دیگر تعمیم داده نشود.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
بیماران با گزارش‌های Oncotype DX در یک سیستم بهداشتی در ایالات متحده میانی.
مداخله/مواجهه
رویکردهای OCR (تست‌رکت، ایزی‌OCR، هیبریدی) برای استخراج خودکار امتیاز تکرار.
مقایسه
انتزاع دستی امتیازات و داده‌های ثبت سرطان محلی.
حجم نمونه
۶۷۵ گزارش Oncotype DX.

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

Breast cancerGenomicsOncologyOptical character recognition
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Computational Network Analysis for Defining Transcriptional Programs.

Cancer cell identity is governed by coordinated transcriptional programs that are frequently rewired during tumorigenesis. Systematic identification of cancer type-specific gene regulatory networks provides a framework for understanding oncogenic state transitions and for prioritizing candidate therapeutic targets. Here, we present a reproducible network-based workflow for reconstructing and analyzing transcriptional regulatory program…

PubMed2027

Topic-Driven Bibliometrics and Trend Intelligence for Stem Cell and Cancer Research.

The rapid growth of biomedical literature has created an urgent need for computational tools that enable researchers to systematically analyze publication trends, identify emerging research themes, and map the evolution of scientific fields. PubMed Atlas is a command-line and web-enabled workflow for topic-driven bibliometrics and trend intelligence using PubMed E-utilities. The pipeline executes PubMed queries, retrieves matching PMID…

PubMed2026

Nursing Students' Reports of Patient Safety Incidents and Reasons During Clinical Placements: A Secondary Analysis of the International Data.

Clinical placements expose nursing students to patient safety incidents and provide important opportunities for learning about safe care. This study explored how nursing students in four countries recognized and interpreted patient safety incidents encountered or witnessed during clinical practice, including perceived contributing factors. A secondary qualitative content analysis was conducted using narrative data from 1442 undergradua…

PubMed2026

Frequency of Clinically Relevant Drug-Drug Interactions Between Tyrosine Kinase Inhibitors and Proton Pump Inhibitors in Patients With Cancer Using Real-World Data.

BACKGROUND: Proton pump inhibitors (PPIs) raise stomach pH, leading to reduced bioavailability of many tyrosine kinase inhibitors (TKIs), thereby affecting treatment outcomes. To what extent this interaction occurs in clinical practice remains underexplored. OBJECTIVE: To determine the frequency of clinically relevant interactions between TKIs and PPIs in clinical practice and the duration of concomitant prescription. METHODS: A retros…