PubMed چکیده/رکورد

Automated Extraction of Genetic Eligibility Criteria from Clinical Trial Records Using LLMs - A Technical Case Report.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

INTRODUCTION: Accurate interpretation of clinical trial eligibility criteria is essential for applications such as patient-trial matching and clinical decision support, particularly in precision oncology. However, relevant information, including genetic mutation requirements, is typically embedded in unstructured text within trial registries such as ClinicalTrials.gov, limiting accessibility for automated processing. METHODS: This paper reports on development, integration and evaluation of a system for extracting and structuring mutational eligibility criteria from trial records, focusing on identifying mutated genes, distinguishing inclusion and exclusion criteria, and assigning them to individual study arms. To address challenges like ambiguous abbreviations and context-dependent meaning, we combine large language models (LLMs) for context-aware extraction with rule-based validation against HUGO Gene Nomenclature. The system was implemented using local LLMs and integrated into the Community Annotated Trial Search (CATS) platform. RESULTS: Applied to 4,918 clinical trials, the system generated structured representations of genetic eligibility criteria for 1,010 studies. Expert review of 42 trials showed that 88.1% of studies were correctly annotated, with a precision of 80% at the level of individual eligibility criteria. Failures were mainly due to hallucinated genes and misinterpreted abbreviations, highlighting challenges in biomedical text processing. DISCUSSION: The findings indicate that LLM-based extraction is a promising approach for structuring complex eligibility criteria, particularly when combined with strategies to improve precision. The integration into an operational system demonstrates practical feasibility, while the observed limitations emphasize the need for careful dataset design, error mitigation strategies, and continued refinement to achieve reliable automation in clinical applications.

نتیجه فارسی

این مقاله یک سیستم را برای استخراج خودکار معیارهای واجد شرایط ژنتیکی از سوابق آزمایش‌های بالینی با استفاده از مدل‌های زبانی بزرگ (LLM) گزارش می‌دهد. سیستم با استفاده از LLMهای محلی در پلتفرم CATS ادغام شد. در بررسی ۴۲ آزمایش، ۸۸.۱٪ مطالعات به درستی برچسب‌گذاری شدند و دقت استخراج معیارها ۸۰٪ بود. محدودیت‌ها عمدتاً به دلیل ژن‌های توهمی و تفسیر اشتباه مخفف‌ها بود.

  • سیستم برای استخراج معیارهای واجد شرایط ژنتیکی از سوابق آزمایش‌های بالینی طراحی شد.
  • استفاده از LLMهای محلی و ادغام در پلتفرم CATS.
  • دقت ۸۰٪ در سطح معیارهای واجد شرایط فردی گزارش شد.
  • ژن‌های توهمی و تفسیر اشتباه مخفف‌ها به عنوان عوامل اصلی شکست ذکر شدند.
  • استخراج مبتنی بر LLM برای کاربردهای بالینی عملیاتی نشان داده شد.

ترجمه فارسی چکیده

مقدمه: تفسیر دقیق معیارهای واجد شرایط آزمایش‌های بالینی برای کاربردهایی مانند تطبیق بیمار-آزمایش و پشتیبانی تصمیم‌گیری بالینی، به‌ویژه در سرطان‌شناسی دقیق، ضروری است. با این حال، اطلاعات مرتبط، از جمله الزامات جهش ژنتیکی، معمولاً در متن ساختاریافته‌ای در ثبت‌های آزمایش‌های بالینی مانند ClinicalTrials.gov گنجانده شده‌اند که دسترسی آن‌ها را برای پردازش خودکار محدود می‌کند. روش‌ها: این مقاله گزارشی از توسعه، ادغام و ارزیابی یک سیستم برای استخراج و ساختاردهی معیارهای واجد شرایط جهش‌زا از سوابق آزمایش، با تمرکز بر شناسایی ژن‌های جهش‌یافته، تفکیک معیارهای ورود و خروج و اختصاص آن‌ها به شاخه‌های مطالعه، ارائه می‌دهد. برای مقابله با چالش‌هایی مانند مخفف‌های مبهم و معنای وابسته به زمینه، ما از مدل‌های زبانی بزرگ (LLM) برای استخراج آگاهانه از زمینه و اعتبارسنجی مبتنی بر قوانین در برابر نام‌گذاری ژن HUGO استفاده می‌کنیم. سیستم با استفاده از LLMهای محلی پیاده‌سازی شده و در پلتفرم CATS ادغام شد. نتایج: با کاربرد بر روی ۴,۹۱۸ آزمایش بالینی، سیستم نمایش‌های ساختاریافته از معیارهای واجد شرایط ژنتیکی برای ۱,۰۱۰ مطالعه تولید کرد. بررسی تخصصی ۴۲ آزمایش نشان داد که ۸۸.۱٪ از مطالعات به درستی برچسب‌گذاری شدند، با دقت ۸۰٪ در سطح معیارهای واجد شرایط فردی. شکست‌ها عمدتاً به دلیل ژن‌های توهمی و مخفف‌های تفسیر اشتباه ناشی از پردازش متن زیست‌پزشکی بود. بحث: یافته‌ها نشان می‌دهند که استخراج مبتنی بر LLM یک رویکرد امیدوارکننده برای ساختاردهی معیارهای واجد شرایط پیچیده است، به‌ویژه زمانی که با استراتژی‌هایی برای بهبود دقت ترکیب می‌شود. ادغام در یک سیستم عملیاتی، قابلیت عملی را نشان می‌دهد، در حالی که محدودیت‌های مشاهده‌شده، نیاز به طراحی دقیق مجموعه داده، استراتژی‌های کاهش خطا و اصلاح مداوم برای دستیابی به اتوماسیون قابل اعتماد در کاربردهای بالینی را برجسته می‌کند.

روش پژوهش

سیستم با استفاده از LLMهای محلی پیاده‌سازی شد و در پلتفرم CATS ادغام گردید. برای استخراج، از LLMها برای استخراج آگاهانه از زمینه و اعتبارسنجی مبتنی بر قوانین در برابر نام‌گذاری ژن HUGO استفاده شد.

محدودیت‌ها

شکست‌ها عمدتاً به دلیل ژن‌های توهمی و تفسیر اشتباه مخفف‌ها ناشی از پردازش متن زیست‌پزشکی بود. محدودیت‌ها نیاز به طراحی دقیق مجموعه داده و استراتژی‌های کاهش خطا را برجسته می‌کنند.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
مطالعات بالینی (۴,۹۱۸ مورد)
مداخله/مواجهه
استخراج خودکار معیارهای واجد شرایط ژنتیکی با استفاده از LLM
مقایسه
اعتبارسنجی مبتنی بر قوانین در برابر نام‌گذاری ژن HUGO
حجم نمونه
۴,۹۱۸ آزمایش بالینی (برای اعمال سیستم) و ۴۲ آزمایش (برای بررسی تخصصی)

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

Clinical TrialData MiningGenesHealth Information InteroperabilityPrecision OncologyTerminology
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …