PubMed دسترسی آزاد

Clinical evaluation and regression test of a commercial deep-learning auto-segmentation model.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Commercial deep-learning segmentation (DLS) tools are increasingly used in clinical practice. Software updates may alter segmentation performance, highlighting the need for systematic clinical evaluation before implementation. PURPOSE: This study presents our experience in clinically evaluating and regression-testing a RayStation DLS model with the aim of guiding commissioning and routine quality assurance of commercial DLS tools in clinical environments. METHODS: U-Net DLS model for normal-tissue structures, originally commissioned in RayStation version 11B, was regression-tested after upgrading to 2024A. Previously commissioned CT datasets for head and neck (n = 18), thorax (n = 18), male pelvis (n = 18), and breast (n = 21) were re-evaluated in RayStation 2024A for regression testing following the software upgrade, while 20 new abdominal cases were included for initial commissioning. Clinical segmentations served as the reference standard. Geometric performance was evaluated using Dice Similarity Coefficient (DSC), 95% Hausdorff Distance (HD95), and Mean Surface Distance (MSD). Dosimetric evaluation used the original clinical treatment plans to compare Dmean and Dmax between DLS and clinical contours, normalized to prescription dose; breast was excluded from dosimetric analysis because its DLS model was not commissioned for clinical use. RESULTS: Regression testing showed that, in comparison with the previously commissioned data, mean [± standard deviation (SD)] DSC values were 0.81 ± 0.09 for the head and neck, 0.84 ± 0.13 for the thorax, and 0.89 ± 0.06 for male pelvis. Corresponding MSD values were 0.83 ± 0.34 mm, 2.47 ± 1.81 mm, and 1.76 ± 0.68 mm; and HD95 values were 4.75 ± 2.38 mm, 8.45 ± 3.38 mm, and 13.21 ± 7.05 mm. Significant paired differences between versions were limited to selected structures, including the mandible, cochlea, and eye in head and neck; lungs and esophagus in thorax; and prostate in the male pelvis (p < 0.05). Breast segmentation results showed low performance (DSC 0.66 ± 0.15, MSD 4.8 ± 1.93 mm, and HD95 22.2 ± 9.6 mm) compared to clinical reference, so the breast structure models are not recommended for clinical use. Abdominal structures (RayStation 2024A vs clinical) had DSC 0.95 ± 0.014; MSD 0.82 ± 0.32 mm; and HD95 5.4 ± 2.9 mm. The mean dose differences between DLS and clinical segmentations were 0.03 Gy for the head and neck, -0.92 Gy for the thorax, 0.27 Gy for the male pelvis, and 0.02 Gy for the abdominal sites, and the corresponding maximum dose differences were 10.96 Gy, 13.23 Gy, 3.39 Gy, and 6.31 Gy. CONCLUSIONS: We described a procedure for practical quality assurance of auto-segmentation tools for clinical use after periodic software upgrades. Regression testing of head and neck, thorax, and male pelvis sites showed consistent performance after the software upgrade and the DLS models for abdominal site were commissioned in this study for clinical use; while the DLS models for breast site were not recommended.

نتیجه فارسی

این مطالعه یک روش برای نظارت کیفیت ابزارهای خودبخشی تجاری پس از به‌روزرسانی‌های نرم‌افزاری ارائه می‌دهد. تست بازگشت نشان داد که مدل‌های DLS برای سر و گردن، قفسه سینه و لگن مردانه پس از ارتقا عملکرد ثابتی دارند. مدل DLS برای سایت abdominal در این مطالعه برای استفاده بالینی کاریابی شد. در مقابل، مدل‌های DLS برای سایت پستان عملکرد پایینی داشتند و برای استفاده بالینی توصیه نمی‌شوند.

  • تست بازگشت مدل DLS U-Net پس از ارتقا RayStation 2024A انجام شد.
  • عملکرد مدل‌های DLS برای سر و گردن، قفسه سینه و لگن مردانه پس از ارتقا ثابت بود.
  • مدل DLS برای سایت abdominal در این مطالعه برای استفاده بالینی کاریابی شد.
  • مدل‌های DLS برای سایت پستان عملکرد پایینی داشتند و توصیه نمی‌شوند.
  • تفاوت‌های دوز میانگین و حداکثر بین خودبخشی DLS و بالینی گزارش شد.

ترجمه فارسی چکیده

ابزارهای خودبخشی یادگیری عمیق (DLS) تجاری به طور فزاینده‌ای در عمل بالینی استفاده می‌شوند. به‌روزرسانی‌های نرم‌افزاری ممکن است عملکرد خودبخشی را تغییر دهد، که نیاز به ارزیابی بالینی سیستماتیک قبل از پیاده‌سازی را برجسته می‌کند. این مطالعه تجربه ما در ارزیابی بالینی و تست بازگشت یک مدل DLS RayStation را برای هدایت کاریابی و نظارت کیفیت روتین ابزارهای DLS تجاری در محیط‌های بالینی ارائه می‌دهد. مدل DLS U-Net برای ساختارهای بافت نرمال، که در ابتدا در نسخه 11B RayStation کاریابی شده بود، پس از ارتقا به نسخه 2024A مورد تست بازگشت قرار گرفت. مجموعه‌های داده CT کاریابی شده قبلی برای سر و گردن (n = 18)، قفسه سینه (n = 18)، لگن مردانه (n = 18) و پستان (n = 21) در RayStation 2024A برای تست بازگشت پس از ارتقا نرم‌افزار، در حالی که ۲۰ مورد abdominal جدید برای کاریابی اولیه شامل شد. خودبخشی‌های بالینی به عنوان استاندارد مرجع عمل کردند. عملکرد هندسی با ضریب شباهت دیس (DSC)، فاصله هافسدورف ۹۵٪ (HD95) و فاصله سطح متوسط (MSD) ارزیابی شد. ارزیابی دوزیمتری از برنامه‌های درمان بالینی اصلی برای مقایسه Dmean و Dmax بین خودبخشی DLS و خودبخشی‌های بالینی، با نرمال‌سازی به دوز تجویزی استفاده کرد؛ پستان از تحلیل دوزیمتری حذف شد زیرا مدل DLS آن برای استفاده بالینی کاریابی نشده بود. نتایج تست بازگشت نشان داد که در مقایسه با داده‌های کاریابی شده قبلی، مقادیر میانگین DSC به ترتیب 0.81 ± 0.09 برای سر و گردن، 0.84 ± 0.13 برای قفسه سینه و 0.89 ± 0.06 برای لگن مردانه بود. مقادیر متناظر MSD به ترتیب 0.83 ± 0.34 میلی‌متر، 2.47 ± 1.81 میلی‌متر و 1.76 ± 0.68 میلی‌متر و مقادیر HD95 به ترتیب 4.75 ± 2.38 میلی‌متر، 8.45 ± 3.38 میلی‌متر و 13.21 ± 7.05 میلی‌متر بود. تفاوت‌های جفت‌شده معنی‌داری بین نسخه‌ها محدود به ساختارهای انتخابی، از جمله تمپلان، حلزون گوش و چشم در سر و گردن؛ ریه‌ها و مری در قفسه سینه و پروستات در لگن مردانه بود (p < 0.05). نتایج خودبخشی پستان عملکرد پایینی را نشان داد (DSC 0.66 ± 0.15، MSD 4.8 ± 1.93 میلی‌متر و HD95 22.2 ± 9.6 میلی‌متر) در مقایسه با مرجع بالینی، بنابراین مدل‌های ساختاری پستان برای استفاده بالینی توصیه نمی‌شوند. ساختارهای abdominal (RayStation 2024A در برابر بالینی) DSC 0.95 ± 0.014؛ MSD 0.82 ± 0.32 میلی‌متر و HD95 5.4 ± 2.9 میلی‌متر داشتند. تفاوت‌های دوز میانگین بین خودبخشی DLS و خودبخشی‌های بالینی به ترتیب 0.03 Gy برای سر و گردن، -0.92 Gy برای قفسه سینه، 0.27 Gy برای لگن مردانه و 0.02 Gy برای سایت‌های abdominal بود و تفاوت‌های دوز حداکثر متناظر به ترتیب 10.96 Gy، 13.23 Gy، 3.39 Gy و 6.31 Gy بود. ما یک روال برای نظارت کیفیت عملی ابزارهای خودبخشی برای استفاده بالینی پس از به‌روزرسانی‌های دوره‌ای نرم‌افزار را توصیف کردیم. تست بازگشت سایت‌های سر و گردن، قفسه سینه و لگن مردانه عملکرد ثابتی پس از به‌روزرسانی نرم‌افزار را نشان داد و مدل‌های DLS برای سایت abdominal در این مطالعه برای استفاده بالینی کاریابی شدند؛ در حالی که مدل‌های DLS برای سایت پستان توصیه نمی‌شوند.

روش پژوهش

مدل DLS U-Net برای ساختارهای بافت نرمال، که در نسخه 11B کاریابی شده بود، پس از ارتقا به نسخه 2024A مورد تست بازگشت قرار گرفت. مجموعه‌های داده CT کاریابی شده قبلی برای سر و گردن (n = 18)، قفسه سینه (n = 18)، لگن مردانه (n = 18) و پستان (n = 21) در RayStation 2024A برای تست بازگشت، در حالی که ۲۰ مورد abdominal جدید برای کاریابی اولیه شامل شد. عملکرد هندسی با DSC، HD95 و MSD و ارزیابی دوزیمتری با مقایسه Dmean و Dmax ارزیابی شد.

محدودیت‌ها

پستان از تحلیل دوزیمتری حذف شد زیرا مدل DLS آن برای استفاده بالینی کاریابی نشده بود. نتایج خودبخشی پستان عملکرد پایینی را نشان داد.

استخراج ساختاریافته از متن منبع

نمای PICO و پیامدها

جمعیت
مجموعه‌های داده CT کاریابی شده قبلی برای سر و گردن (n = 18)، قفسه سینه (n = 18)، لگن مردانه (n = 18) و پستان (n = 21) و ۲۰ مورد abdominal جدید.
مداخله/مواجهه
مدل DLS U-Net برای ساختارهای بافت نرمال در RayStation 2024A.
مقایسه
خودبخشی‌های بالینی به عنوان استاندارد مرجع.
حجم نمونه
مجموعاً ۷۵ مورد (۱۸ سر و گردن، ۱۸ قفسه سینه، ۱۸ لگن مردانه، ۲۱ پستان و ۲۰ abdominal).

متن کامل اصلی

نسخه دارای مجوز در منبع علمی در دسترس است.

لینک مستقیم از metadata منبع گرفته شده و در تب جدید باز می‌شود.

باز کردن متن کامل

کلیدواژه‌ها

auto‐segmentationcommercial softwaredeep learningquality assurance
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2027

Computational Network Analysis for Defining Transcriptional Programs.

Cancer cell identity is governed by coordinated transcriptional programs that are frequently rewired during tumorigenesis. Systematic identification of cancer type-specific gene regulatory networks provides a framework for understanding oncogenic state transitions and for prioritizing candidate therapeutic targets. Here, we present a reproducible network-based workflow for reconstructing and analyzing transcriptional regulatory program…

PubMed2027

Topic-Driven Bibliometrics and Trend Intelligence for Stem Cell and Cancer Research.

The rapid growth of biomedical literature has created an urgent need for computational tools that enable researchers to systematically analyze publication trends, identify emerging research themes, and map the evolution of scientific fields. PubMed Atlas is a command-line and web-enabled workflow for topic-driven bibliometrics and trend intelligence using PubMed E-utilities. The pipeline executes PubMed queries, retrieves matching PMID…

PubMed2026

Nursing Students' Reports of Patient Safety Incidents and Reasons During Clinical Placements: A Secondary Analysis of the International Data.

Clinical placements expose nursing students to patient safety incidents and provide important opportunities for learning about safe care. This study explored how nursing students in four countries recognized and interpreted patient safety incidents encountered or witnessed during clinical practice, including perceived contributing factors. A secondary qualitative content analysis was conducted using narrative data from 1442 undergradua…

PubMed2026

Frequency of Clinically Relevant Drug-Drug Interactions Between Tyrosine Kinase Inhibitors and Proton Pump Inhibitors in Patients With Cancer Using Real-World Data.

BACKGROUND: Proton pump inhibitors (PPIs) raise stomach pH, leading to reduced bioavailability of many tyrosine kinase inhibitors (TKIs), thereby affecting treatment outcomes. To what extent this interaction occurs in clinical practice remains underexplored. OBJECTIVE: To determine the frequency of clinically relevant interactions between TKIs and PPIs in clinical practice and the duration of concomitant prescription. METHODS: A retros…