PubMed چکیده/رکورد

Detecting Misspelled Drug Names Using Transformer-Based Language Models: Model Development and External Validation.

استودیوی صوتی مقاله

پخش حرفه‌ای فارسی و انگلیسی

در حال بررسی نسخه‌های صوتی ذخیره‌شده…

صوت تولیدشده با هوش مصنوعی است. برای کاربرد علمی یا درمانی، متن و منبع اصلی را بررسی کنید.
خواندن هوشمند فارسی و انگلیسی در حال آماده‌سازی صداهای مرورگر…
تنظیم صدای طبیعی و سرعت

صداهایی که در نامشان «Natural»، «Neural» یا «Online» دیده می‌شود معمولاً طبیعی‌ترند. انتخاب صدا به صداهای نصب‌شده در ویندوز و مرورگر شما بستگی دارد.

چکیده اصلی

BACKGROUND: Misspellings in medication names can compromise patient safety, reduce data utility, and impede large-scale data initiatives that integrate medication information from electronic health records (EHRs). Existing methods for detecting misspelled medical terms are mostly dictionary-based and can lead to high false-positive rates when correctly spelled but previously unseen (out-of-vocabulary) terms are encountered. OBJECTIVE: We aimed to develop and validate domain-specific, transformer-based language models for detecting misspelled drug names, with an emphasis on performance for unseen terms. METHODS: Using RxNorm as a standardized drug vocabulary, we created an RxNorm-augmented training corpus and developed two BERT (Bidirectional Encoder Representations from Transformers)-based models-BERTDrug and CharBERTDrug-for misspelling detection. Specifically, we randomly split 69,824 RxNorm drug names into training, development, and test sets (3:1:1) and generated k misspellings per name using text-perturbation techniques (k optimized for training; fixed at 1 for development and test sets). The models were fine-tuned on the training and development sets and evaluated using the RxNorm test set and 3586 drug names from the Long-Term Care Data Cooperative (LTCDC) database (external validation). The RxNorm test set and out-of-vocabulary LTCDC dataset (1922 terms), neither overlapping with the RxNorm training data, were used to evaluate performance on unseen terms. SpellChecker served as a dictionary-based baseline, while fastTextML and BioWordVecML, which used different subword embeddings as inputs for machine learning, served as additional baselines. Additionally, we compared model performance with GPT-4o, a generative large language model (LLM), using 2200 randomly sampled test terms. RESULTS: On the RxNorm test set, BERTDrug and CharBERTDrug outperformed the baseline models across most metrics. BERTDrug achieved the best overall performance (F1-score=0.859; area under the receiver operating characteristic curve [ROC-AUC]=0.947), followed by CharBERTDrug (F1-score=0.833; ROC-AUC=0.906). Both models also outperformed the baseline models on the out-of-vocabulary LTCDC dataset across most metrics, with CharBERTDrug performing best (F1-score=0.696; ROC-AUC=0.788), followed by BERTDrug (F1-score=0.669; ROC-AUC=0.786). In the secondary analysis, both models exceeded GPT-4o on most metrics (except Recall) for 2000 RxNorm terms. BERTDrug performed best (ROC-AUC=0.951; F1-score=0.855), followed by CharBERTDrug (ROC-AUC=0.911; F1-score=0.831) and GPT-4o (ROC-AUC=0.856; F1-score=0.721). In contrast, among 200 randomly selected LTCDC terms, GPT-4o performed best on most metrics except precision and specificity. CONCLUSIONS: Domain-specific language models improved detection of misspellings in out-of-vocabulary drug names and outperformed baseline models in both internal and external evaluations. The comparison with a generative LLM suggests that domain shift may substantially reduce the advantages conferred by domain-specific training. With further fine-tuning on diverse data that capture the terminology, formatting conventions, and spelling patterns encountered across real-world clinical settings, these models could be adapted for use in other clinical databases and EHR systems to improve medication data quality for research and to support future safety-focused applications.

متن کامل اصلی

متن در JumpToDate ذخیره نشده است.

برای بررسی دسترسی کتابخانه‌ای یا خرید، رکورد اصلی را باز کنید.

رفتن به منبع اصلی

کلیدواژه‌ها

clinical repositoriesdata augmentation using RxNormlarge language modelsmedication namesmisspelling detectiontransformer-based language models
در همین زیرشاخه

مقاله‌های مرتبط

PubMed2026

Clinical effectiveness and safety of metadoxine in the management of acute alcohol intoxication: A single-center retrospective cohort study.

BACKGROUND: Acute alcohol intoxication (AAI) is a common emergency with no specific antidote. Metadoxine has shown potential but lacks sufficient real-world evidence, particularly in Chinese populations. OBJECTIVES: To evaluate the clinical efficacy and safety of metadoxine in patients with acute alcohol intoxication. METHODS: This single-center retrospective cohort study included 124 patients with AAI admitted to an emergency departme…

PubMed2026

D3MI: an efficient and powerful federated imputation method for bias reduction in the analysis of distributed incomplete data by accounting for within-site correlation and between-site heterogeneity.

BACKGROUND: Electronic health records (EHRs) collected from diverse healthcare institutions offer a rich and representative data source for clinical research. Federated learning enables analysis of these distributed data without sharing sensitive patient-level information, preserving privacy. However, missing data remain a major challenge and can introduce substantial bias if not properly addressed. Very few distributed imputation meth…

PubMed2026

Extraction of Pain Severity and Functional Interference From Clinical Narratives Using Domain-Informed Large Language Models: Protocol for a Development and Validation Study.

BACKGROUND: Chronic pain is a leading cause of disability and requires multidimensional assessment of pain intensity and functioning, yet electronic health records rarely capture these measures systematically. By contrast, surveys collecting patient-reported outcomes can assess pain over multiple dimensions but remain resource-intensive and difficult to scale for continuous population-level monitoring. OBJECTIVE: The objective of this …

PubMed2026

From data entry to digital transformation: Allied health perspectives on standardised electronic medical records data.

BACKGROUND: Electronic medical records (EMRs) currently rely on standardised data fields to support secondary data use for clinical care, performance monitoring, and system-level reporting. However, utilisation of standardised data capture and reporting within allied health remains underdeveloped in practice. Greater understanding of how allied health clinicians and managers perceive the purpose, value, and impact of standardised data …