Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and environmental information within a One Health framework, while addressing critical challenges in governance, equity, and interoperability. The discussion covers the entire genomic surveillance workflow, from sample collection and sequencing to bioinformatic analysis and phylogenetic inference, and highlights the transformative role of artificial intelligence (AI) in predictive surveillance. By analyzing global initiatives, operational barriers, and emerging technologies, this chapter underscores the necessity of sustainable, equitable, and interoperable genomic systems to proactively address current and future infectious disease threats.
Journal of medical systemsTim Wilhelmi, Vanessa Bartsch, Marius Platt, Asarnusch Rashid, Johannes Hornig, Martin Krusche, Axel J Hueber, Daniel Fink, Alexander Pfeil, Gabriel Dischereit…
Patients with rheumatic diseases have persistent information needs that are not fully addressed in routine care. We developed and evaluated guideline-grounded, large language model (LLM) chatbots to support patient self-management and education in rheumatology.Ten disease-specific chatbots based on German guidelines were co-developed and deployed through 13 rheumatology centres and six patient organisations. Chatbot users rated responses and completed a questionnaire. User questions, feedback, and response characteristics were analysed using category-based coding and a six-dimensional LLM-as-a-judge assessment, with LLM-based ratings compared with rheumatologist ratings in random subsets.Between September 2025 and January 2026, 6291 questions were recorded. Thirteen question categories were identified, most commonly disease-specific questions (50.2%), medication and monitoring (37.3%) and diagnostics (28.2%). The chatbots were unable to answer in 263 interactions (4.2%). Of 2671 responses rated by users, 2481 (92.9%) received a positive rating. Insufficient detail was the most common reason for negative ratings (125/190, 65.8%). Among 602 questionnaire respondents, 84.6% reported that the chatbot was easy to use, 84.1% that answers were easy to understand, and 80.2% that it was a useful addition to patient education. In the LLM-based evaluation, 95.3% of answers were rated as completely safe and 79.1% as completely correct. Guideline adherence was assessed separately, with 45.0% rated as fully adherent; agreement with physician assessment was weak.Guideline-grounded chatbots received predominantly positive user feedback in real-world use, while LLM-based evaluation suggested that most responses were safe and correct. User questions and feedback may help guide iterative improvements to source content and patient education materials. Further studies are needed to evaluate educational effectiveness and independently validate response quality and clinical safety.
Journal of medical Internet researchLuis Madeira, Cíntia Águas, Meryam Schouler-Ocak, Jerzy Samochowiec, Silvana Galderisi, Mariana Pinto da Costa
AI is increasingly incorporated into psychiatric triage, risk prediction, passive monitoring, clinical documentation, and patient-facing conversational systems. These applications may improve access, continuity, efficiency, and pattern recognition, but they also redistribute epistemic authority and complicate responsibility when harm occurs. European regulation is developed in relation to market access, data governance, risk management, and product safety, yet remains fragmented regarding civil liability, organizational negligence, and the psychiatric standard of care. This Viewpoint examines how liability and standard of care should be understood when AI becomes part of psychiatric reasoning in Europe. It advances one central thesis: psychiatric AI requires justified integration supported by layered accountability within, but not determined by, European regulation. It presents a targeted doctrinal and normative synthesis of binding European Union instruments, regulatory guidance, selected national governance materials, and psychiatric, bioethical, legal, and digital mental health literature. It distinguishes binding law from guidance and policy, and separates ex ante regulation from ex post liability, and from professional standards of care. Four illustrative domains are analyzed: conversational or therapeutic chatbots, suicide prediction, digital phenotyping and passive monitoring, and large language model documentation. Psychiatric AI raises distinctive concerns because psychiatric judgment depends heavily on testimony, contextual meaning, therapeutic trust, risk interpretation, privacy, and liberty-sensitive decisions. Existing European instruments, including the AI Act, Medical Device Regulation, General Data Protection Regulation, revised Product Liability Directive, and European Health Data Space Regulation, establish governance duties, but do not provide a harmonized fault-based liability framework for AI-assisted health care. Regulatory compliance may inform later legal assessment, but it does not determine whether psychiatric care was reasonable. The proposed standard of justified integration requires knowledge of intended use and model limits, assessment of local and patient-level applicability, active clinical interpretation, disclosure when AI use is material to consent or trust, documentation in high-stakes decisions, and organizational audit. Accountability should be distributed across developers, deployers, and clinicians according to control and preventability. Mixed-fault scenarios are therefore likely to be common. The augmented-clinician model and layered accountability are offered as normative proposals rather than settled European legal standards. Clinicians should remain responsible for contextual, patient-centered judgment; developers for design, validation, documentation, and foreseeable misuse; and deployers for procurement, training, workflow integration, local validation, monitoring, and escalation. Future empirical research should evaluate effects on clinician reliance, documentation burden, patient outcomes, coercive interventions, therapeutic trust, and feasibility across differently resourced services.
Laboratory medicineYusuf Yesil, Alpay Medetalibeyoglu, Evin Ademoglu
INTRODUCTION: Routine laboratory panels are nearly universal, but the panels' joint information is underused. We evaluated contemporaneous classification of International Statistical Classification of Diseases, Tenth Revision (ICD-10) code groups from same-encounter laboratory results. METHODS: We developed 17 eXtreme Gradient Boosting (XGBoost) classifiers in 242 648 adult internal medicine encounters using age, sex, and results from 34 assays, with ICD-10 codes used only as outcomes. Stratified 10-fold cross-validation assessed discrimination, calibration, and decision curve performance. Leakage-free recalibration and fold-specific thresholds targeted 95% specificity. Using an 80-patient temporal holdout cohort, we compared the model with 3 large language models accessed through an application programming interface and findings from 1 specialist physician. RESULTS: The mean cross-validated area under the curve (AUC) was 0.893. With identical preprocessing, the mean AUC was 0.892 for XGBoost and 0.809 for logistic regression. Recalibration changed the mean intercept and slope to -0.002 and 0.999, respectively, and the Brier score from 0.099 to 0.044. The holdout micro-averaged AUC was 0.880 vs 0.729 to 0.739 for language models; paired differences were 0.141 to 0.151 (all P < .005). Matched-specificity sensitivity differences favored the model but were imprecise. DISCUSSION: Routine panels contain substantial information about contemporaneous ICD-10 coding, supporting prospective multicenter evaluation rather than a diagnostic claim.
PituitaryDanyal Z Khan, Zhehua Mao, Anjana Wijekoon, Adrito Das, Simon C Williams, Ann Blandford, Abhiney Jain, Lauren Harris, Anouk Borg, Neil L Dorward, Matthew J Cla…
INTRODUCTION: Precise anatomical navigation is fundamental to safe endoscopic pituitary surgery, a high-stakes procedure characterised by a challenging learning curve. While traditional navigation systems often rely on workflow-disrupting probes or static preoperative imaging, advancements in computer vision AI (CVAI) now enable dynamic, real-time pixel-level anatomical segmentation directly from live surgical video. Our group has previously conducted a series of preclinical human-computer interaction studies to refine the system's design, alongside digital and high-fidelity physical simulations demonstrating the potential benefit of AI assistance in improving surgical performance, training, and safety. Building on this foundation, the current study represents a first-in-human evaluation of real-time pixel-level CVAI anatomical segmentation in the neurosurgical operating room - assessing feasibility, human factors and clinical outcomes, while iteratively improving the system. METHOD: Guided by the DECIDE-AI and IDEAL frameworks, this single-centre evaluation comprises an initial proof-of-concept of CVAI anatomical segmentation in endoscopic transsphenoidal pituitary surgery. The AI model utilised a DINOv3-derived vision transformer architecture, deployed via a high-performance edge computing unit to achieve low-latency real-time inference without reliance on cloud infrastructure. Feasibility and functionality were assessed via structured questionnaire, prospective observation, and blinded retrospective review of the recordings of the endoscopic surgical video feed and wider operating room environment. Continuous multi-stakeholder feedback through validated human factors surveys drove iterative technical refinements between cases. Routine clinical outcomes, aligning with the standard pituitary surgery core outcome set, were collected. RESULTS: Eight patients with pituitary adenomas were enrolled. The CVAI system was successfully deployed in six cases, demonstrating acceptable real-time pixel-level sella segmentation accuracy. Deployment failed pre-operatively in two cases owing to a single platform-level boot-configuration issue. Iterative refinement between cases was driven by our experience and surgical team feedback. This resulted in the integration of additional anatomical structure segmentations (e.g., carotid arteries), enhanced model accuracy via training dataset expansion, and hardware firmware upgrades. Multi-stakeholder surveys demonstrated satisfactory system feasibility, usability, and acceptability among the surgical team. Both prospective observation and retrospective video review confirmed the absence of adverse events, including no significant distraction to the primary surgeon, and there were no AI-related clinical complications. CONCLUSION: This first-in-human early clinical evaluation (IDEAL Stage 1) of real-time pixel-level AI anatomical segmentation in live neurosurgery demonstrates feasibility, showcases iterative system evolution, and reports clinical and human factors outcomes. Future work will include a larger single-centre case series (IDEAL Stage 2a) with more surgical teams to further iterate the system and explore its impact on safety, training and workflow. As the underpinning AI models improve and integrate with other intra-operative navigational technologies, such tools will likely be the cornerstone of intra-operative surgical decision support systems.
JMIR human factorsLouise Nørgaard Olsen, Philipp Harbig, Anna Bay Laurberg, Jacob Laurberg, Morten Haaning Charles
BACKGROUND: Administrative workload in general practice limits time for direct patient care. AI-assisted documentation has been proposed as a way to reduce the documentation burden, but evidence from routine primary care settings remains limited. OBJECTIVE: This study aimed to evaluate general practitioners' (GPs) acceptance of AI-assisted documentation and its association with documentation time and clinical note quality in routine Danish general practice. METHODS: We conducted a quantitative pragmatic pre-post quality improvement evaluation in Danish general practice. A total of 20 GPs documented 239 consultations before and 236 consultations after implementation of an AI-assisted documentation system. Documentation quality, structure, clinical clarity, and documentation time categories were self-assessed using standardized audit forms completed immediately after each consultation. Technology acceptance and usability were assessed using the technology acceptance model (TAM) and the System Usability Scale (SUS). RESULTS: Self-assessed documentation structure increased from 3.99 to 4.45, while self-reported documentation time categories decreased from 2.85 to 2.29. Technology acceptance and usability were high (TAM domain means 3.76-4.19; SUS mean 77.5). GP-level paired analyses showed moderate improvements in structure and clarity and a reduction in documentation time. Combined blinded external assessments showed higher postimplementation scores for quality, structure, and clinical clarity, although reviewer-specific ratings diverged, and interrater reliability was low. The association between documentation time and perceived quality was negligible. TAM and SUS indicated high clinician acceptance. CONCLUSIONS: AI-assisted documentation was associated with lower self-reported documentation time categories while maintaining or modestly improving perceived clinical note quality. These findings support the feasibility of AI-assisted documentation in primary care, while highlighting the need for controlled studies with objective time measurement and longer follow-up.
Journal of global healthJiayu Xu, Zhihan Zhang, Guanran Zhang, Yanlin Qu, Zhenyu Wu, Xiaodong Sun, Da He, Huixun Jia
BACKGROUND: Retinopathy of prematurity (ROP) is one of the leading causes of childhood blindness globally, but particularly in middle-income countries, where neonatal survival has improved faster than access to specialist ophthalmic care. Although telemedicine and artificial intelligence (AI)-assisted screening show promise for improving access to ROP screening, their economic value within decentralised health systems remains uncertain. To explore this issue, we evaluated the cost-effectiveness of community-based AI-assisted ROP screening in China compared with telemedicine and traditional bedside screening. METHODS: We developed a decision tree model to estimate lifetime health outcomes and societal costs for a hypothetical cohort of 100,000 preterm infants eligible for ROP screening. We compared three strategies: community-based AI-assisted screening, community-based telemedicine screening, and bedside screening at specialist healthcare facilities. Separate parameter sets were applied for urban and rural settings. Costs were assessed from a societal perspective. RESULTS: Community-based AI-assisted screening was the most cost-effective strategy in both urban and rural settings. In the base-case analysis, telemedicine screening dominated bedside screening (meaning it was less costly and more effective), and AI-assisted screening further dominated telemedicine screening. AI-assisted screening incurred the lowest lifetime costs (USD 1,507 per infant in urban and USD 2,581 in rural areas) and the greatest health benefits (29.7138 quality-adjusted life-years (QALYs) in urban and 29.5835 QALYs in rural areas). Compared with bedside screening, AI-assisted screening was dominant, with incremental cost-effectiveness ratios of USD -7,307 per QALY gained in urban and USD -3,688 per QALY gained in rural settings. Findings were robust across scenario and sensitivity analyses, with AI-assisted screening ceasing to be cost-effective only at extremely low treatment-requiring ROP incidence (<0.7%) or low screening coverage (11%) in rural areas. CONCLUSIONS: Community-based AI-assisted ROP screening is likely the most cost-effective strategy in China. By expanding access to early eye care while reducing societal costs, it may support equitable service delivery and offer a scalable model for other regions facing similar resource constraints.
International journal of dermatologyTim J Zeuner, Titus J Brinker
A growing number of artificial intelligence (AI) systems for melanoma diagnosis now match dermatologist-level performance. Early detection is a driving factor, yet it is rarely validated beyond aggregate diagnostic accuracy. More targeted studies are necessary so these claims become verifiable and the promise of early detection can actually be delivered to patient care.
Adaptive neuromodulation offers a path toward personalized neuropsychiatric treatment by allowing therapy to respond to clinically meaningful neurophysiological and behavioral changes. Its clearest clinical precedents are responsive neurostimulation for epilepsy and adaptive deep brain stimulation for movement disorders. Both rely on signals that can be measured reliably and linked directly to stimulation. Psychiatric disorders pose a different problem: symptoms are heterogeneous, evolve over multiple timescales, and are observed only indirectly through noisy and incomplete measurements. Artificial intelligence can provide data-driven components of an adaptive neuromodulation system, connecting multimodal sensing, inference of patient state, prediction of treatment response, and constrained therapeutic adaptation. Interpretability is required throughout to support clinician and patient trust. Clinical evidence, however, remains limited to small proof-of-concept studies. Translation to practice will require replicated biomarkers and prospective multisite validation. It will also require explicit management of uncertainty and system designs that match adaptation to symptom and treatment timescales. Clinician and patient oversight must be preserved throughout, alongside safety constraints and governance across the treatment life cycle. This review presents a system-design view of how these components combine into a single adaptive architecture for psychiatric neuromodulation.
The clinical teacherW Baraza, Georgina Staples, Cameron Wells, Louise Barbier, Isi Tonga, Parry Singh, Mike Puttick, Craig Webster
INTRODUCTION: Self-directed learning (SDL) and reflection are critically important for clinicians. There is limited evidence describing the utilization of artificial intelligence (AI) to support medical SDL. AI simulated patients (AISP) may allow for high-fidelity, reflective educational interactions. We aimed to determine how SDL with AISP interaction compared to standard SDL in the development of clinical reasoning skills. METHODS: This was a single-institution randomized controlled study in early clinical years' undergraduate students, learning about gallstone disease. For 90% power and a 20% noninferiority margin, 36 participants were randomized either to an intervention group using an AISP or to a control group using current educational resources. Both groups underwent a standardized clinical skills assessment. Quantitative and qualitative data were collected. RESULTS: Twenty-seven participants completed the study. The intervention group demonstrated statistical noninferiority compared with the control group (assessment scores of 13.1 vs. 12.1, respectively, p = 0.11). In addition, there were fewer borderline grades, and significantly more distinction grades in the intervention group than control (6 vs. 1, respectively, p = 0.037), suggesting the AISP may have improved performance in clinical management. The qualitative data suggested widespread satisfaction and enthusiasm for AISPs. CONCLUSION: Our results suggest that AISPs are comparable to traditional resources in developing undergraduate clinical skills, and may contribute to improved academic performance in undergraduate surgery. Globally, there is an appetite for their introduction into the clinical curriculum. Real patient interaction is the standard. However, AISPs may have a supplementary role and are likely to become increasingly prominent in medical curricula.
Health promotion journal of Australia : official journal of Australian Association of Health Promotion ProfessionalsSarah Turner, Sara Dingle, Melissa Ensink, Emily Tomlinson, Tristan Duncan, Carah Figueroa, Shane Kavanagh
ISSUE ADDRESSED: Artificial intelligence (AI) is transforming public health and health promotion (PHHP) practice, creating a need for graduates who can use AI ethically, critically and effectively. Although AI is increasingly being incorporated into higher education, the focus has been on individual units or assessment tasks, with limited guidance available for embedding AI capabilities across entire courses. This study responded to this gap by developing a curriculum framework for integrating AI capability across postgraduate PHHP programmes. METHODS: An iterative co-design approach was undertaken within a large Australian university to develop an AI curriculum framework for the Master of Public Health and Master of Health Promotion programmes. The framework was informed by competency-based education and integrated curriculum design principles. It was also aligned with local institutional AI principles and mapped to international accrediting competencies. RESULTS: The resulting 'ECHA' framework comprises four principles: Ethical and Responsible Use of AI; Critical Appraisal; Human-Centred Decision-Making and Professional Judgement; and Applied AI Literacy and Safe Practice in PHHP. Each principle is accompanied by discipline-specific learning outcomes, curriculum content and core graduate skills. CONCLUSIONS: The new framework offers a practical, programme-wide model for integrating AI capability into postgraduate PHHP education. Embedding AI within existing professional competency frameworks supports coherent curriculum design while preserving the human-centred values underpinning professional practice. SO WHAT?: This paper provides a coordinated, competency-informed approach to course-wide AI curriculum design, intended to assist universities in preparing PHHP graduates to use AI ethically, critically and responsibly.
Acta anaesthesiologica ScandinavicaConstance Thyra Giertz Christiansen, Mathilde Nellemann, Moritz Kilian German Denneborg, Sofie Amalie Bosholdt, Janus Christian Jakobsen, Rasmus Tingkær Hessel…
BACKGROUND: Pre-anaesthetic assessment is essential for safe perioperative care but is resource-intensive and applied variably in clinical practice. Increasing surgical demand and workforce constraints have prompted interest in alternative models, including digital approaches and AI-supported clinical decision support tools. However, limited knowledge exists on clinicians' perspectives, values, and perceived information requirements in relation to pre-anaesthetic assessment, including which information they consider essential prior to anaesthesia. METHODS: This international, cross-sectional survey aims to explore clinicians' perspectives on pre-anaesthetic assessment practices and their attitudes towards AI-supported clinical decision support tools. The survey will be conducted in two phases, beginning with Nordic centres before broader international expansion, using two linked instruments: an organisational survey completed once by a local site investigator at each site, and an individual survey completed by anaesthesia personnel recruited via site-based convenience sampling. The primary outcome is clinicians' perspectives on pre-anaesthetic assessment, comprising perceived importance, information not readily available from chart review considered essential prior to anaesthesia, and use in subsequent anaesthetic management; secondary outcomes include airway assessment practices and attitudes towards AI-supported tools, including willingness to adopt them. Exploratory outcomes include organisational workflows and time-use in pre-anaesthetic assessment. Responses will be summarised descriptively across clinician and organisational subgroups. CONCLUSION: This survey will provide an international overview of clinicians' perspectives on pre-anaesthetic assessment and its role in clinical decision-making. The findings may contribute to the development of more efficient, stratified, and clinically aligned assessment models while recognising that implementation of digital and AI-supported approaches will also depend on organisational, technical, ethical, legal, and patient-related factors.
Artificial intelligence systems are increasingly drawn upon to inform nursing practice, education, and health policy, often on the assumption that they produce stable, broadly consistent knowledge. This paper challenges that assumption by examining how seven widely accessible large language models, developed in distinct socio-technical and geopolitical contexts, interpret a shared global nursing challenge. Through comparative qualitative analysis of responses to a standardized zero-shot prompt on the global nursing shortage, four competing policy logics were identified, namely workforce, efficiency, equity, and mobility, with patterns that appeared consistent with aspects of the institutional environments associated with those systems' development. Each response was internally coherent and presented with substantial authority, yet none acknowledged the situatedness of its own framing. The paper argues that such outputs in nursing should be understood as situated artifacts rather than as neutral knowledge, and offers a framework of questions to guide discretion for nurses, educators, and policymakers engaging with these tools. Recognizing this plurality is necessary for safe, contextually grounded decision-making and for governance that addresses both interpretive variability and accuracy.
Physics in medicine and biologyBohua Wan, Todd McNutt, Harry Quon, Junghoon Lee
Objective. Radiation therapy (RT) is critical in head and neck cancer (HNC) treatment but often causes radiation-induced toxicities, such as xerostomia (RIX). While deep learning (DL)-based models show promise in predicting these toxicities, their black-box nature hinders clinical applications. This study aims to develop a robust DL model to predict RIX 12 months after RT and to leverage model interpretability techniques, specifically class activation maps (CAMs) and voxel-wise dose gradient maps (GMs), to guide personalized treatment plan optimization.Approach. A 3D ResNet-based model was trained using planning CTs, dose volumes, and salivary gland contours obtained from a retrospective cohort of 839 HNC patients. To address anatomical variations, we normalized each patient's volumes to a common reference frame using atlas normalization. To improve spatial correspondence between deep features and patient anatomy, we integrated blur pooling and adaptive average pooling, mitigating downsampling-induced voxel shifts. Model interpretability was achieved using Grad-CAM++ and GMs. Personalized plan optimization was performed on predicted RIX-positive cases by generating avoidance contours from both CAMs and GMs to guide dose reduction.Main results. The proposed model was tested on 30 independent test cases. It achieved an AUC of 0.77 with balanced sensitivity (0.71) and specificity (0.83). Among 9 predicted RIX-positive cases, GM-guided optimization converted the predictions to RIX-negative in 7 cases (77.8%), compared with 6 cases (66.7%) for CAM-guided optimization, and achieved a greater reduction in mean model-predicted RIX probability (26% versus 20%).Significance. The proposed atlas-normalized model achieved robust discrimination, and a 9-case plan-refinement analysis showed that CAM- and GM-derived spatial information could be translated into clinically constrained avoidance objectives that reduced the model-predicted RIX probability while preserving specified target and organs at risk dose requirements. These results demonstrate the feasibility of the proposed strategy for guiding clinical interventions.
Rhode Island medical journal (2013)Mohamad Y Fares, Tarishi Parmar, Peter Boufadel, Mohammad Daher, Jonathan Berg, Adam Z Khan, Brian W Hill, John G Horneff, Joseph A Abboud
OBJECTIVES: Chatbots have been increasingly recognized as modern tools that provide patients with reliable health-related information. This study aimed to evaluate and compare ChatGPT-3.5 and GPT-4's ability to answer glenohumeral osteoarthritis-related questions. METHODS: Fifteen questions were derived from the 2020 AAOS Clinical Practice Guidelines for the Surgical Management of Glenohumeral Joint Osteoarthritis. Questions were categorized into three groups: risk factors, implant/intraoperative considerations, and pain/functional outcomes. ChatGPT-3.5 and GPT-4 were prompted with these questions, and responses were evaluated by four fellowship-trained shoulder and elbow surgeons. Each response was rated on a scale (scores:1-5) based on relevance, accuracy, clarity, completeness, and evidence-based support. Data was analyzed descriptively and statistically to compare the scores between ChatGPT-3.5 and GPT-4. RESULTS: Average score for ChatGPT-3.5 was 19.7/25, with "Risk Factor" prompts achieving the highest mean score. GPT-4 averaged 18.7/25, with "Functional Outcomes" prompts scoring highest. However, there were no statistically significant differences between different prompt themes for GPT-3.5 and GPT-4. "Clarity" category received the highest score for GPT-3.5, while "Relevance" was highest for GPT-4. Both models scored lowest on "Evidence-based" prompts. On the Flesch- Kincaid scale, GPT-3.5 responses had a significantly higher score of 18.3 compared to GPT-4's 15.4, indicating a more difficult reading level in GPT-3.5's responses. CONCLUSION: Both ChatGPT-3.5 and GPT-4 performed adequately in providing well-informed medical responses to patient queries about glenohumeral osteoarthritis. Future chatbot versions should focus on providing evidence-based content through systematic and reliable reviews of literature, in an accessible readable manner.
Pharmacoepidemiology and drug safetyByeong Yeob Choi
BACKGROUND: In observational pharmacoepidemiology, estimating average treatment effects (ATEs) is often challenging due to a lack of practical positivity. In highly selective clinical settings, certain patients almost always or never receive treatment, causing ATE estimators to rely on unstable extrapolation. Incremental propensity score interventions (IPSIs) offer a stochastic alternative by shifting each patient's probability of treatment, providing a more clinically realistic framework that circumvents positivity violations. METHODS: We illustrate the IPSI approach, including key identification results and inferential procedures. Using observational data from a cohort of 996 patients undergoing percutaneous coronary intervention (PCI), we evaluated the effect of shifting each patient's probability of receiving abciximab by a predetermined amount on six-month mortality. Propensity scores (PSs) and outcome predictions were estimated using a machine learning ensemble (Super Learner) with 10-fold sample splitting. RESULTS: The ATE estimate suggested that abciximab administration reduced the 6-month mortality risk by 5.9 percentage points compared with PCI alone (risk difference = -0.059, 95% CI: -0.104 to -0.015). However, the practical interpretability of the ATE estimate may be limited because it implicitly assumes that patients with a near-certain probability of treatment could realistically be assigned to withhold abciximab. In contrast, shifting each patient's treatment propensity by odds ratios ranging from 0.1 to 10 showed that 6-month mortality would be significantly reduced under a strong treatment policy promoting abciximab administration. CONCLUSIONS: IPSIs provide a robust and practical alternative to conventional causal inference methods in pharmacoepidemiology settings where treatment assignment is highly selective and the strict positivity is violated.
Pharmacoepidemiology and drug safetyJames M Gwinnutt, Rodrigo de Oliveira, Miriam J Haviland, Julien H Shippee, Elizabeth Eldridge, Lenon Mendes Pereira, Emily Bratton, Jay Nanavati, Christina De…
Large language models (LLMs) represent a type of generative artificial intelligence (GenAI) that generate and interpret text, with some LLMs able to process multimodal content (e.g., images, audio, video), and can be deployed as part of agents to perform users' tasks. LLMs can perform natural language processing functions such as summarization, translation, and extraction giving them the potential to enhance and scale pharmacoepidemiological and real-world research by performing tasks such as literature review, data extraction, and medical writing. Despite the growing integration of GenAI tools into research workflows, their technical foundations and methodological implications remain unfamiliar to many pharmacoepidemiologists, who are often responsible for the reliability and accuracy of research that relies on these tools. This paper aims to inform pharmacoepidemiologists about the capabilities and limitations of LLMs to support responsible integration into the field of pharmacoepidemiology, providing an intuitive overview of how LLMs work, focusing on training and text generation, and reviews current and emerging applications in drug effectiveness and safety research and epidemiology. The article addresses challenges associated with LLM use in real-world evidence generation, including concerns regarding reproducibility, bias, hallucinations, plagiarism, data privacy, and the need for validation.
Liver international : official journal of the International Association for the Study of the LiverHaoshuang Fu, Yanan Du, Gangde Zhao, Yuelin Xiao, Shuying Song, Xinya Zang, Tianhui Zhou, Yan Zhuang, Ruidong Mo, Rongtao Lai, Qing Xie
BACKGROUND & AIMS: Drug-induced liver injury (DILI) would progress to chronicity or death. Large language models (LLMs) may enhance clinical decision-making, yet their utility relative to physicians in DILI remains unclear. Therefore, we evaluated their performance in predicting DILI outcomes. METHODS: We enrolled 943 DILI patients from three centers as internal and external cohorts. Based on 12-month follow-up, outcomes were classified as recovery, 6/12-month chronicity, and death. LLMs (Gemini-2.5 Pro, GPT-5.1, DeepSeek-3.2), hepatologists (Junior, middle, senior), and models (Hy's Law, nHy's Law, MELD Score) estimated probabilities of outcomes. LLM-Rules (VOTE, OR, AND) were applied to enhance stability. Model performance was assessed. RESULTS: For 6-month chronicity, the senior achieved highest AUROC (0.61) with an accuracy of 70%. Gemini-2.5 Pro and GPT-5.1 yielded AUROCs of 0.60 and 0.59, respectively, outperforming junior and middle hepatologists. Gemini-2.5 Pro demonstrated strongest agreement with senior (κ = 0.43). LLMs all exhibited lower accuracy and specificity than hepatologists. A similar result was observed in 12-month chronicity. For overall mortality, the senior achieved highest AUROC (0.87) with an accuracy of 83%. Gemini-2.5 Pro and GPT-5.1 achieved AUROCs of 0.86, outperforming junior hepatologist, Hy's Law, and nHy's Law. GPT-5.1 achieved strongest agreement with senior (κ = 0.25). LLM-Rules demonstrated stability for predicting outcomes across cohorts. OR and AND rules improved sensitivity and specificity, respectively. CONCLUSIONS: GPT-5.1 and Gemini-2.5 Pro showed AUROCs approaching senior hepatologists for DILI outcomes with limited accuracy and specificity. LLM-Rules demonstrated stable performance across cohorts with improved sensitivity or specificity, supporting the potential of multi-LLM approaches as clinician-supervised complementary tools.
American journal of public healthKathryn Heley, Linnea I Laestadius, Jeffrey K Hom, Andy J King
The rapid emergence of photorealistic generative artificial intelligence (AI) has introduced synthetic visual content into health information ecosystems at an unprecedented scale, lowering barriers to creation while eroding the traditional reliability of visual evidence. Synthesizing perspectives from public health, communication, and psychology, we argue that photorealistic AI-generated images and videos (AIVs) have transformative and destabilizing potential with clear public health implications, creating a visual information environment where long-standing assumptions about visual evidence no longer hold. We reviewed relevant research and identified risks posed by AIVs, including mis- and disinformation, fraudulent medical representations, amplified stigma, and a possible collapse of visual trust. We applied the host‒agent‒vector‒environment (HAVE) model as a strategic framework to map determinants of AIV-related harms and identify multilevel intervention points. Strategic frameworks like HAVE are most effective when paired with a normative, deliberative layer that addresses the ethical dimensions of possible responses. We therefore propose a public health ethics deliberation guide designed to work alongside HAVE, encouraging responses to photorealistic AIVs that are strategically targeted, proportionate, equitable, and ethically grounded. (Am J Public Health. 2026;116(10):1539-1549. https://doi.org/10.2105/AJPH.2026.308675).
Pediatric annalsJennifer L Fang, Mark W Kaczor, Whitney S Thompson, Christopher A Collura
Telemedicine, rapid genomic sequencing (rGS), artificial intelligence (AI), and emerging interventions at the limits of viability are reshaping neonatal critical care and transforming the future of neonatology. Telemedicine is expanding neonatal expertise beyond tertiary centers, improving access and continuity of care through prenatal consultation, tele-resuscitation, virtual rounds, and home monitoring. Advances in rGS are shifting neonatal diagnostics toward earlier, precision-based identification of genetic disease, enabling more individualized treatment and prognostication. AI applications in the neonatal intensive care unit are demonstrating promise in early disease detection, imaging interpretation, predictive analytics, and clinical decision support, with the potential to improve outcomes while reducing clinical burdens. Neonatology continues to consider ethical challenges at the limits of viability. The potential for artificial womb technology raises complex questions regarding fetal status, viability thresholds, and the future implications of ectogenesis prompting the urgent need for evolving ethical frameworks.
Physiotherapy research international : the journal for researchers and clinicians in physical therapyEzgi Güçlü, Nuray Girgin
BACKGROUND AND PURPOSE: Robot-assisted gait training (RAGT) is well established for post-stroke gait rehabilitation, but its potential effects on psychological and behavioral outcomes are less clear. This study investigated the effects of adding RAGT to conventional rehabilitation on balance, gait, kinesiophobia, and movement confidence in individuals with post-stroke hemiparesis. METHODS: This single-blind, parallel-group randomized controlled trial included 60 individuals with post-stroke hemiparesis (50-75 years), randomly allocated to an RAGT group (n = 30) or control group (n = 30). Ethical approval was obtained from the Clinical Research Ethics Committee of Istanbul Yeni Yüzyıl University (Approval No. 20.01.2022/05; approval date: 20 January 2022). Both groups received conventional rehabilitation for 8 weeks; the RAGT group additionally received 24 sessions of RAGT. Kinesiophobia was a prespecified study outcome assessed using the Kinesiophobia Causes Scale (KCS); balance, gait, and balance confidence were also assessed. All 60 randomized participants completed follow-up and were analyzed in their assigned groups. RESULTS: Significant group × time interactions were observed for several outcomes, including BBS, TUG duration, 10MWT walking speed, ABC, and KCS total score (p < 0.05). The between-group difference in change for KCS total score was -0.36 (95% CI: -0.51 to -0.21; partial eta squared = 0.292). In post hoc analyses adjusting each outcome for its baseline value, significant group effects remained for BBS, TUG duration, 10MWT walking speed, ABC, KCS biological domain, and KCS total score (p< = 0.031), whereas 10MWT step count and the KCS psychological domain were no longer statistically significant. DISCUSSION: Adding RAGT to conventional rehabilitation was associated with greater improvements in several balance, mobility, walking-speed, balance-confidence, and kinesiophobia outcomes compared with conventional rehabilitation alone. These findings suggest potential additional physical and psychological benefits of incorporating RAGT into post-stroke rehabilitation. However, because the RAGT group received greater overall treatment exposure, the observed between-group differences cannot be attributed solely to the robotic component. The principal contribution of this study is the concurrent evaluation of kinesiophobia and movement confidence alongside physical outcomes.
Evidence-based clinical decision support artificial intelligence (AI) is rapidly expanding, but its safe and effective use depends on rigorous validation, trustworthy evidence sources and careful integration into clinical workflows. Current available systems show strong potential to improve diagnostic accuracy, reduce clinician workload and possibly benefit patient care, but challenges remain before its real-world adoption. We must be responsible in its integration to ensure AI truly strengthens clinical judgment. This article is a viewpoint of AI tools for clinical decision support, addressing a rapidly evolving field and provides insights that are useful for clinicians and educators in clinical settings. When used within appropriate medical training, AI may help augment diagnostic accuracy and improve efficiency. Nevertheless, while AI offers promising features, it also presents ethical and reliability challenges, which may negatively affect the medical professional identity.
American journal of public healthImran Hossain Mithu, Mona Arora, Onicio B Leal Neto, Paloma I Beamer
We examined environmental and public health implications of energy-intensive artificial intelligence (AI) infrastructure, emphasizing health equity, workers, and vulnerable communities. We analyzed evidence on AI infrastructure across its life cycle, including data-center electricity demand, fossil-fuel power, water use, heat, noise, critical-mineral extraction, occupational exposures, and electronic waste (e-waste). AI infrastructure can expose nearby communities and workers to air pollution, heat, noise, water stress, hazardous mining conditions, and e-waste toxicants. In 2023, US data centers consumed 4.4% of national electricity, much from fossil fuel combustion, contributing to climate change and air pollution. Data centers also consume large water volumes and generate e-waste, while demand for high-performance hardware intensifies upstream environmental and occupational hazards. These risks fall disproportionately on low-income, Indigenous, and other marginalized communities near data centers, mines, power plants, and e-waste recycling sites. Responsible AI governance should treat energy-intensive AI as a public health and health equity issue requiring health impact assessments, exposure surveillance, equitable siting, accountability, community participation, and fair distribution of AI's health-related benefits and burdens. (Am J Public Health. 2026;116(10):1529-1538. https://doi.org/10.2105/AJPH.2026.308683).
Hernia : the journal of hernias and abdominal wall surgeryBruno Amantini Messias, Guilherme Costa E Silva, Pedro Henrique de Freitas Amaral, Diogo Parente Falcão, Sergio Roll, Jaques Waisberg
PURPOSE: Despite technical advances, the expansion of minimally invasive approaches and the development of novel biomaterials, recurrence and complications following abdominal wall hernia (AWH) repair remain frequent. This variability likely reflects anatomical and technical factors and host biological heterogeneity. Precision herniology is emerging as a translational paradigm integrating biomarkers, genetics, quantitative imaging, and artificial intelligence to refine risk stratification and therapeutic planning. METHODS: We conducted a critical narrative review with translational scope. PubMed/MEDLINE, Scopus, Web of Science, and Embase were searched from January 2015 to June 2026. We prioritized adult studies across four axes: (i) biomarkers of collagen, extracellular matrix, and cellular mechanisms; (ii) genetics and susceptibility to AWH; (iii) immunonutritional markers; (iv) artificial intelligence (AI) and machine learning (ML) for predicting complexity and complications. RESULTS: The literature reveals convergent, albeit heterogeneous, signals indicating that host biology may contribute to variability in hernia phenotype and surgical outcomes. Alterations in collagen turnover and in the MMP-TIMP axis, fibroblast heterogeneity, inflammatory signaling, polygenic architecture related to connective tissue integrity, immunonutritional markers associated with perioperative vulnerability, and quantitative imaging metrics analyzed by AI/ML models represent biologically plausible, clinically relevant domains. CONCLUSION: Available data support the plausibility that host biological heterogeneity contributes to the formation, healing, prosthetic integration, and recurrence of AWH, thereby broadening the traditional paradigm centered predominantly on defect anatomy. The main value of this work is the critical integration of molecular and cellular biomarkers, genetics, quantitative imaging, and AI/ML into a unified conceptual framework for future precision herniology research.
INTRODUCTION: Evidence-based dosing guidance for medications in critically ill patients with acute kidney injury (AKI) and receiving continuous kidney replacement therapy (CKRT) is limited. Freely available large language models (LLMs) can generate confident, human-like outputs. The accuracy and reproducibility of LLMs in providing precision drug dosing recommendations in the setting of AKI and CKRT have not been evaluated. OBJECTIVES: We sought to characterize literature concordance and internal consistency of LLM-generated dosing recommendations for cefepime and meropenem in patients with AKI and receiving CKRT. METHODS: Six investigators queried the freely available versions of six LLMs (ChatGPT, Claude, Google Gemini, Microsoft Copilot, OpenEvidence, and Perplexity) from July to September 2025 using three standardized vignettes asking for dosing recommendations and rationale to meet prespecified pharmacodynamic targets: (i) adult with AKI not on dialysis receiving cefepime, (ii) child receiving CKRT and cefepime, and (iii) toddler receiving high-effluent CKRT and meropenem. To assess inter-iteration consistency, each investigator also prompted one LLM three times for each vignette. Responses were parsed for concordance with recommendations in primary literature. LLM "reasoning" was evaluated for use of pharmacokinetic (PK) equations, citation accuracy versus confabulation, acknowledgment of uncertainty, and recommendations for therapeutic drug monitoring (TDM) for efficacy or safety. RESULTS: LLM-generated dosing recommendations varied widely. Mean concordance with literature-based recommendations was 63% (range: 28%-94%). LLMs produced variable responses to the same user upon multiple iterations, with variance in daily maintenance doses recommended ranging from 0 to 1000%. OpenEvidence universally cited relevant sources, whereas other LLMs leveraged sources inconsistently or confabulated them. Each recommended TDM, though only Claude consistently acknowledged its own uncertainty. CONCLUSION: Freely available LLMs produce highly variable and often discordant antibiotic dosing recommendations for patients with AKI and receiving CKRT. Although valuable for hypothesis generation and literature retrieval, LLM outputs should not be used in isolation for drug dosing in critically ill patients with kidney dysfunction.
JMIR medical educationLaura Brunelli, Federico Fonda, Silvio Brusaferro
The continuous development of AI presents unprecedented opportunities for public health. This development prompts educators and researchers to consider how AI can be applied to and integrated into a rapidly evolving global landscape in which AI is reshaping health systems. In turn, this generates an urgent need for structured proposals that bridge theory and practice. In this viewpoint, we present perspectives and proposals on AI applications in public health. We operationalize AI literacy as a core public health competency, defined through knowledge, skills, and attitudes and scaled across professional roles. For education, we describe how AI can support learners, teachers, and organizations through personalized learning experiences, dynamic scenarios, immersive simulations, content preparation, and communication strategies. For research, we analyze how AI can be embedded as a transversal enabler supporting the World Health Organization's essential public health functions. We emphasize that AI should enhance, rather than replace, human capabilities and propose that AI literacy should be recognized as a core public health competency. We suggest that harnessing AI for public health education and research is, therefore, less a technological challenge than a collective responsibility shared by educators, researchers, and institutions.
Clinical journal of oncology nursingRichard Nishan Boyajian, Molly Hagen, Sugato Bagchi, Shilpa Mahatma, Ashleigh Kowtoniuk, Quoc-Dien Trinh, Tarun Kumar, Paul L Nguyen
BACKGROUND: The U.S. oncology workforce is not increasing at the same pace as the cancer survivor population. Cancer care requires extended coordination across multiple disciplines. Oncology teams need technological solutions to support patient surveillance and meet growing care demands. OBJECTIVES: The purpose of this study was to develop and test an artificial intelligence (AI) approach to identify and track care states in patients with prostate cancer. METHODS: The study examined data from a retrospective cohort of 180 patients to develop ground truth care state assignments with parameters set by clinicians, then developed an AI model that set its own parameters for care state assignment and tested it on this cohort dataset. The study assessed the agreement between assignments made by the AI model and clinician specialists. FINDINGS: The AI model produced care state assignments for all 180 patients in less than one second, and agreement with the clinician-assigned care state was 95%.
JMIR cancerAlexandra H Smick, Martins Ayoola, Rajeshree Rajpara, Isabel Lazo, Dana M Chase
BACKGROUND: Patients with newly diagnosed gynecologic cancers often seek information online, but the quality of available resources may be inconsistent. Although GPT-4 may offer an alternative to traditional internet search engines, its performance remains largely under-studied in gynecologic oncology. OBJECTIVE: This study aimed to compare the completeness, accuracy, and reference quality of responses generated by GPT-4 with those generated by Google in clinical scenarios involving a new diagnosis of a gynecologic cancer. METHODS: Clinical scenarios representing early- and advanced-stage endometrial, ovarian, and cervical cancers were developed by gynecologic oncologists using publicly available patient education materials. Each scenario included 4 standardized questions addressing etiology, prognosis, treatment, and treatment efficacy. GPT-4 and Google were queried for each question, with new sessions for GPT-4 and private browsing for Google to minimize bias. Responses were independently evaluated by 4 gynecologic oncology experts who were blinded to each other's ratings. Accuracy was scored on a 6-point Likert scale; completeness and reference quality were scored on 3-point scales. Reference quality was categorized as low (commercial), medium (institutional or government), or high (peer reviewed). Optional free-text reviewer comments were collected and summarized descriptively to provide context for the quantitative findings. Descriptive statistics and Wilcoxon signed-rank tests were used for analysis. RESULTS: Across 6 clinical scenarios and 21 standardized questions (N=84 total responses), GPT-4 outperformed Google across all evaluated domains. The median accuracy score was 6.00 (IQR 5.00-6.00) for GPT-4 and 5.00 (IQR 4.00-6.00) for Google (P=.04). The median completeness score was 3.00 (IQR 3.00-3.00) for GPT-4 and 2.00 (IQR 1.00-3.00) for Google (P=.009). Reference quality was also higher for GPT-4, with a median score of 3.00 (IQR 3.00-3.00) compared to 2.00 (IQR 2.00-2.00) for Google (P=.009). Reviewer comments noted that GPT-4 provided more accurate, comprehensive, and personalized responses, while Google returned less detailed content from general consumer health websites rather than peer-reviewed sources. CONCLUSIONS: GPT-4 may serve as a reliable and high-quality resource for information in gynecologic oncology. Compared to Google, it delivered more accurate and complete content, with higher-quality references. Further research is needed to assess the readability and accessibility of GPT-4-generated content across diverse patient populations.
Journal of medical Internet researchZiyan Wu, Honglin Xu, Futai Feng, Shulan Zhang, Rongrong Cheng, Tianqi Shi, Siyu Wang, Yongzhe Li
BACKGROUND: Large language models (LLMs) are increasingly used as health information intermediaries. Whether they provide comparable accuracy and communication quality across languages has direct implications for health information equity; however, systematic bilingual evaluations remain limited. OBJECTIVE: This study aimed to provide a preliminary bilingual benchmark evaluating whether 11 LLMs deliver comparable accuracy and communication quality when answering identical consumer health questions in English and Chinese. METHODS: We conducted a controlled evaluation of 11 LLMs (GPT-4.5, Claude Sonnet 4, Gemini 2.5 Flash, Grok 3, DeepSeek R1, Qwen 3, Doubao, Kimi k1.5, Hunyuan T1, ERNIE X1 Turbo, and ChatGLM 4) using 150 binary consumer health questions from the Text Retrieval Conference Health Misinformation Track (2019, 2021, and 2022). All models were accessed through official public-facing web interfaces during May 2025. Models were assessed under 2 full-benchmark prompting conditions (no-context and expert), evaluating accuracy, comprehensiveness, precision, and understandability. Four post hoc error-correction strategies (chain-of-thought [CoT], retrieval-augmented generation [RAG], CoT+RAG, and error attribution) were applied to baseline-incorrect responses. Composite ranking used the technique for order of preference by similarity to ideal solution (TOPSIS), with sensitivity analysis across 3 weighting schemes. Generalized estimating equations and linear mixed models with Benjamini-Hochberg false discovery rate (FDR) correction were applied using a full 3-way interaction specification (model×language×prompt). RESULTS: English and Chinese inputs showed comparable overall accuracy under no-context conditions (1572/1650, 95.27% vs 1548/1650, 93.82%), with no significant language main effect (β=0.00; P=.99). No language main effects for any individual model remained significant after FDR correction. TOPSIS analysis identified ChatGPT and Qwen as the most consistently top-ranked models (tier 1 in 12/12 condition×weight-scheme combinations). A model-specific language interaction emerged for communication quality: DeepSeek showed a significant English-language decrement in understandability (β=-0.73; FDR=-0.016), while its decrements in precision and comprehensiveness were not significant after correction. One 3-way interaction survived: Grok showed a disproportionate accuracy reduction when English input and expert prompting were combined (β=-1.88; FDR=-0.022). Among post hoc correction strategies, error attribution achieved the highest correction rate (Δ55.56%), although this condition provided models with privileged information. CONCLUSIONS: Contemporary LLMs achieved high binary accuracy on consumer health questions in both English and Chinese, with no significant aggregate language effect. The only robust model-specific language interaction was DeepSeek's English understandability decrement, independently confirmed by TOPSIS tier analysis. These findings suggested that cross-linguistic communication quality concerns were model-specific rather than universal and warrant targeted monitoring.
Canadian journal of surgery. Journal canadien de chirurgieJoëlle Deschênes-Bilodeau, Peter Staunton, Charles Desgagné, John Antoniou
BACKGROUND: Patients have been resorting to online content as their main source of medical knowledge, notably on orthopedic interventions. However, Web-based research findings present a reverse correlation between quality of content and popularity. We sought to evaluate whether ChatGPT could provide an alternative and safe source of medical information for arthroplasty patients. METHODS: We gave 5 commonly Googled questions related to hip and knee arthroplasty to board-certified arthroplasty surgeons, fellows, and orthopedic surgery residents, as well as to ChatGPT 4.0. We anonymized all answers, which were then analyzed by an independent board-certified arthroplasty surgeon. We scored the answers for accuracy of content (6-point Likert scale) and completeness (3-point Likert scale), then compared the performance of all groups. RESULTS: Among human responders, the mean accuracy grade was 75% (standard deviation [SD] 17%), with a mean completeness grade of 69% (SD 22%). The fellows represented the strongest subgroup, with 75% of their answers scored above 5/6 for accuracy (mean 82%, SD 16%). We found that all human-generated answers had a statistically significant correlation between accuracy and completeness. ChatGPT had a mean accuracy grade of 93% (SD 12%), with a mean completeness of 93% (SD 14%). CONCLUSION: ChatGPT appears to be a safe tool for patients to access general arthroplasty information online. It outperformed all human responders on both accuracy and completeness of answers. The strongest human responder group was the arthroplasty fellows. Further work is required to clarify the tool's performance against other easily accessible online patient information sources.
International journal of medical informaticsMonika Shekhawat, Ashok Kumar Nagar, Arpita A Gupta, Alma Rachel Koshy, V Sai Sankalp Naidu, Pura Krishnamurthy Kiran, Neeraja Sathyakkagari
INTRODUCTION: Cancer is a major public health challenge in India, with an estimated 1.4 million new cases each year. Despite this burden, the ratio of oncologists to new cancer patients is critically low, about one specialist per 1600 new diagnoses per year. The AI-driven CDSS could help offset the oncologist shortage in LMICs. OneRx.AI is an RAG platform grounded in international and national oncology guidelines. We evaluated its Expert-rated accuracy, clinical safety, Guideline concordance and decision impact, confidence, efficiency, and inter-rater reliability (IRR) in Indian practice. METHODOLOGY: A prospective, multicentre, expert-validation study conducted at three tertiary oncology centres in India. Three Oncologist-clinicians independently evaluated de-identified queries across five cancer types. Individuals and common-prompt evaluation phases were analysed separately. Co-primary endpoints were assessed using one-sample z-tests with 95% Wilson confidence intervals, and inter-rater reliability was measured using the pre-specified Gwet's AC1 statistic. RESULTS: Of 124 submitted clinical queries, 110 (88.7%) met the eligibility criteria. Expert-rated fully correct clinical accuracy was 82.7%, expert-rated clinical safety was 100%, and fully concordant guideline recommendations were 94.5%. Composite rates for Expert-rated clinical accuracy and guideline concordance were 100%. Expert-rated Clinical accuracy showed Gwet's AC1 = 0.707. Expert-rated Safety and guideline concordance received uniform ratings across evaluators. Citation accuracy was 86.4%, decision-influence rate 92.7%, physician-confidence increased by a mean of 1.69 (p < 0.0001), and median time to clinical recommendation decreased from 389.2 s to 47.2 s, representing an 87.9% reduction (p < 0.0001). CONCLUSION: OneRx.AI demonstrated high expert-rated clinical performance in this exploratory, ceiling-affected evaluation, supporting progression to a larger Stage II study.
Translational psychiatryChristian Rauschenberg, Frederike Schirmbeck, Janik Fechtelpeter, Eva Wierzba, Selina Hiller, Katharina Kahr, Anna Kessler, Lale Hornbacher, Christian Goetzl, …
Ecological Momentary Interventions (EMIs) using machine learning (ML)-based assignment algorithms may improve mental health outcomes by delivering more person-tailored content, but evidence is pending. The study aimed to determine whether ML-based assignment of EMI components augments effects on momentary mental health outcomes when compared to random assignment in youth from the general population and psychological counselling services. A within-subject micro-randomized trial was conducted. Participants were randomly assigned up to seven times daily (1:1 ratio; up to 210 decision points) to either an ML-based (experimental condition) or a random (active control condition) assignment of EMI components. Proximal outcomes were time-lagged changes in positive affect, momentary resilience, and negative affect at tn+1. Feasibility and safety were assessed. Distal outcomes included psychological distress, resilience, and emotion regulation. A total of 49 youths (mean age 20.6; 78% female) were included. At baseline, participants reported mild-to-moderate psychological distress on average (K10 mean = 23.8, SD = 7.1), with more than one third reporting moderate or severe distress. An initial, outcome-specific signal favoring ML-based over random assignment was observed for momentary resilience (B = 0.147, 95% confidence interval (CI), 0.004 - 0.290, p = 0.044), whereas there was no evidence of beneficial effects on positive or negative affect. Feasibility indicators supported delivery of the AI4U training, with favorable ratings of satisfaction, acceptability, and usability; no serious adverse events were reported. Uncontrolled pre-post comparisons suggested a small reduction in psychological distress (d = -0.23) and improvements in resilience (d = 0.55) and adaptive emotion regulation (d = 0.53). Taken together, this study demonstrates the feasibility and preliminary safety of a ML-based adaptive EMI in youth and provides an initial, outcome-specific signal that ML-based assignment may improve momentary resilience relative to random assignment, whilst underscoring the need for larger, adequately powered MRTs and formal validation of whether forecasting performance translates into policy value.
World journal of urologyTheo Clark, Samuel Murphy, Solomon Bracey, Ali Talyshinskii, Arman Tsaturyan, Steffi Kar Kei Yuen, Carlotta Nedbal, Selcuk Guven, Vineet Gauhar, Frederic Panth…
BACKGROUND: Patients with urological conditions often undergo recurrent computed tomography (CT) imaging, which results in cumulative radiation exposure which can be deleterious. Artificial intelligence (AI)- based technologies, including Deep Learning Image Reconstruction (DLIR), have emerged as a potential strategy to reduce radiation dose in CT imaging while maintaining image quality. This systematic review aimed to quantify the benefits of AI-based techniques in urological CT imaging. METHODS: A systematic search of Ovid MEDLINE, Embase, and Scopus was conducted. Studies assessing AI-based techniques in urological CT were included. Primary outcomes of interest were radiation dose metrics (CTDIvol, DLP, effective dose), with other outcomes of interest including image quality metrics along with diagnostic performance of AI techniques. RESULTS: Eleven studies met the inclusion criteria. All studies demonstrated a radiation dose reduction aided with AI techniques, with most reporting reductions of 60-80%. Consistently, Image quality was maintained or improved, with reduced image noise and increased signal-to-noise ratio. Limited evidence from diagnostic studies showed at-least comparable performance between AI-based reconstruction and conventional iterative reconstruction, but with unclear superiority at equivalent dose levels. CONCLUSION: AI-based techniques show a clear ability to allow radiation dose reduction in urological CT imaging, whilst maintaining or improving image quality. However, the current evidence base is limited by a lack of diagnostic outcomes, and further research into AI techniques is required in order to quantify their clinical effectiveness.
Oral and maxillofacial surgeryDaniel Stephan, Sophia Schumacher, Bilal Al-Nawas, Peer W Kämmerer, Daniel G E Thiem
PURPOSE: Medical documentation is essential for clinical communication but is often intended for professional audiences, limiting patient understanding. This linguistic complexity can reduce health literacy and hinder shared decision-making. Large language models offer new opportunities to simplify medical texts while maintaining factual accuracy, thereby improving accessibility and comprehension. This study aimed to evaluate whether ChatGPT can simplify medical reports while preserving clinical content and to assess whether these simplifications improve patient comprehension, readability, and perceived communication quality. METHODS: Five document types were analysed, including MRI, CT, surgical, pathology reports, and discharge summaries. Each original physician-written document was simplified using a standardized ChatGPT prompt instructing full content preservation and patient-oriented phrasing. Ten simplifications per text were reviewed for completeness. Readability was measured using Flesch Reading Ease (FRE) and LIX indices. A total of 576 participants without medical knowledge evaluated either a simplified or original report version and completed a standardized questionnaire assessing clarity, structure, and applicability, followed by comprehension testing. RESULTS: Simplified reports achieved significantly higher readability (FRE 48.4 ± 5.0 vs. 22.9 ± 5.2; p < 0.0001; LIX 48.3 ± 3.2 vs. 59.2 ± 2.5; p = 0.004). Across all document types, patients rated simplified texts significantly higher in clarity, structure, and usefulness (p < 0.001), with significantly improved comprehension accuracy (e.g., MRI 80.3% vs. 53.7%; p < 0.001). No loss of medical information was observed. CONCLUSIONS: ChatGPT appears to be capable of simplifying complex medical documents while preserving clinically relevant information, leading to improvements in both perceived readability and objective understanding under the conditions of this study. AI-driven text simplification thus represents a promising tool to enhance patient communication and health literacy.
Einstein (Sao Paulo, Brazil)Agnaldo Rodrigues da Costa, Carlos Henrique Sartorato Pedrotti, Tarso Augusto Duenhas Accorsi
In the aftermath of the COVID-19 pandemic, telemedicine has evolved from a contingency-based alternative into a structural pillar of healthcare delivery, increasingly being integrated with artificial intelligence. Algorithms for triage, clinical decision support, remote monitoring, and generative language models now actively participate in medical care. This integration is more than just functional; it reorganizes the foundational aspects of medicine-ontological, ethical, and epistemological-by redefining what qualifies as valid knowledge, who has legitimate authority, and how care is organized when an algorithmic agent facilitates the clinical interaction. This article critically examines the emergence of the patient-clinician-algorithm triad in telemedicine, and analyzes its implications for autonomy, informed consent, epistemic trust, privacy, justice, responsibility, and the phenomenology of care. It argues that the future of medicine will not be human or algorithmic but human-algorithmic, provided clinical judgment, explainability, and moral responsibility remain central.
Neurological sciences : official journal of the Italian Neurological Society and of the Italian Society of Clinical NeurophysiologyYunting Wu, Ran Zhang, Fei Wu, Minghui Liu, Yuying Wang, Weige Sun, Weixin Cai
OBJECTIVE: To compare and rank the efficacy of brain-computer interface-coupled robotic rehabilitation (BCI-robot), robot-assisted therapy, motor imagery (MI), and conventional rehabilitation for upper-limb recovery after stroke, and to explore whether treatment effects differed according to key clinical and intervention-related characteristics. METHODS: A systematic search was conducted in 11 databases from inception to March 31, 2026. Randomized controlled trials enrolling adults with post-stroke upper-limb motor impairment were included. The primary outcome was the Fugl-Meyer Assessment of Upper Extremity (FMA-UE). Secondary outcomes included the Action Research Arm Test (ARAT), Wolf Motor Function Test (WMFT), and Modified Barthel Index (MBI). Direct pairwise meta-analyses were performed using mean differences (MDs) with 95% confidence intervals (CIs), and a frequentist network meta-analysis was conducted to synthesize direct and indirect evidence and rank interventions using P-scores. RESULTS: Twenty-three reports representing 22 independent study cohorts were included, comprising 727 participants. Network meta-analysis showed no significant inconsistency, and the consistency model was adopted. BCI-robot ranked highest for improving FMA-UE according to the P-score analysis (P-score = 0.992), followed by robot-assisted therapy (0.641), MI (0.240), and conventional rehabilitation (0.127). In direct comparisons, BCI-robot produced a statistically significant improvement in FMA-UE versus conventional rehabilitation (MD = 6.80, 95% CI 4.29 to 9.30). The point estimate exceeded the 5.25-point (minimal clinically important difference) MCID, although the lower confidence limit did not. Compared with robot-assisted therapy, BCI-robot showed a statistically significant but small improvement (MD = 2.04, 95% CI 0.29 to 3.78), with both the point estimate and the entire 95% CI remaining below the MCID. Follow-up-duration subgroup analyses suggested that BCI-robot remained favorable over conventional rehabilitation in both < 3-month and ≥ 3-month subgroups, whereas no significant follow-up advantage over robot-assisted therapy was observed. For secondary outcomes, BCI-robot significantly improved WMFT versus conventional rehabilitation (MD = 8.94, 95% CI 5.22 to 12.65) and MBI versus conventional rehabilitation (MD = 3.37, 95% CI 0.10 to 6.63). A significant benefit for MBI versus robot was also observed (MD = 8.61, 95% CI 1.76 to 15.46), although this result was not robust in sensitivity analysis. No significant advantage was found for ARAT. Exploratory subgroup analyses suggested that, compared with robot-assisted therapy, the additional benefit of BCI-robot was more apparent in studies delivering more than four sessions per week. CONCLUSIONS: BCI-robot ranked highest for improving FMA-UE and may provide a clinically important benefit over conventional rehabilitation, although the magnitude remains uncertain. Its additional benefit over robot-assisted therapy was small and unlikely to be clinically important. Evidence for longer-term effects and for activity-level and daily-function outcomes remains limited, and the findings should be interpreted cautiously.
JMIR research protocolsPınar Kaya, Esra Tekeci, Elif Hocaoğlu, Ramazan Ünal, Gokhan Ozkocak
BACKGROUND: Stroke is a leading cause of long-term disability worldwide. Persistent lower-extremity motor and somatosensory impairments after stroke commonly limit walking and balance despite rehabilitation. Virtual reality (VR)-integrated robotic rehabilitation may support structured, goal-directed ankle-foot practice; however, evidence for ankle-foot-focused sensorimotor protocols remains limited. In particular, approaches that combine robot-assisted motor training with plantar tactile localization and VR-supported joint position sense training to target plantar sensory and proprioceptive function are scarce. OBJECTIVE: This study aims to evaluate the effectiveness of a structured, VR-integrated, robot-assisted ankle-foot sensorimotor rehabilitation protocol compared with a content-matched manual training protocol in individuals with chronic stroke and to examine its effects on clinical and sensorimotor outcomes. METHODS: This is an assessor-blinded, 2-arm, parallel-group randomized controlled trial. Thirty individuals with chronic stroke will be randomized 1:1 to the robot-assisted training group or the manual training group. All participants will receive conventional rehabilitation. In addition, the robot-assisted training group will receive a structured robot-assisted ankle-foot training program integrated with VR and assist-as-needed control, whereas the manual training group will receive the same structured ankle-foot training protocol delivered manually by a physiotherapist. Interventions will be delivered 3 times per week for 6 weeks, for a total of 18 sessions. Total session duration will be time matched between groups at 50 to 60 minutes per session. The primary outcome will be the change in 10-meter walk test-derived walking speed from baseline to 6 weeks. The secondary outcomes will include 2-minute walk test distance, ankle range of motion, joint position sense, plantar tactile sensation, muscle tone, motor performance, static and dynamic balance, and stroke-specific quality of life. RESULTS: This study is funded by the Scientific and Technological Research Council of Turkey (TÜBİTAK) under the 1002-A Rapid Support Program (grant 225S390). Recruitment began in March 2026 and is planned to continue until May 2027. As of March 25, 2026, four participants have been enrolled and are currently receiving the assigned intervention. Data analysis will begin once all enrolled participants have completed post-intervention assessments. CONCLUSIONS: This trial will provide evidence on whether a structured robot-assisted ankle-foot sensorimotor training program offers additional benefit compared with a content-matched manual training protocol in individuals with chronic stroke. The findings may contribute to the development of standardized, individualized, and sensorimotor-oriented rehabilitation protocols for improving walking and balance after stroke.
Journal of medical Internet researchSeongwoo Yang, Ju Hyun Jin, Seng Chan You, Min Jung Kim, Kyung Won Kim
BACKGROUND: Pediatric health care requires distinct considerations, including caregiver involvement and developmental differences in cognition and communication as children gain autonomy, particularly as pediatric health care chatbots gradually emerge. Because childhood and adolescence are formative periods for health behaviors and self-management practices, pediatric chatbots also warrant evaluation against long-term rather than immediate outcomes. OBJECTIVE: This study aimed to characterize and synthesize the available evidence on pediatric health care chatbots evaluated for health-related outcomes. Furthermore, by identifying gaps in the existing literature, we sought to propose specific considerations for the design, evaluation, and implementation of pediatric health care chatbots. METHODS: PubMed, Embase, Scopus, PsycINFO, the Cochrane Library, and the Web of Science were systematically searched without publication year restrictions. Randomized controlled trials, mixed methods, and observational studies that evaluated health care chatbots for children (aged <19 y) or caregivers and assessed health-related outcomes were included. Nonoriginal papers, end-of-life or palliative care studies, and non-English publications were excluded. Study quality was assessed using the Mixed Methods Appraisal Tool and the Oxford Levels of Evidence 2. RESULTS: A total of 9 studies were included, with 5 (55.6%) involving pediatric participants only, while 4 (44.4%) involved caregivers. Six (66.7%) studies lacked a comparator, and only 3 (33.3%) chatbots were AI-based. Health and psychosocial outcomes were mixed, often showing null findings in objective clinical metrics despite some subjective improvements. Behavioral and cognitive outcomes generally showed favorable changes but relied heavily on subjective evaluations. Although chatbots demonstrated explicit developmental tailoring, with designs shifting from caregiver-mediated approaches in early childhood to autonomous, privacy-focused platforms for adolescents, definitive conclusions regarding their robust associations with health-related outcomes cannot be drawn. This is primarily due to pervasive methodological limitations, including the lack of active comparator groups, reliance on short-term metrics, and significant study heterogeneity. CONCLUSIONS: Pediatric health care chatbots are emerging across diverse health care contexts, but the current evidence remains limited and heterogeneous. This review identified developmentally relevant considerations, including caregiver involvement, age-appropriate communication, and developmental differences, that may warrant explicit attention in future chatbot design, evaluation, and implementation.
Flatfoot (pes planus) is common in children and adults. Although often asymptomatic, it may alter lower-limb biomechanics and contribute to pain or injury. Patients and caregivers increasingly use artificial intelligence chatbots such as chat generative pretrained transformer (ChatGPT) and Google Gemini for medical information, yet their responses on flatfoot remain insufficiently studied. This study compared the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions. In this cross-sectional comparative study, 15 standardized patient-oriented questions were developed from clinical guidelines, peer-reviewed literature, and educational resources from established orthopedic societies, then refined by 2 orthopedic surgeons. Each question was submitted to ChatGPT (OpenAI) and Google Gemini (Google LLC) in July 2025 using identical wording in independent chat sessions. Responses were anonymized and independently rated by 2 board-certified orthopedic surgeons for factual accuracy, added value, and omissions using an investigator-developed 5-point rubric. Discrepant assessments were reviewed by a 3rd senior orthopedic surgeon. Primary analyses used mean scores of the 2 reviewers. Readability was assessed using the Flesch-Kincaid grade level and Flesch reading ease score. Paired differences were analyzed using the Wilcoxon signed-rank test, with P < .05 considered significant. Both models produced generally accurate responses, with mean factual accuracy scores above 4/5. Gemini scored higher than ChatGPT for factual accuracy (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01), added value (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02), and omissions (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), with higher scores indicating fewer clinically relevant omissions. Gemini also generated longer responses (165 ± 20 vs 120 ± 15 words, P < .001) and showed better readability, with lower Flesch-Kincaid grade level (9.2 ± 1.1 vs 11.8 ± 1.4, P < .001) and higher Flesch reading ease score (59.4 ± 6.2 vs 42.5 ± 5.3, P < .001). Both models generated generally accurate answers to standardized flatfoot questions. Gemini performed better across evaluation domains and readability measures. As these findings are time-specific, repeated assessments and direct patient and caregiver evaluations are needed before broader patient-facing use can be recommended.
Digital therapeutics (DTx) are emerging as evidence-based software interventions, but current AI-driven personalization approaches lack dedicated safety-focused frameworks and face challenges due to scarce long-term outcome data and unpredictable model behaviors. We propose SAFE_DTx, a safety-first architectural framework for DTx that integrates well-established principles of predictive modeling and constrained decision-making to prioritize patient safety. SAFE_DTx's 2-module architecture comprises an AI feedback prediction module that forecasts short-term patient responses and a constrained planning module that selects the next intervention under explicit safety constraints. By decoupling these components and enforcing clear safety guardrails, the framework enables dynamic, real-time adaptation to individual patient feedback while staying within evidence-based safety limits. This modular design also enhances transparency in the decision-making process, and an in silico evaluation demonstrates its preliminary architectural feasibility, showing greater engagement and no safety violations compared to baseline strategies within the simulated environment. SAFE_DTx's safety-by-design architecture aligns with emerging regulatory emphasis on AI transparency and patient safety. It directly addresses key clinical challenges in AI-driven DTx personalization by ensuring that tailored interventions do not compromise patient safety.