انفورماتیک پزشکی
مقالهها، منابع و پژوهشهای تازه حوزه انفورماتیک پزشکی
ورود به زیرشاخهداده، اطلاعات، کتابداری و شواهد سلامت
برای رسیدن به فهرست متمرکزتر، یک مسیر تخصصی را انتخاب کنید.
مقالهها، منابع و پژوهشهای تازه حوزه انفورماتیک پزشکی
ورود به زیرشاخهمقالهها، منابع و پژوهشهای تازه حوزه بیوانفورماتیک
ورود به زیرشاخهمقالهها، منابع و پژوهشهای تازه حوزه مدیریت اطلاعات سلامت
ورود به زیرشاخهمقالهها، منابع و پژوهشهای تازه حوزه پزشکی مبتنی بر شواهد
ورود به زیرشاخهمقالهها، منابع و پژوهشهای تازه حوزه مرور نظاممند و متاآنالیز
ورود به زیرشاخهمقالهها، منابع و پژوهشهای تازه حوزه کتابداری و اطلاعرسانی پزشکی
ورود به زیرشاخهThe intestinal epithelium is maintained by stem cells at the crypt base, however, acute or chronic inflammation can severely disrupt stem cell function and tissue homeostasis. To characterize mucosal inflammation and tissue damage, we evaluated time-course gene expression changes using a dextran sulfate sodium (DSS)-induced mouse model of colitis. By applying normalization, Z-score transformation, and spline curve modeling, we trace the temporal expression dynamics of key intestinal genes. This computational approach offer insights into the molecular responses during inflammation and may help identifying new biomarkers for chronic inflammatory diseases.
Human induced pluripotent stem cell (iPSC)-derived microglia (iMG) provide an in vitro experimental system for studying human microglial biology, neuroinflammation, and genetic risk mechanisms associated with neurological disease. This chapter describes a standardized, scalable, and reproducible protocol for the differentiation of human iPSCs into functional microglia-like cells, with particular emphasis on applications in transcriptional and epigenomic network analysis. The protocol supports high-viability floating iMG production, compatibility with pooled CRISPR perturbation approaches, and downstream multiomic profiling, including single-cell RNA sequencing, chromatin accessibility assays, and proteomics. Detailed procedures are provided for iPSC maintenance, hematopoietic progenitor cell generation, microglial maturation, functional genomics integration, and quality control.
Single-cell RNA sequencing (scRNA-seq) has facilitated the studies of cellular heterogeneity in many different biological contexts, including how B cells function and mediate immune responses to infection, malignancies, and autoimmunity. scRNA-seq as applied to B cells requires specific bioinformatics strategies, which consider the distinct biology of B cells and biological processes unique to these cells, such as somatic hypermutation (SHM) and class switch recombination (CSR). We present here a protocol that analyses scRNA-seq alignments and count matrices, as well as B-cell receptor sequencing data obtained in parallel to scRNA-seq. We highlight computational strategies to ensure cell clusters align with biologically relevant B-cell subsets and cell states. We provide practical guidance on how CSR and SHM information can be extracted from the data and utilized to build robust computational models of B-cell maturation.
Cardiac organoids are increasingly used to model human cardiac development and disease, but their small size often limits molecular characterization, especially if different approaches and protocols are necessary to extract and quantify metabolites and lipids. Here, we present a multistep workflow for combined targeted free amino-acid (FAA) based metabolomic and lipidomic profiling from a single pooled cardiac organoid sample. The protocol covers organoid harvesting, detergent-assisted lysis, and a modified Folch extraction that generates an organic phase for lipid analysis and an aqueous phase for FAA metabolite analysis. Lipids are quantified by Liquid Chromatography Electrospray Ionization Tandem Mass Spectrometry (LC-ESI-MS/MS) in Multiple Reaction Monitoring (MRM) mode using class-matched external standards and an internal standard to support calibration and reduce technical variability. Free amino acids and derivatives are analyzed from the same sample after filtration, drying, and AccQ-Tag derivatization, with norvaline as an internal standard. Together, this approach maximizes information yield from limited material and enables integrated analysis of metabolic and lipid pathways within the same biological specimen, facilitating organoid-based studies of cardiac maturation, disease modeling, and pharmacological responses.
Cancer cell identity is governed by coordinated transcriptional programs that are frequently rewired during tumorigenesis. Systematic identification of cancer type-specific gene regulatory networks provides a framework for understanding oncogenic state transitions and for prioritizing candidate therapeutic targets. Here, we present a reproducible network-based workflow for reconstructing and analyzing transcriptional regulatory programs across human cancer types using publicly available expression datasets. We describe procedures for curating and preprocessing microarray data from the Gene Expression Omnibus, implementing random forest classification, and reconstructing gene regulatory networks using the CellNet platform. Detailed guidance is provided for evaluating classifier performance, quantifying network influence scores, integrating transcription factor, target interaction resources, and performing functional enrichment analyses. In addition, we outline approaches for comparing cancer-specific networks with corresponding normal tissue profiles to identify candidate drivers of malignant cell identity and potential prognostic biomarkers. Together, these protocols provide investigators with a scalable computational framework for defining cancer type-specific transcriptional states and for systematically interrogating regulatory mechanisms underlying tumor heterogeneity.
Clustered regularly interspaced short palindromic repeats (CRISPR) and associated (Cas) systems have revolutionized the field of genome engineering by providing versatile, efficient, and programmable tools for precise genetic manipulation. Originally identified as an adaptive immune mechanism in prokaryotes, CRISPR/Cas systems have been extensively repurposed for a wide range of applications across molecular biology, biotechnology, and medicine. This chapter provides a comprehensive overview of the molecular mechanisms underlying CRISPR/Cas immunity. Furthermore, the classification of CRISPR/Cas systems into distinct types and subtypes is discussed, highlighting their structural and functional diversity. Advances in genome editing technologies, including CRISPR-mediated knockout, base editing, and prime editing, are explored with an emphasis on their mechanisms and applications. The chapter also examines emerging CRISPR-based platforms for transcriptional regulation, epigenome editing, and RNA targeting, which enable precise and reversible modulation of gene expression without altering genomic DNA. In addition, the transformative impact of CRISPR technologies on functional genomics is addressed, particularly through high-throughput screening approaches that facilitate the identification of gene function and genetic vulnerabilities. CRISPR-based diagnostic tools and therapeutic strategies are also reviewed, underscoring their potential in disease detection and treatment. Despite significant progress, challenges such as off-target effects, delivery limitations, and safety concerns remain critical considerations. Overall, this chapter highlights the expanding capabilities of CRISPR/Cas systems and their growing importance in both fundamental research and clinical applications.
Conventional treatments often face challenges such as the limited ability to penetrate the blood-brain barrier (BBB). The Doxorubicin-loaded graphene oxide/magnetite (DOX/GO/Fe3O4) nanocomplex offers a promising platform due to GO's high surface area and pH-sensitive release, and Fe3O4's magnetic properties. This protocol describes the methodology for evaluating the cytotoxicity of free DOX versus the DOX/GO/Fe3O4 nanocomplex in the A-172 glioblastoma cell line, followed by advanced bioinformatics analysis to identify gene networks and indirect pathways that enhance the nanomaterial's biocompatibility. The methodology integrates the MTT assay, real-time PCR for apoptosis genes (Casp3, Bax, and Bcl-2), and advanced analysis, including protein-protein interaction (PPI) networking, clustering, and promoter motif analysis. The analysis indicated that miR-92a-2-5p is a potential therapeutic target for preventing myocardial damage and enhancing biocompatibility. The findings highlight key regulatory pathways that indirectly boost nanodrug biocompatibility through the modulation of secondary components like miRNAs and cellular stress mechanisms.
Understanding drug mechanisms of action (MOA) and predicting drug-target interactions (DTIs) are fundamental challenges in modern drug discovery and development, hindered by high costs, long development timelines, and limited knowledge of compound activity and molecular targets. Here, we present two deep learning-based computational protocols designed to address these challenges. The first framework employs directed message passing neural networks (D-MPNN) to predict drug MOA from chemical-genetic interaction profiles (CGIPs), by learning how molecular structures perturb biological pathways through systematic profiling across genetically sensitized strains. The second framework, iNGNN-DTI, utilizes interpretable nested graph neural networks combined with pretrained molecule models to predict DTIs, leveraging cross-attention mechanisms to provide insights into binding determinants. We highlight the application of these methods to key therapeutic areas, including antibacterial drug discovery and drug repurposing for COVID-19 therapeutics. Each protocol provides comprehensive guidance on data preparation, model implementation, validation strategies, and result analysis. These computational approaches offer scalable, cost-effective tools for accelerating therapeutic development by bridging chemical structure, molecular interactions, and systems-level biological responses.
The plasma membrane (PM) is the primary interface between plant cells and their environment, and its resident proteins mediate key processes such as extracellular signal perception and downstream cellular reprogramming. Yet, PM proteins are typically underrepresented in total protein extracts, and existing enrichment strategies are often laborious and require extensive optimization. Here, a simple and robust workflow is described for enriching PM proteins from Arabidopsis thaliana seedlings using total microsomal membranes obtained by differential centrifugation as starting material. Sequential low- and high-speed spins are used to isolate total microsomal membranes and progressively deplete contaminating organelles, thereby increasing the relative abundance of PM proteins. Coupled with the rich genetic toolkit available in Arabidopsis, this protocol provides an accessible platform for systematic characterization of the PM proteome.
Affinity purification-mass spectrometry (AP-MS) is a powerful proteomic approach for dissecting the interaction network between virus and host. Traditional AP-MS employs overexpression of viral proteins as baits to enrich host interactors. However, overexpressed viral proteins may mislocalize to inappropriate cellular compartments and trigger endoplasmic reticulum stress by overwhelming the protein-folding machinery, which leads to false identification of host factors. To overcome these limitations, we introduce an AP-MS strategy based on direct infection with an epitope-tagged chikungunya virus (CHIKV/myc-E2), which we used to successfully uncover two new antiviral factors in CHIKV cellular reservoirs-macrophages. In this protocol, we will describe this technique step by step: (1) design and construction of myc-tagged virus by advanced multi-fragment assembly, (2) in vitro transcription and preparation of infectious myc-tagged virus stocks, and (3) immunoprecipitation of myc-tagged viral protein and its interactome for mass spectrometry analysis. This strategy enables accurate identification of viral interactors in a physiologically relevant context, providing a framework for future proteomic studies using tagged viruses.
The availability of a high-quality genome assembly facilitates the analysis of fungal genomes. This chapter outlines the tools and steps involved in genome sequence assembly and annotation of a plant pathogen, Botrytis cinerea. We describe the use of Illumina short-read and Oxford Nanopore long-read sequencing data to assemble the B. cinerea genome. The steps include the pre-processing of sequencing data, genome assembly using Flye, scaffolding with NtLink, and polishing with Racon, Medaka, and NextPolish. The quality of the final assembly is evaluated using BUSCO, which serves as a benchmark for the completeness of a genome. We also provide details on the identification and masking of repetitive elements using the EarlGrey pipeline, as well as the gene prediction and annotation process with Funannotate. The methodologies and insights described can be applied to genome research in other fungal species.
Alterations in chromatin state, mediated through histone modifications and the incorporation of histone variants, are fundamental to establishing transcriptional networks and cell identity. Recent advances in low-input epigenome profiling methods, such as CUT&Tag and CUT&RUN, have enabled the study of chromatin states from very limited starting materials. In this chapter, we describe procedures for generating CUT&Tag libraries to profile histone modifications and histone variants in early-developing zebrafish embryos.
Global genomic surveillance has emerged as a foundational pillar of public health in the twenty-first century, enabling real-time tracking of pathogen evolution and informing outbreak response. This chapter examines the strategic architecture of global genomic surveillance, focusing on its application to arboviruses such as chikungunya virus (CHIKV). It explores the integration of genomic data with epidemiological, clinical, and environmental information within a One Health framework, while addressing critical challenges in governance, equity, and interoperability. The discussion covers the entire genomic surveillance workflow, from sample collection and sequencing to bioinformatic analysis and phylogenetic inference, and highlights the transformative role of artificial intelligence (AI) in predictive surveillance. By analyzing global initiatives, operational barriers, and emerging technologies, this chapter underscores the necessity of sustainable, equitable, and interoperable genomic systems to proactively address current and future infectious disease threats.
Recent advances in multiomics technologies have revolutionized the study of immune-driven diseases by enabling high-dimensional, single-cell resolution analyses. This chapter provides a practical guide for integrating CyTOF (mass cytometry) and single-cell RNA sequencing (scRNA-seq) data to address key challenges in this field. The integration of these modalities allows for consistent and reproducible cell type annotation, the transfer of annotations to assist in characterizing difficult-to-identify populations, and the transcriptional characterization of rare and heterogeneous subpopulations. Using tools such as OMIQ and R, the chapter outlines workflows for preprocessing, normalization, and scaling of CyTOF data, as well as dimensionality reduction and clustering techniques. The integration process involves creating Seurat objects, identifying common features, and using anchor-based methods to link CyTOF and scRNA-seq datasets. The chapter also discusses the use of multimodal deep learning techniques for rare subpopulation detection and emphasizes the importance of reproducibility and standardization in multiomics integration. By leveraging these methodologies, researchers can gain deeper insights into cellular heterogeneity and function, ultimately enhancing the understanding of immune-driven diseases. The chapter concludes by addressing integration challenges and proposing future directions for improving model interpretability and capturing nonlinear molecular interactions.
Mass spectrometry-based proteomics allows the unbiased identification and quantification of proteins and phosphopeptides in biological materials. The nature of walled plant cells requires specific protocols for effective and efficient protein isolation, and, in general, the plant sciences can benefit from more accessible, optimized proteomics workflows. Advances in MS instrumentation now allow the measurement of large numbers of samples, shifting constraints in proteomics toward the accurate, high-throughput preparation of samples. Here, we describe a high-throughput (phospho)proteomics protocol that enables processing of samples using different filter types in a 96-well format.
Aberrant three-dimensional genome organization is a hallmark of cancer, often driving oncogene activation through mechanisms such as enhancer hijacking. High-throughput chromosome conformation capture (Hi-C) maps these interactions on a genome-wide scale. Unlike earlier dilution-based methods, in situ Hi-C performs proximity ligation within intact nuclei, minimizing random ligation noise and enabling fine-scale structure detection. This chapter describes an optimized in situ Hi-C protocol tailored for cancer cell lines using MboI digestion and biotin-mediated pull-down to generate high-complexity libraries. We further outline a computational workflow that extends beyond standard topological mapping of compartments and topologically associating domains to identify cancer-specific aberrations. Specifically, we focus on detecting chromosomal rearrangements (structural variants) and characterizing the distinct circular topology of extrachromosomal DNA. This integrated experimental and analytical framework provides the necessary tools to dissect the spatial dysregulation underlying tumor evolution.
High-throughput sequencing of total RNA has permitted the detection of novel mycoviruses in fungi with different types of genomes, including mycoviruses with double-stranded RNA, single-stranded positive- or negative-stranded RNA, or single-stranded DNA genomes. However, in silico detection of mycoviruses is not always sufficient to guarantee their presence in sequenced samples, especially in the case of the discovery of unique mycoviruses, and additional analyses are required to validate in vivo the data obtained by bioinformatics analysis. This chapter provides comprehensive protocols for the extraction of total RNA from the plant pathogenic fungus Botrytis cinerea for next-generation sequencing (NGS), outlines the bioinformatics pipeline designed and followed to detect mycoviruses in the sequenced samples, and details the detection in vivo and the complete molecular characterization of the mycoviruses identified in silico.
Single-cell transcriptomics has revolutionized our understanding of cellular heterogeneity by enabling high-resolution gene expression profiling at the individual cell level. However, traditional single-cell RNA sequencing (scRNA-seq) lacks direct protein quantification, limiting comprehensive immunophenotyping. Cellular Indexing of Transcriptomes and Epitopes by sequencing (CITE-seq) overcomes this limitation by integrating antibody-derived tag (ADT) quantification with scRNA-seq, allowing simultaneous measurement of surface protein and gene expression from the same cell. This multimodal approach enhances immune cell characterization, revealing new functional states and rare subpopulations in complex biological systems. Here, we provide a detailed protocol for performing CITE-seq, from sample preparation to sequencing and data analysis. We highlight key experimental considerations, discuss challenges related to antibody selection and batch effects, and provide troubleshooting strategies to ensure robust and reproducible results. The integration of transcriptomic and proteomic data through CITE-seq provides unparalleled insights into cellular function, with broad applications in immunology, oncology, and systems biology.
Studying the transcriptome and the proteome of cells is essential for gaining a detailed understanding of cellular behavior, development, drug action, and disease progression. Spatial biology emerges to advance our ability to study the expression of molecules within tissues while preserving their natural spatial context. These cutting-edge technologies enable the mapping of thousands of individual cells in their original environment by detecting the location and biological quantity of cellular contents, such as RNAs and proteins. Here, we present a protocol to combine spatial transcriptomics (Xenium) and spatial proteomics (PhenoCycler-Fusion) within 8 days on the same tissue section to successfully study the expression of hundreds of RNAs and tens of proteins simultaneously. The combination of these two technologies and consequent integration of the two data layers together with high-resolution H&E images allows for the extraction of a maximum of information from a single tissue section. Application of this protocol and the resulting integrated data will help researchers to understand complex biological processes and disease mechanisms, supporting more nuanced research in molecular biology and pathology.
Single-cell RNA sequencing (scRNA-seq) has revolutionized the ability to resolve cellular heterogeneity within complex tissues, enabling the identification of discrete cell states. Here, we present an in silico analytical pipeline designed to characterize intestinal stem cells (ISC), transit-amplifying (TA) progenitors, and BEST4⁺ enterocyte precursors from human scRNA-seq datasets, with a focus on inflammatory contexts such as inflammatory bowel disease (IBD). The pipeline integrates dataset acquisition, quality control, normalization, dimensionality reduction, unsupervised clustering, and cell type annotation using a reference cell atlas. We implemented iterative subsetting and re-clustering of ISC and TA compartments to identify inflammation-associated subpopulations and epithelial biomarkers. While demonstrated in the context of IBD, this computational framework is broadly applicable to other tissues and pathological conditions where stem/progenitor dynamics underpin disease progression and tissue repair.
Gene regulatory networks (GRNs) represent the complex interplay of transcription factors, regulatory elements, and target genes that orchestrate cellular identity and function, playing a crucial role in the differentiation and maintenance of stem cells. This chapter provides an overview of experimental and computational methodologies for inferring GRNs, with particular emphasis on single-cell approaches. We first review key experimental techniques for detecting transcription factor binding sites, chromatin accessibility, and DNA motifs, alongside essential databases that support GRN reconstruction. We then introduce computational inference methods that can be categorized into four principal frameworks: correlation-based approaches, regression and machine learning models, probabilistic and deep learning methods, and integrative or message-passing frameworks. To illustrate practical application, we present a case study applying the pySCENIC workflow to a peripheral blood mononuclear cell single-cell RNA sequencing dataset from mouse, demonstrating how regulon-based analysis can reveal cell-type-specific regulatory programs. This chapter aims to serve as a practical guide for researchers seeking to understand and implement GRN inference methodologies in stem cell biology and related fields.
.: Nonalcoholic fatty liver disease (NAFLD) is a significant global health concern, impacting roughly 25% of people and leading to chronic liver conditions. It involves excess fat accumulation in the liver without significant alcohol intake and can develop into nonalcoholic steatohepatitis (NASH), fibrosis, or cirrhosis. While liver biopsy remains the gold standard for diagnosis, its invasive nature and associated risks restrict its routine use. Noninvasive biomarkers, such as serum ALT, AST, and various composite scores, are available; however, their clinical usefulness is often limited by variable sensitivity and specificity across different populations and disease stages. To overcome these limitations, this chapter offers a comprehensive, reproducible protocol for identifying and clinically validating circulating long noncoding RNA (lncRNA) biomarkers for NAFLD and NASH. The workflow integrates bioinformatic analysis of four Gene Expression Omnibus (GEO) transcriptomic datasets (two human and two murine cohorts) with network-based inference to construct a NAFLD-related lncRNA-miRNA-mRNA coregulatory network. This is followed by candidate prioritization based on cross-dataset evidence and a literature review. Candidate lncRNAs are then experimentally validated in patient-derived blood samples using quantitative PCR (qPCR), and their diagnostic performance is quantified using receiver operating characteristic (ROC) analysis, both as individual markers and multi-lncRNA panels. Circulating lncRNAs are detected in diverse biofluids, remain stable under standard preanalytical conditions, and are often tissue-specific. This integrated approach facilitates the development of more precise, scalable, and noninvasive biomarkers for NAFLD/NASH. The chapter further emphasizes essential translational steps, including preanalytical standardization, analytical validation, and validation in independent patient cohorts with relevant clinical endpoints.
Proteomics, the large-scale study of proteins, enables the identification, quantification, and functional characterization of proteins, revealing post-translational modifications and protein interactions that are not apparent from transcriptomic data. Human organoids, which recapitulate the structural and functional complexity of native epithelial tissues, provide powerful tools to study disease mechanisms and personalize therapies. However, their culture poses challenges for efficient protein extraction and reproducible analysis. Here, we present a proteomics workflow optimized to maximize protein recovery from Matrigel-encased organoids. Samples were processed using S-Trap microcolumns to minimize losses, followed by liquid chromatography-mass spectrometry (LC-MS) in data-independent acquisition (DIA/SWATH-MS) mode for comprehensive, untargeted quantification. Library-free computational analysis using DIA-NN, combined with differential expression analysis, enabled sensitive detection of key proteins in intestinal organoids.
Nanopore sequencing is transforming viral genomics through real-time, portable, long-read analysis of RNA and DNA. Unlike traditional short-read platforms, it detects nucleotide sequences by measuring ionic current changes as nucleic acids pass through nanoscale pores, enabling direct single-molecule sequencing and base modification detection. Its simplicity, flexibility, and capacity for ultra-long reads make it ideal for resolving complex genomic regions, structural variants, and full viral genomes. These advantages have accelerated its use in pathogen surveillance and outbreak response, especially in resource-limited settings. For chikungunya virus (CHIKV), nanopore sequencing allows rapid, culture-independent recovery of complete genomes from clinical and vector samples, enabling real-time tracking of viral diversity, evolution, and spread. Experiences from Ebola, Zika, and COVID-19 have demonstrated the power of portable sequencing, now applied to CHIKV monitoring. Advances in tools such as Guppy, Dorado, Minimap2, and Medaka enhance read quality, consensus accuracy, and downstream analyses. Despite challenges in basecalling and error correction, robust quality control pipelines ensure reliable results. Ongoing improvements in chemistry, flow cell design, and machine learning will further enhance fidelity and throughput, establishing nanopore sequencing as a cornerstone of CHIKV genomic surveillance and epidemic preparedness.
Proteomics has been revealed as a key set of technologies that provide a detailed description of the molecular processes involved in the development of a specific phenotype. "Omics" technologies can collect an incredible amount of information. Among them, proteomics is an invaluable tool for defining specific biological information by studying the complete set of proteins under specific conditions, the proteome; or specific subsets of proteins, the subproteome. It is a crucial instrument for describing protein post-translational modifications, the functional annotation of the genome, and the detection of orphan genes. Protein extraction procedures are necessary to obtain B. cinerea protein extracts of sufficient quality to be analyzed by LC-MS/MS, avoiding contaminants that interfere with the identification process. After experimental design, collect the samples and replicates as defined in each experimental approach; we will describe protocols and procedures for the next steps of proteome and subproteome extraction and LC-MS analysis.
The Immuno-Oncology Biological Research (IOBR) package is an R-based analysis tool for exploring the tumor microenvironment (TME) and its influence on anti-tumor immunity. Built for high-throughput data-spanning both transcriptomic and genomic profiles-IOBR integrates six analytical modules, including transcriptomic data preprocessing, TME profiling, TME pattern identification, ligand-receptor interaction analysis, genome-TME interaction assessment, and visualization. In this chapter, we walk through a multi-omics workflow using example datasets, illustrating data preparation, distribution analyses, result interpretation, and graphical output. IOBR is open source and is available at https://github.com/IOBR/IOBR and a detailed GitBook ( https://iobr.github.io/book/ ) offers a complete manual and analysis guide for each function.
Teratoma formation is the gold standard assay for evaluating the developmental pluripotency of human and mouse embryonic stem cells (ESCs) and induced pluripotent stem cells (iPSCs). Following subcutaneous injection into immunodeficient mice, pluripotent stem cells spontaneously differentiate into derivatives representing all three embryonic germ layers-ectoderm, mesoderm, and endoderm. Beyond serving as a functional assay for pluripotency, teratomas provide a unique three-dimensional model system for studying early human development and lineage specification in vivo. This chapter describes comprehensive protocols for teratoma formation in immunodeficient mice, tissue processing for multiple downstream genomic applications, and multi-omics profiling approaches. We detail methods for embryonic stem cell culture, teratoma generation via subcutaneous injection, tissue dissection and processing for chromatin immunoprecipitation followed by sequencing (ChIP-Seq), RNA sequencing (RNA-Seq), single-cell multiome profiling combining chromatin accessibility (ATAC-Seq) and gene expression (scRNA-Seq), and histological analysis using hematoxylin and eosin (H&E) staining. Additionally, we provide bioinformatics workflows for analyzing the resulting genomic datasets to characterize the epigenetic and transcriptional landscapes of teratoma-derived tissues. These methods enable comprehensive molecular characterization of developmental processes and provide valuable resources for stem cell biologists studying pluripotency, differentiation, and early embryonic development.
The rapid growth of biomedical literature has created an urgent need for computational tools that enable researchers to systematically analyze publication trends, identify emerging research themes, and map the evolution of scientific fields. PubMed Atlas is a command-line and web-enabled workflow for topic-driven bibliometrics and trend intelligence using PubMed E-utilities. The pipeline executes PubMed queries, retrieves matching PMIDs, downloads full metadata records in batches, parses structured information (title, abstract, authors/affiliations, MeSH terms, publication types, grants, keywords, DOI), and stores normalized data in a local SQLite database for rapid querying and visualization. A Streamlit dashboard provides interactive exploration of publication trends, journal distributions, MeSH term summaries, geographic distributions, and recent article browsing with direct PubMed links. This protocol describes the installation, configuration, and operation of PubMed Atlas for cancer stem cell and stem cell transcriptional network research, and other fields, enabling investigators to conduct reproducible bibliometric analyses and identify knowledge gaps in rapidly evolving fields.
Cell-type identification is a crucial step in single-cell RNA-seq (scRNA-seq) data analysis, for which supervised methods are preferred due to their accuracy and efficiency. The quality of the reference data plays an important role in cell-type identification performance, but systematic strategies for selecting and reconstructing reference data remain limited. We present Target-Oriented Reference Construction (TORC), a widely applicable strategy for constructing reference data from available labeled cells given a target dataset. TORC alleviates the differences in data distribution and cell-type composition between the reference and the target. TORC combines initial supervised prediction, optional reference expansion using target cells with high-confidence predicted labels, and reference reconstruction guided by estimated cell-type compositions. Here, we provide detailed, step-by-step instructions describing the input requirements, configurable parameters, and practical considerations for applying TORC in real scRNA-seq analyses. TORC is available at https://github.com/weix21/TORC , where an example implementation using an MLP-based classifier is provided.
Colonic inflammation induces profound alterations in the intestinal mucosa that are evident at both macroscopic and microscopic levels, including disruption of epithelial barrier integrity, crypt fission, and immune cell infiltration. Recent advances in single-cell and spatial RNA sequencing have greatly expanded our understanding of the cellular diversity of the colonic mucosa, enabling pathological features to be directly linked to underlying cellular and molecular mechanisms. This chapter demonstrates the application of single-cell RNA sequencing (scRNA-seq) analysis techniques to interrogate stem cell dynamics in the context of inflammatory bowel disease. Using publicly available datasets, the chapter provides a step-by-step workflow implemented in Python, covering data access, loading, integration of multiple datasets, and initial preprocessing. Cells are mapped to large single-cell atlas references to infer cell identities, followed by a pseudobulk analysis strategy to assess inflammation-associated changes in cell phenotypes. Finally, trajectory inference approaches are applied to explore potential mechanisms governing the specification and modulation of key cell types during inflammation. Accompanied by an online resource containing fully annotated scripts, this chapter offers guided instruction in contemporary scRNA-seq analysis workflows. It is intended as an accessible introduction for researchers seeking to develop practical skills in single-cell data analysis that can be readily applied to their own biological questions.