Search bioRxivSearch

Biology subjects

Hou, Y.

Publications and source records attributed to Hou, Y..

16 recordsLinked to original sources

BioIMA: a one-click desktop tool for standardized extraction of phenotypic traits from biological images

Standardized extraction of quantitative phenotypes from images is increasingly important across plant biology, from ecological and evolutionary studies to genetics, breeding, and functional genomics. However, as large image datasets are increasingly used for trait analysis, many biologically relevant traits, including size, shape, color, and spatial patterning, are still measured manually or using fragmented semi-automated workflows. These limitations reduce throughput, reproducibility, and accessibility, especially for researchers without computational expertise. Here, we present BioIMA, an open-source desktop tool for rapid and standardized phenotyping from biological images. BioIMA integrates foundation model-based segmentation with automated trait computation, allowing users to extract quantitative measurements from images through an intuitive graphical interface and without model training. To validate its performance, we quantified a set of knot morphological traits in two Populus species, as these measurements are typically time-consuming to perform manually. Automatic measurements showed strong agreement with manual ImageJ-based measurements (R2 > 0.95), while reducing per-image processing time by approximately 75% (from ~15 s to ~4 s). BioIMA was further applied to diverse plant datasets, including Helianthus and Rhododendron images with varying morphologies and background conditions. Although developed for plant phenotyping, BioIMA may also be extended to other biological samples where region-based size, shape, or color traits are of interest. By combining accessibility and standardization in a lightweight local application, BioIMA provides a practical community resource for image-based phenotyping in ecological and evolutionary studies.

bioinformatics

Synergistic antifungal effect of amphotericin B-loaded PLGA nanoparticles based ultrasound against C. albicans biofilms

C. albicans is human opportunistic pathogens that cause superficial and life-threatening infections. An important reason for the failure of current antifungal drugs is related to biofilm formation mostly associated with implanted medical device. The present study aims to investigate the synergistic antifungal efficacy of low-frequency and low-intensity ultrasound combined with amphotericin B-loaded PLGA nanoparticles (AmB-NPs) on C.albicans biofilms. AmB-NPs were prepared by a double emulsion method and demonstrated the lower toxicity than free AmB, after which biofilms were established and treated with ultrasound and AmB-NPs separately or jointly in vitro and in vivo. The results demonstrated the activity, biomass, and proteinase and phospholipase activities of biofilms were decreased significantly after the combination treatment of AmB-NPs with 42 KHz ultrasound irradiation at an intensity of 0.30 W/cm2 for 15 min compared to the control, the AmB alone or the ultrasound alone treatment (P < 0.01), and the morphology of biofilms was altered remarkably after jointly treatment under CLSM observation and detection, especially thickness thinning and structure loosing. Furthermore, the same synergistic effects were proved in a subcutaneous catheter biofilm rat model. The result of colony forming units of catheter fungus loading exhibited a significant reduction after AmB-NPs and ultrasound jointly treatment for 7 days continuous therapy, and the CLSM images revealed that the biofilm on the catheter surface was substantially eliminated. Our study may provide a new noninvasive, safe and effective application to C.albicans biofilm infection therapy.

microbiology

Integrated Analysis Revealed Hub Genes in Breast Cancer

The aim of this study was to identify the hub genes in breast cancer and provide further insight into the tumorigenesis and development of breast cancer. To explore the hub genes in breast cancer, we performed an integrated bioinformatics analysis. Two gene expression profiles were downloaded from the GEO database. The differentially expressed genes (DEGs) were identified by using the \"limma\" package. Then, we performed Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis to explore the functional annotation and potential pathways of the DEGs. Next, protein-protein interaction (PPI) network analysis and weighted gene coexpression network analysis (WGCNA) were conducted to screen for hub genes. To confirm the reliability of the identified hub genes, we obtained TCGA-BRCA data by using WGCNA to screen for genes that were strongly related to breast cancer. By combining the results from the GEO and TCGA datasets, we finally identified 15 real hub genes in breast cancer. Finally, we performed an overall survival analysis to explore the connection between the expression of hub genes and the overall survival time of breast cancer patients. We found that for all hub genes, higher expression was associated with significantly shorter overall survival times among breast cancer patients.

bioinformatics

Sequencing of the MHC region defines HLA-DQA1 as the major independent risk for anti-citrullinated protein antibodies (ACPA)-positive rheumatoid arthritis in Han population

The strong genetic contribution of the major histocompatibility complex (MHC) to rheumatoid arthritis (RA) susceptibility has been generally attributed to HLA-DRB1. However, due to the high linkage disequilibrium in the MHC region, it is difficult to define the real or/and additional independent genetic risks using the conventional HLA genotyping or chip-based microarray technology. By the capture sequencing of entire MHC region for discovery and HLA-typing for validation in 2,773 subjects of Han ancestry, we identified HLA-DQ1:160D as the strongest independent genetic risk for anti-citrullinated protein antibodies (ACPA)-positive RA in Han population (P = 6.16 x 10-36, OR=2.29). Further stepwise conditional analysis revealed that DR{beta}1:37N has an independent protective effect on ACPA-positive RA (P = 5.81 x 10-16, OR=0.49). The DQ1:160 coding allele DQA1*0303 displayed high impact on joint radiographic severity, especially in patients with early disease and smoking (P = 3.02 x 10-5). Interaction analysis by comparative molecular modeling revealed that the negative charge of DQ1:160D stabilizes the dimer of dimers, leading to an increased T cell activation. The electrostatic potential surface analysis indicated that the negative charged DR{beta}1:37N encoding alleles could bind with epitope P9 arginine, thus may result in a decreased RA susceptibility.\n\nIn this study, we provide the first evidence that HLA-DQA1, instead of HLA-DRB1, is the strongest and independent genetic risk for ACPA-positive RA in Chinese Han population. Our study also illustrates the value of MHC deep sequencing for fine mapping disease risk variants in the MHC region.

genetics

PIRD: Pan immune repertoire database

MotivationT and B cell receptors (TCRs and BCRs) play a pivotal role in the adaptive immune system by recognizing an enormous variety of external and internal antigens. Understanding these receptors is critical for exploring the process of immunoreaction and exploiting potential applications in immunotherapy and antibody drug design. Although a large number of samples have had their TCR and BCR repertoires sequenced using high-throughput sequencing in recent years, very few databases have been constructed to store these kinds of data. To resolve this issue, we developed a database.\n\nResultsWe developed a database, the Pan Immune Repertoire Database (PIRD), located in China National GeneBank (CNGBdb), to collect and store annotated TCR and BCR sequencing data, including from Homo sapiens and other species. In addition to data storage, PIRD also provides functions of data visualisation and interactive online analysis. Additionally, a manually curated database of TCRs and BCRs targeting known antigens (TBAdb) was also deposited in PIRD.\n\nAvailability and ImplementationPIRD can be freely accessed at https://db.cngb.org/pird.

immunology

Comprehensive analysis of immune evasion in breast cancer by single-cell RNA-seq

The tumor microenvironment is composed of numerous cell types, including tumor, immune and stromal cells. Cancer cells interact with the tumor microenvironment to suppress anticancer immunity. In this study, we molecularly dissected the tumor microenvironment of breast cancer by single-cell RNA-seq. We profiled the breast cancer tumor microenvironment by analyzing the single-cell transcriptomes of 52,163 cells from the tumor tissues of 15 breast cancer patients. The tumor cells and immune cells from individual patients were analyzed simultaneously at the single-cell level. This study explores the diversity of the cell types in the tumor microenvironment and provides information on the mechanisms of escape from clearance by immune cells in breast cancer.\n\nOne Sentence SummaryLandscape of tumor cells and immune cells in breast cancer by single cell RNA-seq

cancer biology

Single-cell Transcriptomic Landscape of Nucleated Cells in Umbilical Cord Blood

Umbilical cord blood (UCB) transplant is a therapeutic option for both pediatric and adult patients with a variety of hematologic diseases such as several types of blood cancers, myeloproliferative disorders, genetic diseases, and metabolic disorders. However, the level of cellular heterogeneity and diversity of nucleated cells in the UCB has not yet been assessed in an unbiased and systemic fashion. In the current study, nucleated cells from UCB were subjected to single-cell RNA sequencing, a technology enabled simultaneous profiling of the gene expression signatures of thousands of cells, generating rich resources for further functional studies. Here, we report the transcriptomic maps of 19,052 UCB cells, covering 11 major cell types. Many of these cell types are comprised of distinct subpopulations, including distinct signatures in NK and NKT cell types in the UCB. Pseudotime ordering of nucleated red blood cells (NRBC) identifies wave-like activation and suppression of transcription regulators, leading to a polarized cellular state, which may reflect the NRBC maturation. Progenitor cells in the UBC also consist two subpopulations with divergent transcription programs activated, leading to specific cell-fate commitment. Collectively, we provide this comprehensive single-cell transcriptomic landscape and show that it can uncover previously unrecognized cell types, pathways and gene expression regulations that may contribute to the efficacy and outcome of UCB transplant, broadening the scope of research and clinical innovations.

genomics

Deconvolution of single-cell multi-omics layers reveals regulatory heterogeneity

Integrative analysis of multi-omics layers at single cell level is critical for accurate dissection of cell-to-cell variation within certain cell populations. Here we report scCAT-seq, a technique for simultaneously assaying chromatin accessibility and the transcriptome within the same single cell. We show that the combined single cell signatures enable accurate construction of regulatory relationships between cis-regulatory elements and the target genes at single-cell resolution, providing a new dimension of features that helps direct discovery of regulatory patterns specific to distinct cell identities. Moreover, we generated the first single cell integrated maps of chromatin accessibility and transcriptome in human pre-implantation embryos and demonstrated the robustness of scCAT-seq in the precise dissection of master transcription factors in cells of distinct states during embryo development. The ability to obtain these two layers of omics data will help provide more accurate definitions of \"single cell state\" and enable the deconvolution of regulatory heterogeneity from complex cell populations.

genomics

Negative Cooperativity between Gemin2 and RNA Determines RNA Selection and Release of the SMN Complex in snRNP Assembly

The assembly of snRNP cores, in which seven Sm proteins, D1/D2/F/E/G/D3/B, form a ring around the nonameric Sm site of snRNAs, is the early step of spliceosome formation and essential to eukaryotes. It is mediated by the PMRT5 and SMN complexes sequentially in vivo. SMN deficiency causes neurodegenerative disease spinal muscular atrophy (SMA). How the SMN complex assembles snRNP cores is largely unknown, especially how the SMN complex achieves high RNA assembly specificity and how it is released. Here we show, using crystallographic and biochemical approaches, that Gemin2 of the SMN complex enhances RNA specificity of SmD1/D2/F/E/G via a negative cooperativity between Gemin2 and RNA in binding SmD1/D2/F/E/G. Gemin2, independent of its N-tail, constrains the horseshoe-shaped SmD1/D2/F/E/G from outside in a physiologically relevant, narrow state, enabling high RNA specificity. Moreover, the assembly of RNAs inside widens SmD1/D2/F/E/G, causes the release of Gemin2/SMN allosterically and allows SmD3/B to join. The assembly of SmD3/B further facilitates the release of Gemin2/SMN. This is the first to show negative cooperativity in snRNP assembly, which provides insights into RNA selection and the SMN complexs release. These findings reveal a basic mechanism of snRNP core assembly and facilitate pathogenesis studies of SMA.

molecular biology

Nanjinganthus:An Unexpected Flower from the Jurassic of China

The origin of angiosperms has been the focus of intensive botanical debate for well over a century. The great diversity of angiosperms in the Early Cretaceous makes the Jurassic rather expected to elucidate the origin of angiosperm. Former reports of early angiosperms are frequently based on a single specimen, making many conclusions tentative. Here, based on observations of 284 individual flowers preserved on 28 slabs in various states and orientations, we describe a fossil flower, Nanjinganthus dendrostyla gen. et sp. nov., from the South Xiangshan Formation (Early Jurassic) of China. The large number of specimens and various preservations allows us to give an evidenced interpretation of the flower. The complete enclosure of ovules in Nanjinganthus is fulfilled by a combination of an invaginated and ovarian roof. Characterized by its actinomorphic flower with a dendroid style, cup-form receptacle, and angio-ovuly, Nanjinganthus is a bona fide angiosperm from the Jurassic. Nanjinganthus re-confirms the existence of Jurassic angiosperms and provides first-hand raw data for new analyses on the origin and history of angiosperms.

paleontology

Overlooked polyploidies in lycophytes generalize their roles during the evolution of vascular plants

Seed plants and lycophytes constitute the extant vascular plants. As a model lycophyte, Selaginalla moellendroffii was deciphered its genome, previously proposed to have avoided polyploidies, as key events contributing to the origination and fast expansion of seed plants. Here, using a gold-standard streamline recently proposed to deconvolute complex genomes, we reanalyzed the S. moellendroffii genome. To our surprise, we found clear evidence of multiple paleo-polyploidies, with one being recent (~ 13-15 millions of years ago or Mya), another one occurring about ~125-142 Mya, during the evolution of lycophytes, and at least 2 or 3 events being more ancient. Besides, comparison of reconstructed ancestral genomes of lycophytes and angiosperms shows that lycophytes were likely much more affected by paleo-polyploidies than seed plants. The present analysis here provides clear and solid evidence that polyploidies have contributed the successful establishment of all vascular plants on earth.

evolutionary biology

Pan-cancer study of heterogeneous RNA aberrations

We present the most comprehensive catalogue of cancer-associated gene alterations through characterization of tumor transcriptomes from 1,188 donors of the Pan-Cancer Analysis of Whole Genomes project. Using matched whole-genome sequencing data, we attributed RNA alterations to germline and somatic DNA alterations, revealing likely genetic mechanisms. We identified 444 associations of gene expression with somatic non-coding single-nucleotide variants. We found 1,872 splicing alterations associated with somatic mutation in intronic regions, including novel exonization events associated with Alu elements. Somatic copy number alterations were the major driver of total gene and allele-specific expression (ASE) variation. Additionally, 82% of gene fusions had structural variant support, including 75 of a novel class called \"bridged\" fusions, in which a third genomic location bridged two different genes. Globally, we observe transcriptomic alteration signatures that differ between cancer types and have associations with DNA mutational signatures. Given this unique dataset of RNA alterations, we also identified 1,012 genes significantly altered through both DNA and RNA mechanisms. Our study represents an extensive catalog of RNA alterations and reveals new insights into the heterogeneous molecular mechanisms of cancer gene alterations.

genomics

Layered structure and complex mechanochemistry of a strong bacterial adhesive

While designing adhesives that perform in aqueous environments has proven challenging for synthetic adhesives, microorganisms commonly produce bioadhesives that efficiently attach to a variety of substrates, including wet surfaces that remain a challenge for industrial adhesives. The aquatic bacterium Caulobacter crescentus uses a discrete polar polysaccharide complex, the holdfast, to strongly attach to surfaces and resist flow. The holdfast is extremely versatile and has an impressive adhesive strength. Here, we use atomic force microscopy (AFM) to unravel the complex structure of the holdfast and characterize its chemical constituents and their role in adhesion. We used purified holdfasts to dissect the intrinsic properties of this component as a biomaterial, without the effect of the bacterial cell body. Our data support a model where the holdfast is a heterogeneous material composed of two layers: a stiff nanoscopic core, covered by a sparse, flexible brush layer. These two layers contain not only N-acetyl-D-glucosamine (NAG), the only yet identified component present in the holdfast, but also peptides and DNA, which provide structure and adhesive character. Biochemical experiments suggest that, while polypeptides are the most important components for adhesive force, the presence of DNA mainly impacts the brush layer and initial adhesion, and NAG plays a primarily structural role within the core. Moreover, our results suggest that holdfast matures structurally, becoming more homogeneous over time. The unanticipated complexity of both the structure and composition of the holdfast likely underlies its distinctive strength as a wet adhesive and could inform the development of a versatile new family of adhesives.

biophysics

TASI: A software tool for spatial-temporal quantification of tumor spheroid dynamics

Spheroid cultures derived from explanted cancer specimens are an increasingly utilized resource for studying complex biological processes like tumor cell invasion and metastasis, representing an important bridge between the simplicity and practicality of 2D monolayer cultures and the complexity and realism of in vivo animal models. Temporal imaging of spheroids can capture the dynamics of cell behaviors and microenvironments, and when combined with quantitative image analysis methods, enables deep interrogation of biological mechanisms. This paper presents a comprehensive open-source software framework for Temporal Analysis of Spheroid Imaging (TASI) that allows investigators to objectively characterize spheroid growth and invasion dynamics. TASI performs spatiotemporal segmentation of spheroid cultures, extraction of features describing spheroid morpho-phenotypes, mathematical modeling of spheroid dynamics, and statistical comparisons of experimental conditions. We demonstrate the utility of this tool in an analysis of non-small cell lung cancer spheroids that exhibit variability in metastatic and proliferative behaviors.

bioinformatics

Multiplexed sgRNA Expression Allows Versatile Single Non-repetitive DNA Labeling and Endogenous Gene Regulation

The CRISPR/Cas9 system has made significant contribution to genome editing, gene regulation and chromatin studies in recent years. High-throughput and systematic investigations into the multiplexed biological systems and disease conditions require simultaneous expression and coordinated functioning of multiple sgRNAs. However, current co-transfection based sgRNA co-expression systems remain poorly efficient and virus-based transfection approaches are relatively costly and labor intensive. Here we established a vector-independent method allowing multiple sgRNA expression cassettes to be assembled in series into a single plasmid. This synthetic biology-based strategy excels in its efficiency, controllability and scalability. Taking the flexibility advantage of this all-in-one sgRNA expressing system, we further explored its applications in single non-repetitive genomic locus imaging as well as coordinated gene regulation in live cells. With its strong potency, our method will greatly facilitate the understandings in genome structure, function and dynamics, and will contribute to the systemic investigations into complex physiological and pathological conditions.

synthetic biology

Long-read whole genome sequencing identifies causal structural variation in a Mendelian disease

Current clinical genomics assays primarily utilize short-read sequencing (SRS), which offers high throughput, high base accuracy, and low cost per base. SRS has, however, limited ability to evaluate tandem repeats, regions with high [GC] or [AT] content, highly polymorphic regions, highly paralogous regions, and large-scale structural variants. Long-read sequencing (LRS) has complementary strengths and offers a means to discover overlooked genetic variation in patients undiagnosed by SRS. To evaluate LRS, we selected a patient who presented with multiple neoplasia and cardiac myxomata suggestive of Carney complex for whom targeted clinical gene testing and whole genome SRS were negative. Low coverage whole genome LRS was performed on the PacBio Sequel system and structural variants were called, yielding 6,971 deletions and 6,821 insertions > 50bp. Filtering for variants that are absent in an unrelated control and that overlap a coding exon of a disease gene identified three deletions and three insertions. One of these, a heterozygous 2,184 bp deletion, overlaps the first coding exon of PRKAR1A, which is implicated in autosomal dominant Carney complex. This variant was confirmed by Sanger sequencing and was classified as pathogenic using standard criteria for the interpretation of sequence variants. This first successful application of whole genome LRS to identify a pathogenic variant suggests that LRS has significant potential to identify disease-causing structural variation. We recommend larger studies to evaluate the diagnostic yield of LRS, and the development of a comprehensive catalog of common human structural variation to support future studies.

genomics