Search bioRxivSearch

Biology subjects

Lee, J.-Y.

Publications and source records attributed to Lee, J.-Y..

5 recordsLinked to original sources

Proteomics of natural bacterial isolates powered by deep learning-based de novo identification.

Metaproteomics has been increasingly utilized for high-throughput molecular characterization in complex environments and has been demonstrated to provide insights into microbial composition and functional roles in soil systems. Despite its potential for the study of microbiomes, significant challenges remain in data analysis, including the creation of a sample-specific protein sequence database as the taxonomic composition of soil is often unknown. Almost all metaproteome analysis tools require this database and their accuracy and sensitivity suffer when the database is incomplete or contains extraneous sequences from organisms which are not present. Here, we leverage a de novo peptide sequencing approach to identify sample composition directly from metaproteomic data. First, we created a deep learning model, Kaiko, to predict the peptide sequences from mass spectrometry data, and trained it on 5 million peptide-spectrum matches from 55 phylogenetically diverse bacteria. After training, Kaiko successfully identified unsequenced soil isolates directly from proteomics data. Finally, we created a pipeline for metaproteome database generation using Kaiko. We tested the pipeline on native soils collected in Kansas, showing that the de novo sequencing model can be employed to construct the sample-specific protein database instead of relying on (un)matched metagenomes. Our pipeline identified all highly abundant taxa from 16S ribosomal RNA sequencing of the soil samples and also uncovered several additional species which were strongly represented only in proteomic data. Our pipeline offers an alternative and complementary method for metaproteomic data analysis by creating a protein database directly from proteomic data, thus removing the need for metagenomic sequencing. Significance StatementProteomic characterization of environmental samples, or metaproteomics, reveals microbial activity critical to our understanding of climate, nutrient cycling and human health. Metaproteomic samples originate from diverse environs, such as soil and oceans. One option for data analysis is a de novo interpretation of the mass spectra. Unfortunately, the current generation of de novo algorithms were primarily trained on data originating from human proteins. Therefore, these algorithms struggle with data from environmental samples, limiting our ability to analyze metaproteomics data. To address this challenge, we trained a new algorithm with data from dozens of diverse environmental bacteria and achieved significant improvements in accuracy across a broad range of organisms. This generality opens proteomics to the world of natural isolates and microbiomes.

bioinformatics

Leveraging chromatin accessibility for transcriptional regulatory network inference in T Helper 17 Cells

Transcriptional regulatory networks (TRNs) provide insight into cellular behavior by describing interactions between transcription factors (TFs) and their gene targets. The Assay for Transposase Accessible Chromatin (ATAC)-seq, coupled with transcription-factor motif analysis, provides indirect evidence of chromatin binding for hundreds of TFs genome-wide. Here, we propose methods for TRN inference in a mammalian setting, using ATAC-seq data to influence gene expression modeling. We rigorously test our methods in the context of T Helper Cell Type 17 (Th17) differentiation, generating new ATAC-seq data to complement existing Th17 genomic resources (plentiful gene expression data, TF knock-outs and ChIP-seq experiments). In this resource-rich mammalian setting, our extensive benchmarking provides quantitative, genome-scale evaluation of TRN inference combining ATAC-seq and RNA-seq data. We refine and extend our previous Th17 TRN, using our new TRN inference methods to integrate all Th17 data (gene expression, ATAC-seq, TF KO, ChIP-seq). We highlight new roles for individual TFs and groups of TFs (\"TF-TF modules\") in Th17 gene regulation. Given the popularity of ATAC-seq, which provides high-resolution with low sample input requirements, we anticipate that application of our methods will improve TRN inference in new mammalian systems, especially in vivo, for cells directly from humans and animal models.

systems biology

WDR11-mediated Hedgehog signalling defects underlie a new ciliopathy related to Kallmann syndrome

WDR11 has been implicated in congenital hypogonadotropic hypogonadism (CHH) and Kallmann syndrome (KS), human developmental genetic disorders defined by delayed puberty and infertility. However, WDR11s role in development is poorly understood. Here we report that WDR11 modulates the Hedgehog (Hh) signalling pathway and is essential for ciliogenesis. Disruption of WDR11 expression in mouse and zebrafish results in phenotypic characteristics associated with defective Hh signalling, accompanied by dysgenesis of ciliated tissues. Wdr11 null mice also exhibit early onset obesity. We found that WDR11 shuttles from the cilium to the nucleus in response to Hh signalling. WDR11 was also observed to regulate the proteolytic processing of GLI3 and cooperate with EMX1 transcription factor to induce the expression of downstream Hh pathway genes and gonadotrophin releasing hormone production. The CHH/KS-associated human mutations result in loss-of-function of WDR11. Treatment with the Hh agonist purmorphamine partially rescued the WDR11-haploinsufficiency phenotypes. Our study reveals a novel class of ciliopathy caused by WDR11 mutations and suggests that CHH/KS may be a part of the human ciliopathy spectrum.

developmental biology

Blazing Signature Filter: a library for fast pairwise similarity comparisons

Identifying similarities between datasets is a fundamental task in data mining and has become an integral part of modern scientific investigation. Whether the task is to identify co-expressed genes in large-scale expression surveys or to predict combinations of gene knockouts which would elicit a similar phenotype, the underlying computational task is often a multi-dimensional similarity test. As datasets continue to grow, improvements to the efficiency, sensitivity or specificity of such computation will have broad impacts as it allows scientists to more completely explore the wealth of scientific data. A significant practical drawback of large-scale data mining is that the vast majority of pairwise comparisons are unlikely to be relevant, meaning that they do not share a signature of interest. It is therefore essential to efficiently identify these unproductive comparisons as rapidly as possible and exclude them from more time-intensive similarity calculations. The Blazing Signature Filter (BSF) is a highly efficient pairwise similarity algorithm which enables extensive data mining within a reasonable amount of time. The algorithm transforms datasets into binary metrics, allowing it to utilize the computationally efficient bit operators and provide a coarse measure of similarity. As a result, the BSF can scale to high dimensionality and rapidly filter unproductive pairwise comparison. Two bioinformatics applications of the tool are presented to demonstrate the ability to scale to billions of pairwise comparisons and the usefulness of this approach.

bioinformatics

Xbra and Smad-1 response elements cooperate in PV.1 promoter to inhibit the early neurogenesis in Xenopus embryos

Crosstalk of signaling pathways plays crucial roles in cell fate determination, cell differentiation and proliferation. Both BMP-4/Smad-1 and FGF/Xbra signaling induce the expression of PV.1, leading to neural inhibition. However, BMP-4/Smad-1 and FGF/Xbra signaling crosstalk in the regulation of PV.1 transcription is still largely unknown. In this study, Smad-1 and Xbra physically interacted and regulated the PV.1 transcriptional activation in a synergistic manner. Xbra and Smad-1 directly bound within the proximal region of the PV.1 promoter and cooperatively enhanced the binding of an interacting partner within the promoter. Maximum cooperation was achieved in the presence of intact DNA binding sites of both Smad-1 and Xbra. Collectively, BMP-4/Smad-1 and FGF/Xbra signal crosstalk was required to activate the PV.1 transcription, synergistically. Suggesting that crosstalk of BMP-4 and FGF signaling facilitates the fine-tuning regulation of PV.1 transcription to inhibit neurogenesis during embryonic development of Xenopus.\n\nSummary statementFGF/Xbra positively regulates the PV.1 expression in the Xenopus via an unknown mechanism. Our study shows that both BMP-4/Smad-1 and FGF/Xbra exhibits a signaling crosstalk to regulate PV.1 transcription activation, promoting to ectoderm and mesoderm formation and inhibiting the early neurogenesis in Xenopus.

developmental biology