Search bioRxivSearch

Biology subjects

Li, R.

Publications and source records attributed to Li, R..

32 records · Page 2Linked to original sources

GNE: A deep learning framework for gene network inference by aggregating biological information

The topological landscape of gene interaction networks provides a rich source of information for inferring functional patterns of genes or proteins. However, it is still a challenging task to aggregate heterogeneous biological information such as gene expression and gene interactions to achieve more accurate inference for prediction and discovery of new gene interactions. In particular, how to generate a unified vector representation to integrate diverse input data is a key challenge addressed here. We propose a scalable and robust deep learning framework to learn embedded representations to unify known gene interactions and gene expression for gene interaction predictions. These low-dimensional embeddings derive deeper insights into the structure of rapidly accumulating and diverse gene interaction networks and greatly simplify downstream modeling. We compare the predictive power of our deep embeddings to the strong baselines. The results suggest that our deep embeddings achieve significantly more accurate predictions. Moreover, a set of novel gene interaction predictions are validated by up-to-date literature-based database entries. GNE is freely available under the GNU General Public License and can be downloaded from Github (https://github.com/kckishan/GNE)

systems biology

Collective feature selection to identify crucial epistatic variants

BackgroundMachine learning methods have gained popularity and practicality in identifying linear and non-linear effects of variants associated with complex disease/traits. Detection of epistatic interactions still remains a challenge due to the large number of features and relatively small sample size as input, thus leading to the so-called \"short fat data\" problem. The efficiency of machine learning methods can be increased by limiting the number of input features. Thus, it is very important to perform variable selection before searching for epistasis. Many methods have been evaluated and proposed to perform feature selection, but no single method works best in all scenarios. We demonstrate this by conducting two separate simulation analyses to evaluate the proposed collective feature selection approach.\n\nResultsThrough our simulation study we propose a collective feature selection approach to select features that are in the \"union\" of the best performing methods. We explored various parametric, non-parametric, and data mining approaches to perform feature selection. We choose our top performing methods to select the union of the resulting variables based on a user-defined percentage of variants selected from each method to take to downstream analysis. Our simulation analysis shows that non-parametric data mining approaches, such as MDR, may work best under one simulation criteria for the high effect size (penetrance) datasets, while non-parametric methods designed for feature selection, such as Ranger and Gradient boosting, work best under other simulation criteria. Thus, using a collective approach proves to be more beneficial for selecting variables with epistatic effects also in low effect size datasets and different genetic architectures. Following this, we applied our proposed collective feature selection approach to select the top 1% of variables to identify potential interacting variables associated with Body Mass Index (BMI) in ~44,000 samples obtained from Geisingers MyCode Community Health Initiative (on behalf of DiscovEHR collaboration).\n\nConclusionsIn this study, we were able to show that selecting variables using a collective feature selection approach could help in selecting true positive epistatic variables more frequently than applying any single method for feature selection via simulation studies. We were able to demonstrate the effectiveness of collective feature selection along with a comparison of many methods in our simulation analysis. We also applied our method to identify non-linear networks associated with obesity.

bioinformatics

Transcriptome Landscape of Human Oocytes and Granulosa Cells Throughout Folliculogenesis

Folliculogenesis is a highly regulated process that involves bidirectional interactions of the oocytes and surrounding granulosa cells (GCs). Little is unknown, however, about the transcriptomic profiles of human oocytes and GCs throughout folliculogenesis. Here we performed a high resolution RNA-Seq of human oocytes and GCs at each follicular stage, which revealed unique transcriptional profiles, stage-specific signature genes, oocyte- and GC-derived genes that reflect ovarian reserve. We identified reciprocal cell-to-cell interactions between oocytes and GCs, including NOTCH, TGF-{beta} signaling and gap junctions and determined the expression patterns of maternal-effect genes involved in folliculogenesis and early embryogenesis. Finally, we demonstrated robust differences between human and mice oocyte transcriptomes. This is the first comprehensive overview of the transcriptomic signatures governing the stepwise human folliculogenesis in-vivo that provides a valuable resource for basic and translational research in human reproductive biology.

cell biology

Multi-hierarchical Profiling the Structure-Activity Relationships of Engineered Nanomaterials at Nano-Bio Interfaces

Increasingly raised concerns (nanotoxicity, clinical translation, etc) on nanotechnology require breakthroughs in structure-activity relationship (SAR) analyses of engineered nanomaterials (ENMs) at nano-bio interfaces. However, current nano-SAR assessments failed to disclosure sufficient information to understand ENM-induced bio-effects. Here we developed a multi-hierarchical nano-SAR assessment for a representative ENM, Fe2O3 by systematically examining cellular metabolite and protein changes. This nano-SAR profile allows visualizing the contributions of 7 basal properties of Fe2O3 to their diverse bio-effects. For instance, while surface reactivity is responsible for Fe2O3-induced cell migration, the inflammatory effects of Fe2O3 nanorods and nanoplates are determined by their aspect ratio and surface reactivity, respectively. We further discovered the detailed mechanisms, including NLRP3 inflammasome pathway and monocyte chemoattractant protein-1 involved signaling. Both effects were further validated in animal lungs. Our findings provide substantial new insights at nano-bio interfaces, which may facilitate the tailored design of ENMs to endow them with desired bio-effects.

pharmacology and toxicology

Tumour purity as a prognostic factor in colon cancer

Tumour purity is defined as the proportion of cancer cells in the tumour tissue. The impact of tumour purity on colon cancer (CC) prognosis, genetic profile and microenvironment has not been thoroughly accessed. Therefore, clinical and transcriptomic data from three public datasets, GSE17536/17537, GSE39582, and TCGA were retrospectively collected (n = 1248). Tumour purity of each sample was inferred by a computational method based on transcriptomic data. Stage III and MMR-deficient (dMMR) CC patients showed a significantly lower tumour purity. Low purity CC conferred worse survival and tumour purity was identified as an independent prognostic factor. Moreover, high tumour purity CC patients benefited more from adjuvant chemotherapy. Subsequent genomic analysis found that the mutation burden was negatively associated with tumour purity with only APC and KRAS significantly more mutated in high purity CC. However, no somatic copy number alteration event was correlated with tumour purity. Furthermore, immune-related pathways and immunotherapy-associated markers (PD-1, PD-L1, CTLA-4, LAG-3, and TIM-3) were highly enriched in low purity samples. Notably, the relative proportion of M2 macrophages and neutrophils, which indicated worse survival in CC, was negatively associated with tumour purity. Therefore, tumour purity exhibited potential value for CC prognostic stratification as well as adjuvant chemotherapy benefit prediction. The relative worse survival in low purity CC may attribute to higher mutation frequency in key pathways and purity related microenvironmental changing.\n\nSummaryLow purity colon cancer patients conferred worse survival and benefited less from adjuvant chemotherapy. The mutation burden was negatively associated with tumour purity. Low purity samples exhibited intense immune phenotype with more M2 macrophages and neutrophils infiltration.

cancer biology

Optimizing Trait Predictability in Hybrid Rice Using Superior Prediction Models and Selective Omic Datasets

Hybrid breeding has dramatically boosted yield and its stability in rice. Genomic prediction further benefits rice breeding by increasing selection intensity and accelerating breeding cycles. With the rapid advancement of technology, other omic data, such as metabolomic data and transcriptomic data, are readily available for predicting genetic values (or breeding values) for agronomically important traits. In the current study, we searched for the best prediction strategy for four traits (yield, 1000 grain weight, number of grains per panicle and number of tillers per plant) of hybrid rice by evaluating all possible combinations of omic datasets with different prediction methods. We conclude that, in rice, the predictions using the combination of genomic and metabolomic data generally produce better results than single-omics predictions or predictions based on other combined omic data. Inclusion of transcriptomic data does not improve predictability possibly because transcriptome does not provide more information for the trait than the sum of genome and metabolome; rather, the computational complexity is substantially increased if transcriptomic data is included in the models. Best linear unbiased prediction (BLUP) appears to be the most efficient prediction method compared to the other commonly used approaches, including LASSO, SSVS, SVM-RBF, SVP-POLY and PLS. Our study has provided a guideline for selection of hybrid rice in terms of which types of omic datasets and which method should be used to achieve higher trait predictability.

genetics

First Efficient Transfection in Choanoflagellates using Cell-Penetrating Peptides

Only recently, based on phylogenetic studies choanoflagellates have been confirmed to form the sister group to metazoan. The mechanisms and genes behind the step from single to multicellular organisation and as a consequence the evolution of metazoan multicellularity could not be verified yet, as no reliable and efficient method for transfection of choanoflagellates was available. Here we present cell-penetrating peptides (CPPs) as an alternative to conventional transfection methods. In a series of experiments with the choanoflagellate Diaphanoeca grandis we proof for the first time that the use of CPPs is a reliable and highly efficient method for the transfection of choanoflagellates. We were able to silence the silicon transporter gene (SIT) by siRNA, and hence, to suppress the lorica (characteristic siliceous basket) formation. High gene silencing efficiency was determined and measured by light microscope and RT-qPCR. In addition, only low cytotoxic effects of CPP were detected. Our new method allows the reliable and efficient transfection of choanoflagellates, finally enabling us to verify the function of genes, thought to be involved in cell adhesion or cell signaling by silencing them via siRNA. This is a step stone for the research on the origin of multicellularity in metazoans.

molecular biology

New insights on human essential genes based on integrated multi-omics analysis

Essential genes are those whose functions govern critical processes that sustain life in the organism. Recent gene-editing technologies have provided new opportunities to characterize essential genes. Here, we present an integrated analysis for comprehensively and systematically elucidating the genetic and regulatory characteristics of human essential genes. First, essential genes act as \"hubs\" in protein-protein interactions networks, in chromatin structure, and in epigenetic modifications, thus are essential for cell growth. Second, essential genes represent the conserved biological processes across species although gene essentiality changes itself. Third, essential genes are import for cell development due to its discriminate transcription activity in both embryo development and oncogenesis. In addition, we develop an interactive web server, the Human Essential Genes Interactive Analysis Platform (HEGIAP) (http://sysomics.com/HEGIAP/), which integrates abundant analytical tools to give a global, multidimensional interpretation of gene essentiality. Our study provides a new view for understanding human essential genes.

genomics

TrackSig: reconstructing evolutionary trajectories of mutation signature exposure

We present a new method, TrackSig, to estimate the evolutionary trajectories of signatures of different somatic mutational processes from DNA sequencing data from a single, bulk tumour sample. TrackSig uses probability distributions over mutation types, called mutational signatures, to represent different mutational processes and detects the changes in the signature activity using an optimal segmentation algorithm that groups somatic mutations based on their estimated cancer cellular fraction (CCF) and their mutation type (e.g. CAG->CTG). We use two different simulation frameworks to assess both TrackSigs reconstruction accuracy and its robustness to violations of its assumptions, as well as to compare it to a baseline approach. We find 2-4% median error in reconstructing the signature activities on simulations with varying difficulty with one to three subclones at an average depth of 30x. The size and the direction of the activity change is consistent in 83% and 95% of cases respectively. There were an average of 0.02 missed and 0.12 false positive subclones per sample. In our simulations, grouping mutations by mutation type (TrackSig), rather than by clustering CCF (baseline strategy), performs better at estimating signature activities and at identifying subclonal populations in the complex scenarios like branching, CNA gain, violation of infinite site assumption, and the inclusion of neutrally evolving mutations. TrackSig is open source software, freely available at https://github.com/morrislab/TrackSig.

bioinformatics

Mechanism of Membrane Recovery in Intra-Cytoplasmic Sperm Injection

ICSI (Intra-cytoplasmic sperm injection) is a broadly utilized technique for artificial fertilization. This approach has been successfully performed in human oocytes as well as others such as mouse and bovine. The piercing through the zona layer and the membrane needs to be achieved with a minimal biological damage to facilitate a rapid healing. Since the injection methodology serves as a crucial factor to success rate of ICSI, a significant amount of research efforts has been devoted to the development of injections. In this paper, we conduct comparative study among the major milestones for injection techniques in ICSI. Technical details are provided for each milestone and each technique is evaluated from engineering perspective. Later, we present a mechanism for healing process of membrane after drilling, which could potentially provide guidance for improvement of injection method. More importantly, we perform coarse-grained molecular dynamics simulation to reveal the mechanism of membrane recovery in intra-cytoplasmic sperm injection.

bioengineering

GDCRNATools: an R/Bioconductor package for integrative analysis of lncRNA, miRNA, and mRNA data in GDC

The large-scale multidimensional omics data in the Genomic Data Commons (GDC) provides opportunities to investigate the crosstalk among different RNA species and their regulatory mechanisms in cancers. Easy-to-use bioinformatics pipelines are needed to facilitate such studies. We have developed a user-friendly R/Bioconductor package, named GDCRNATools, to facilitate downloading, organizing, and analyzing RNA data in GDC with an emphasis on deciphering the lncRNA-mRNA related competing endogenous RNAs (ceRNAs) regulatory network in cancers. Many widely used bioinformatics tools and databases are utilized in our package. Users can easily pack preferred downstream analysis pipelines or integrate their own pipelines into the workflow. Interactive shiny web apps built in GDCRNATools greatly improve visualization of results from the analysis.\n\nAvailabilityGDCRNATools is an R/Bioconductor package that is freely available at https://github.com/Jialab-UCR/GDCRNATools

bioinformatics

TahcoRoll: An Efficient Approach for Signature Profiling in Genomic Data through Variable-Length k-mers

k-mer profiling has been one of the trending approaches to analyze read data generated by high-throughput sequencing technologies. The tasks of k-mer profiling include, but are not limited to, counting the frequencies and determining the occurrences of short sequences in a dataset. The notion of k-mer has been extensively used to build de Bruijn graphs in genome or transcriptome assembly, which requires examining all possible k-mers presented in the dataset. Recently, an alternative way of profiling has been proposed, which constructs a set of representative k-mers as genomic markers and profiles their occurrences in the sequencing data. This technique has been applied in both transcript quantification through RNA-Seq and taxonomic classification of metagenomic reads. Most of these applications use a set of fixed-size k-mers since the majority of existing k-mer counters are inadequate to process genomic sequences with variable-length k-mers. However, choosing the appropriate k is challenging, as it varies for different applications. As a pioneer work to profile a set of variable-length k-mers, we propose TahcoRoll in order to enhance the Aho-Corasick algorithm. More specifically, we use one bit to represent each nucleotide, and integrate the rolling hash technique to construct an efficient in-memory data structure for this task. Using both synthetic and real datasets, results show that TahcoRoll outperforms existing approaches in either or both time and memory efficiency without using any disk space. In addition, compared to the most efficient state-of-the-art k-mer counters, such as KMC and MSBWT, TahcoRoll is the only approach that can process long read data from both PacBio and Oxford Nanopore on a commodity desktop computer. The source code of TahcoRoll is implemented in C++14, and available at https://github.com/chelseaju/TahcoRoll.git.

bioinformatics

Enhancer connectome in primary human cells reveals target genes of disease-associated DNA elements

The challenge of linking intergenic mutations to target genes has limited molecular understanding of diverse human diseases. Here, we show H3K27ac HiChIP generates high-resolution contact maps of active enhancers and target genes in rare primary human T cell subtypes and coronary artery smooth muscle cells. Differentiation of naive T cells to either T helper 17 cells or regulatory T cells create subtype-specific enhancer-promoter interactions, specifically at regions of shared DNA accessibility. These data provide a principled means of assigning molecular functions to autoimmune and cardiovascular disease risk variants, linking hundreds of noncoding variants to putative gene targets. Target genes identified with HiChIP are further supported by CRISPR interference and activation at linked enhancers, by the presence of expression quantitative trait loci, and by allele-specific enhancer loops in patient-derived primary cells. The majority of disease-associated enhancers contact genes beyond the nearest gene in the linear genome, leading to a four-fold increase of potential target genes for autoimmune and cardiovascular diseases.

genomics

Reason’s Enemy Is Not Emotion: Engagement of Cognitive Control Networks Explains Biases in Gain/Loss Framing

In the classic gain/loss framing effect, describing a gamble as a potential gain or loss biases people to make risk-averse or risk-seeking decisions, respectively. The canonical explanation for this effect is that frames differentially modulate emotional processes - which in turn leads to irrational choice behavior. Here, we evaluate the source of framing biases by integrating functional magnetic resonance imaging (fMRI) data from 143 human participants performing a gain/loss framing task with meta-analytic data from over 8000 neuroimaging studies. We found that activation during choices consistent with the framing effect were most correlated with activation associated with the resting or default brain, while activation during choices inconsistent with the framing effect most correlated with the task-engaged brain. Our findings argue against the common interpretation of gain/loss framing as a competition between emotion and control. Instead, our study indicates that this effect results from differential cognitive engagement across decision frames.\n\nSignificance StatementThe biases frequently exhibited by human decision-makers have often been attributed to the presence of emotion. Using a large fMRI sample and analysis of whole-brain networks defined with the meta-analytic tool Neurosynth, we find that neural activity during frame-biased decisions are more significantly associated with default behaviors (and the absence of executive control) than with emotion. These findings point to a role for neuroscience in shaping longstanding psychological theories in decision science.

neuroscience