Search bioRxivSearch

Biology subjects

Lu, J.

Publications and source records attributed to Lu, J..

27 records · Page 2Linked to original sources

Removing Contaminants from Metagenomic Databases

Metagenomic sequencing of patient samples is a very promising method for the diagnosis of human infections. Sequencing has the ability to capture all the DNA or RNA from pathogenic organisms in a human sample. However, complete and accurate characterization of the sequence, including identification of any pathogens, depends on the availability and quality of genomes for comparison. Thousands of genomes are now available, and as these numbers grow, the power of metagenomic sequencing for diagnosis should increase. However, recent studies have exposed the presence of contamination in published genomes, which when used for diagnosis increases the risk of falsely identifying the wrong pathogen.\n\nTo address this problem, we have developed a bioinformatics system for eliminating contamination as well as low-complexity genomic sequences in the draft genomes of eukaryotic pathogens. We applied this software to identify and remove human, bacterial, archaeal, and viral sequences present in a comprehensive database of all sequenced eukaryotic pathogen genomes. We also removed low-complexity genomic sequences, another source of false positives. Using this pipeline, we have produced a database of \"clean\" eukaryotic pathogen genomes for use with bioinformatics classification and analysis tools. We demonstrate that when attempting to find eukaryotic pathogens in metagenomic samples, the new database provides better sensitivity than one using the original genomes while offering a dramatic reduction in false positives.

bioinformatics

Variation in Genome-wide NF-kappaB RELA Binding Sites upon Microbial Stimuli and Identification of a Virus Response Profile

NF-kB transcription factors are master regulators of the innate immune response. Activated downstream of pathogen recognition receptors, they regulate the expression of genes to help fighting infections as well as recruiting the adaptive immune system. NF-kB responds to a wide variety of signals, but the processes by which stimulus-specificity is attained remain unclear. Here, we characterized the response of one NF-kB member, RELA, to four stimuli mimicking infection in human nasopharyngeal epithelial cells. Comparing genome-wide RELA binding, we observed stimulus-specific sites, although most sites overlapped across stimuli. Specifically, the response to Poly I:C - mimicking viral dsRNA and signalling through TLR3 - induced a distinct RELA profile, binding in the vicinity of antiviral genes and correlating with corresponding gene expression. This group of binding sites was also enriched in Interferon Regulatory Factor (IRF) motifs and showed overlapping with IRFs binding sites. A novel NF-kB target, OASL was further validated and showed TLR3-specific activation. This work showed that some RELA DNA binding sites varied in activation response following different stimulations and that interaction with more specialized factors could help achieve this stimulus-specific activity. Our data provide a genomic view of regulated host response to different pathogen stimuli.

genomics

GDCRNATools: an R/Bioconductor package for integrative analysis of lncRNA, miRNA, and mRNA data in GDC

The large-scale multidimensional omics data in the Genomic Data Commons (GDC) provides opportunities to investigate the crosstalk among different RNA species and their regulatory mechanisms in cancers. Easy-to-use bioinformatics pipelines are needed to facilitate such studies. We have developed a user-friendly R/Bioconductor package, named GDCRNATools, to facilitate downloading, organizing, and analyzing RNA data in GDC with an emphasis on deciphering the lncRNA-mRNA related competing endogenous RNAs (ceRNAs) regulatory network in cancers. Many widely used bioinformatics tools and databases are utilized in our package. Users can easily pack preferred downstream analysis pipelines or integrate their own pipelines into the workflow. Interactive shiny web apps built in GDCRNATools greatly improve visualization of results from the analysis.\n\nAvailabilityGDCRNATools is an R/Bioconductor package that is freely available at https://github.com/Jialab-UCR/GDCRNATools

bioinformatics

The Moran coalescent in a discrete one-dimensional spatial model

Among many organisms, offspring are constrained to occur at sites adjacent to their parents. This applies to plants and animals with limited dispersal ability, to colonies of microbes in biofilms, and to other genetically heterogeneous aggregates of cells, such as cancerous tumors. The spatial structure of such populations leads to greater relatedness among proximate individuals while increasing the genetic divergence between distant individuals. In this study, we analyze a Moran coa-lescent in a one-dimensional spatial model where a randomly selected individual dies and is replaced by the progeny of an adjacent neighbor in every generation. We derive a recursive system of equations using the spatial distance among haplotypes as a state variable to compute coalescent probabilities and coalescent times. The coalescent probabilities near the branch termini are smaller than in the unstructured Moran model (except for t = 1, where they are equal), corresponding to longer branch lengths and greater expected pairwise coalescent times. The lower terminal coalescent probabilities result from a spatial separation of lineages, i.e. a coalescent event between a haplotype and its neighbor in one spatial direction at time t cannot co-occur with a coalescent event with a haplotype in the opposite direction at t + 1. The concomitant increased pairwise genetic distance among randomly sampled haplotypes in spatially constrained populations could lead to incorrect inferences of recent diversifying selection or of population bottlenecks when analyzed using an unconstrained coalescent model as a null hypothesis.

genetics

SID-1 domains important for dsRNA import in C. elegans

In the nematode Caenorhabditis elegans, RNA interference (RNAi) triggered by double-stranded RNA (dsRNA) spreads systemically to cause gene silencing throughout the organism and its progeny. We confirm that Caenorhabditis nematode SID-1 orthologs have dsRNA transport activity and demonstrate that the SID-1 paralog CHUP-1 does not transport dsRNA. Sequence comparison of these similar proteins, in conjunction with analysis of loss-of-function missense alleles identifies several conserved 2-7 amino acid microdomains within the extracellular domain that are important for dsRNA transport. Among these missense alleles, we identify and characterize a sid-1 allele, qt95, which causes tissue-specific silencing defects most easily explained as a systemic RNAi export defect. However, we conclude from genetic and biochemical analyses that sid-1(qt95) disrupts only import and speculate that the apparent export defect is caused by the cumulative effect of sequentially impaired dsRNA import steps.Thus, consistent with previous studies, we fail to detect a requirement for sid-1 in dsRNA export, but demonstrate for the first time that SID-1 functions in the intestine to support environmental RNAi.

genetics

Molecular evolution, diversity and adaptation of H7N9 influenza A viruses in China

A novel H7N9 avian influenza virus has caused five human epidemics in China since 2013. The substantial increase in prevalence and the emergence of antigenically divergent or highly pathogenic (HP) H7N9 strains during the current outbreak raises concerns about the epizootic-potential of these viruses. Here, we investigate the evolution and adaptation of H7N9 by combining publicly available data with newly generated virus sequences isolated in Guangdong between 2015-2017. Phylogenetic analyses show that currently-circulating H7N9 viruses belong to distinct lineages with differing spatial distributions. Using ancestral sequence reconstruction and structural modelling we have identified parallel amino-acid changes on multiple separate lineages. Furthermore, we infer mutations in HA primarily occur at sites involved in receptor-recognition and/or antigenicity. We also identify seven new HP strains, which likely emerged from viruses circulating in eastern Guangdong around March 2016 and is further associated with a high rate of adaptive molecular evolution.

evolutionary biology

Driver Pattern Identification Over The Gene Co-Expression Of Drug Response In Ovarian Cancer By Integrating High Throughput Genomics Data

The multiple types of high throughput genomics data create a potential opportunity to identify driver pattern in ovarian cancer, which will acquire some novel and clinical biomarkers for appropriate diagnosis and treatment to cancer patients. However, it is a great challenging work to integrate omics data, including somatic mutations, Copy Number Variations (CNVs) and gene expression profiles, to distinguish interactions and regulations which are hidden in drug response dataset of ovarian cancer. To distinguish the candidate driver genes and the corresponding driving pattern for resistant and sensitive tumor from the heterogeneous data, we combined gene co-expression modules and mutation modulators and proposed the identification driver patterns method. Firstly, co-expression network analysis is applied to explore gene modules for gene expression profiles via weighted correlation network analysis (WGCNA). Secondly, mutation matrix is generated by integrating the CNVs and somatic mutations, and a mutation network is constructed from this mutation matrix. The candidate modulators are selected from the significant genes by clustering the vertex of the mutation network. At last, regression tree model is utilized for module networks learning in which the achieved gene modules and candidate modulators are trained for the driving pattern identification and modulator regulatory exploring. Many of the candidate modulators identified are known to be involved in biological meaningful processes associated with ovarian cancer, which can be regard as potential driver genes, such as CCL11, CCL16, CCL18, CCL23, CCL8, CCL5, APOB, BRCA1, SLC18A1, FGF22, GADD45B, GNA15, GNA11 and so on, which can help to facilitate the discovery of biomarkers, molecular diagnostics, and drug discovery.

bioinformatics

Selective Inhibitory Control Of Pyramidal Neuron Ensembles And Cortical Subnetworks By Chandelier Cells

The neocortex comprises multiple information processing streams mediated by subsets of glutamatergic pyramidal cells (PCs) that receive diverse inputs and project to distinct targets. How GABAergic interneurons regulate the segregation and communication among intermingled PC subsets that contribute to separate brain networks remains unclear. Here we demonstrate that a subset of GABAergic chandelier cells (ChCs) in the prelimbic cortex (PL), which innervate PCs at spike initiation site, selectively control PCs projecting to the basolateral amygdala (BLAPC) compared to those projecting to contralateral cortex (ccPC). These ChCs in turn receive preferential input from local and contralateral CCPCs as opposed to BLAPCs and BLA neurons (the PL-BLA network). Accordingly, optogenetic activation of ChCs rapidly suppresses BLAPCs and BLA activity in freely behaving mice. Thus, the exquisite connectivity of ChCs not only mediates directional inhibition between local PC ensembles but may also shape communication hierarchies between global networks.

neuroscience

Estimation of Pairwise Genetic Distances Under Independent Sampling of Segregating Sites vs. Haplotype Sampling

Genetic distance is a standard measure of variation in populations. When sequencing genomes individually, genetic distances are computed over all pairs of multilocus haplotypes in a sample. However, when next-generation sequencing methods obtain reads from heterogeneous assemblages of genomes (e.g. for microbial samples in a biofilm or cells from a tumor), individual reads are often drawn from different genomes. This means that pairwise genetic distances are calculated across independently sampled sites rather than across haplotype pairs. In this paper, we show that while the expected pairwise distance under whole haplotype sampling (WHS) is the same as with independent locus sampling (ILS), the sample variances of pairwise distance differ and depend on the direction and magnitude of linkage disequilibrium (LD) among polymorphic sites. We derive a weighted LD value that, when positive, predicts higher sample variance in estimated genetic distance for WHS. Weighted LD is positive when on average, the most common alleles at two loci are in positive LD. Using individual-based simulations of an infinite sites model under Fisher-Wright genetic drift, variances of estimated genetic distance are found to be almost always higher under WHS than under ILS, suggesting a reduction in estimation error when sites are sampled independently. We apply these results to haplotype frequencies from a lung cancer tumor to compute weighted LD and the variances in estimated genetic distance under ILS vs. WHS, and find that the the relative magnitudes of variances under WHS vs. ILS are sensitive to sampled allele frequencies.

genetics