Search bioRxivSearch

Biology subjects

Bao, Y.

Publications and source records attributed to Bao, Y..

8 recordsLinked to original sources

Properties of the epigenetic clock and age acceleration

BackgroundThe methylation status of numerous CpG sites in the human genome varies with age. The Horvath epigenetic clock used a wide variety of published DNA methylation data to produce an age prediction that has been widely used to, predict age in unknown samples, and draw conclusions about speed of ageing in various tissues, environments, and diseases. Despite its utility, there are a number of assumptions in the model that require examination. We explore the characteristics of the model in whole blood and multiple brain regions from older people, who are not well represented in the original training data, and in blood from a cross-sectional population study.\n\nResultsWe find that the model systematically underestimates age in tissues from older people. A decrease in slope of the predicted ages were observed at approximately 60 years, indicating that some loci in the model may change differently with age, and that age acceleration measures will themselves be age-dependent. This is seen most strongly in the cerebellum but is also present in other examined tissues, and is consistently observed in multiple datasets. An apparent association of Alzheimers disease with age acceleration disappears when age is used as a covariate. Association tests in the literature use a variety of methods for calculating age acceleration and often do not use age as a covariate. This is a potential cause of misleading findings.\n\nConclusionsAssociations of phenotypes with age acceleration should be evaluated cautiously, and chronological age should be included as a covariate in all analyses.

genomics

Characterization and identification of long non-coding RNAs based on feature relationship

The significance of long non-coding RNAs (lncRNAs) in many biological processes and diseases has gained intense interests over the past several years. However, computational identification of lncRNAs in a wide range of species remains challenging; it requires prior knowledge of well-established sequences and annotations or species-specific training data, but the reality is that only a limited number of species have high-quality sequences and annotations. Here we first characterize lncRNAs by contrast to protein-coding RNAs based on feature relationship and find that the feature relationship between ORF (open reading frame) length and GC content presents universally substantial divergence in lncRNAs and protein-coding RNAs, as observed in a broad variety of species. Based on the feature relationship, accordingly, we further present LGC, a novel algorithm for identifying lncRNAs that is able to accurately distinguish lncRNAs from protein-coding RNAs in a cross-species manner without any prior knowledge. As validated on large-scale empirical datasets, comparative results show that LGC outperforms existing algorithms by achieving higher accuracy, well-balanced sensitivity and specificity, and is robustly effective (>90% accuracy) in discriminating lncRNAs from protein-coding RNAs across diverse species that range from plants to mammals. To our knowledge, this study, for the first time, differentially characterizes lncRNAs and protein-coding RNAs based on feature relationship, which is further applied in computational identification of lncRNAs. Taken together, our study represents a significant advance in characterization and identification of lncRNAs and LGC thus bears broad potential utility for computational analysis of lncRNAs in a wide range of species.

bioinformatics

Leveraging DNA methylation quantitative trait loci to characterize the relationship between methylomic variation, gene expression and complex traits.

Characterizing the complex relationship between genetic, epigenetic and transcriptomic variation has the potential to increase understanding about the mechanisms underpinning health and disease phenotypes. In this study, we describe the most comprehensive analysis of common genetic variation on DNA methylation (DNAm) to date, using the Illumina EPIC array to profile samples from the UK Household Longitudinal study. We identified 12,689,548 significant DNA methylation quantitative trait loci (mQTL) associations (P < 6.52x10-14) occurring between 2,907,234 genetic variants and 93,268 DNAm sites, including a large number not identified using previous DNAm-profiling methods. We demonstrate the utility of these data for interpreting the functional consequences of common genetic variation associated with > 60 human traits, using Summary data-based Mendelian Randomization (SMR) to identify 1,662 pleiotropic associations between 36 complex traits and 1,246 DNAm sites. We also use SMR to characterize the relationship between DNAm and gene expression, identifying 6,798 pleiotropic associations between 5,420 DNAm sites and the transcription of 1,702 genes. Our mQTL database and SMR results are available via a searchable online database (http://www.epigenomicslab.com/online-data-resources/) as a resource to the research community.

genetics

Genome-wide study identifies 611 loci associated with risk tolerance and risky behaviors

Humans vary substantially in their willingness to take risks. In a combined sample of over one million individuals, we conducted genome-wide association studies (GWAS) of general risk tolerance, adventurousness, and risky behaviors in the driving, drinking, smoking, and sexual domains. We identified 611 approximately independent genetic loci associated with at least one of our phenotypes, including 124 with general risk tolerance. We report evidence of substantial shared genetic influences across general risk tolerance and risky behaviors: 72 of the 124 general risk tolerance loci contain a lead SNP for at least one of our other GWAS, and general risk tolerance is moderately to strongly genetically correlated ([Formula] to 0.50) with a range of risky behaviors. Bioinformatics analyses imply that genes near general-risk-tolerance-associated SNPs are highly expressed in brain tissues and point to a role for glutamatergic and GABAergic neurotransmission. We find no evidence of enrichment for genes previously hypothesized to relate to risk tolerance.

genetics

Sulforaphane modulates microRNA expression in colon cancer cells to implicate the regulation of oncogenes CDC25A, HMGA2 and MYC

Colorectal cancer is an increasingly important cause of morbidity and mortality, whose incidence is associated with dietary and lifestyle factors, particularly inversely so with the consumption of cruciferous vegetables. These vegetables contain glucosinolates, from the breakdown of which are derived isothiocyanates, such as sulforaphane. Sulforaphane is well-characterised for wide-ranging tumour-suppressive and chemoprotective activities in vitro, yet deeper elucidation of its biological interactions would aid in better realising its potential in chemoprevention and/or chemotherapy. There is evidence to suggest that sulforaphane modulates microRNA expression in the colon, thus implying the potential for microRNA modulation to play a role in the anti-cancer effects of sulforaphane. Therefore, the effects of sulforaphane on microRNA expression profiles in the colonic adenocarcinoma Caco-2 and non-cancerous colonic CCD-841 cell lines were investigated by small RNA cloning and deep sequencing, followed by Northern Blot validation experiments. Sulforaphane upregulated let-7f-5p and let-7g-5p expression at 24 h in Caco-2 cells, but not in CCD-841. Such treatment also downregulated miR-29b-3p in Caco-2. Dual luciferase assays with a let-7f-5p mimic and inhibitor confirmed the binding of the miRNA to predicted binding sites in the mRNA transcript 3-UTRs of cell division cycle 25A (CDC25A), high-mobility group AT-hook-2 (HMGA2) and MYC. Therefore, we hypothesize that let-7f-5p translationally represses CDC25A, HMGA2 and MYC, thereby playing a role in the tumour-suppressive effects of sulforaphane. The apparent selectivity of let-7f-5p induction towards tumour cells would be therapeutically desirable if applicable in vivo. MiR-29b-3p is predicted to target a number of tumour-suppressing genes, further investigation of which could be informative regarding the potential of sulforaphane to suppress tumour progression.

cancer biology

An efficient experiment design helps to identify differential expressed genes (DEGs) in RNA-seq data in studies of plant qualitative traits

In this study, we conducted comparative transcriptome analysis between homozygous dominant parent and heterozygous F1 hybrid with homozygous recessive parent in qualitative trait study of common wheat (Triticum aestivum L.). Two sets of near-isogenic lines (NILs) were used: one set of NILs carrying powdery mildew resistance and susceptible Pm2 alleles, the other set of NILs carrying different awn inhibition gene B1 alleles. The results demonstrated that 2,932 DEGs were identified between L031 (Pm2Pm2) and Chancellor (pm2pm2), while 1,494 DEGs presented between F1 hybrid (Pm2pm2) and Chancellor, the co-regulated DEGs were 1,028. For the wheat awn inhibition gene B1 test, 720 DEGs were identified between SN051-2 (B1B1) and SN051-1 (b1b1), and 231 DEGs were identified between F1 hybrid (B1b1) and SN051-1, the co-regulated DEGs were 180. Hierarchical clustering analysis of co-regulated DEGs showed that dominant parent and F1 hybrid were clustered as the nearest neighbors, while recessive parent showed an apparent departure. The results showed that the overlapping DEGs between dominant parent and F1 hybrid with recessive parent reduced the number of interested DEGs to only one-quarter (or one-third) of that between dominant and recessive parent, these overlapping loci could provide insights into molecular mechanisms that are affected by causal mutations.

plant biology

The Landscape Of Type VI Secretion Across Human Gut Microbiomes Reveals Its Role In Community Composition

While the composition of the human gut microbiome has been well defined, the forces governing its assembly are poorly understood. Recently, prominent members of this community from the order Bacteroidales were shown to possess the type VI secretion system (T6SS), which mediates contact-dependent antagonism between Gram-negative bacteria. However, the distribution of the T6SS in human gut microbiomes and its role have not yet been characterized. To address this challenge, we construct an extensive catalog of T6SS effector/immunity (E-I) genes from three genetic architectures (GA1-3) found in Bacteroidales genomes. We then use metagenomic analysis to assess the abundances of these genes across a large set of gut microbiome samples. We find that despite E-I diversity across reference strains, each individual microbiome harbors a limited set of E-I genes representing a single E-I genotype. Importantly, for GA1-2, these genotypes are not associated with a specific species, suggesting selection for compatibility. GA3, in contrast, is restricted to B. fragilis, and its low diversity reflects a single B. fragilis strain per sample. We further show that in infant microbiomes GA3 is enriched and B. fragilis strains are replaced over time, suggesting competition for dominance in developing microbiomes. Finally, we find a strong association between the presence of GA3 and increased abundance of Bacteroides, indicating that this system confers a selective advantage in vivo in Bacteroides rich ecosystems. Combined, our findings provide the first comprehensive characterization of the T6SS landscape in the human microbiome, implicating it in both intra- and inter-species interactions.

microbiology

Multivariate Genome-Wide and Integrated Transcriptome and Epigenome-Wide Analyses of the Well-being Spectrum.

Phenotypes related to well-being (life satisfaction, positive affect, neuroticism, and depressive symptoms), are genetically highly correlated (| rg | > .75). Multivariate genome-wide analyses (Nobs = 958,149) of these traits, collectively referred to as the well-being spectrum, reveals 63 significant independent signals, of which 29 were not previously identified. Transcriptome and epigenome analyses implicate variation in gene expression at 8 additional loci and CpG methylation at 6 additional loci in the etiology of well-being. We leverage an anatomically comprehensive survey of gene expression in the brain to annotate our findings, showing that SNPs within genes excessively expressed in the cortex and part of the hippocampal formation are enriched in their effect on well-being.

genetics