Search bioRxivSearch

Biology subjects

Yang, J.

Publications and source records attributed to Yang, J..

At least 73 records · Page 4Linked to original sources

DrugPattern: a web-based tool for drug set enrichment analysis

Set enrichment analysis based methods (e.g. gene set enrichment analysis) have provided great helps in mining patterns in biomedical datasets, however, tools for inferring regular patterns in drug-related datasets are still limited. For the above purpose, here we developed a web-based tool, DrugPattern. DrugPattern first collected and curated 7019 drug sets, including indications, adverse reaction, targets, pathways etc. For a list of interested drugs, DrugPattern then evaluates the significance of the enrichment of these drugs in each of the 7019 drug sets. To validate DrugPattern, we applied it to predict the potential protective roles of oxidized low-density lipoprotein (oxLDL), a widely accepted deleterious factor for the body. We predicted that oxLDL has beneficial effects on some diseases, most of which were supported by literature except type 2 diabetes (T2D), in which oxLDL was previously believed to be a risk factor. Animal experiments further validated that oxLDL indeed has beneficial effects on T2D. These data confirmed the prediction accuracy of our approach and revealed unexpected protective roles for oxLDL in various diseases including T2D. This study provides a tool to infer regular patterns in biomedical datasets based on drug set enrichment analysis.

bioinformatics

GIC: A computational method for predicting the essentiality of long noncoding lncRNAs

Measuring the essentiality of genes is critically important in biology and medicine. Some bioinformatic methods have been developed for this issue but none of them can be applied to long noncoding RNAs (lncRNAs), one big class of biological molecules. Here we developed a computational method, GIC (Gene Importance Calculator), which can predict the essentiality of both protein-coding genes and lncRNAs based on RNA sequence information. For identifying the essentiality of protein-coding genes, GIC is competitive with well-established computational scores. More important, GIC showed a high performance for predicting the essentiality of lncRNAs. In an independent mouse lncRNA dataset, GIC achieved an exciting performance (AUC=0.918). In contrast, the traditional computational methods are not applicable to lncRNAs. As a public web server, GIC is freely available at http://www.cuilab.cn/gic/.

bioinformatics

A systems approach to the characterization and classification of T-cell responses

Types of T-cell responses are categorized on the basis of a limited number of molecular markers selected using a priori knowledge about T-cell immunobiology. We sought to develop a novel systems-based approach for the creation of an unbiased framework enabling assessment of antigenic-peptide specific T-cell responses in vitro. A meta-analysis of transcriptome data from PBMCs stimulated with a wide range of peptides identified patterns of gene regulation that provided an unbiased classification of types of antigen-specific responses. Further analysis yielded new insight about the molecular processes engaged following antigenic stimulation. This led for instance to the identification of transcription factors not previously studied in the context of T-cell differentiation. Taken together this profiling approach can serve as a basis for the unbiased characterization of antigen-specific responses and as a foundation for the development of novel systems-based immune profiling assays.

genomics

Increased anxiety and decreased sociability in adulthood following paternal deprivation involve oxytocin in the mPFC

Early adverse experiences often have devastating consequences on adult emotional and social behavior. However, whether paternal deprivation (PD) during the pre-weaning period affects brain and behavioral development remains unexplored in socially mandarin vole (Microtus mandarinus). We found that PD increased anxiety-like behavior and attenuated social preference in adult males and females; decreased prelimbic cortex OT-immunoreactive fibers and paraventricular nucleus OT positive neurons; reduced levels of medial prefrontal cortex (mPFC) OT receptor protein in females and OT receptor and V1a receptor protein in males. Intra-prelimbic cortical OT injections reversed anxiety-like behavior and social preferences affected by PD, whereas injections of OT and OT receptor antagonist blocked this reversal. These findings demonstrate that PD leads to increased anxiety-like behavior and attenuated social preferences with involvement of the mPFC OT system. The prelimbic cortex OT system may be an important target for the treatment of disorders related to early adverse experiences.

neuroscience

Causal associations between risk factors and common diseases inferred from GWAS summary data

Health risk factors such as body mass index (BMI), serum cholesterol and blood pressure are associated with many common diseases. It often remains unclear whether the risk factors are cause or consequence of disease, or whether the associations are the result of confounding. Genetic methods are useful to infer causality because genetic variants are present from birth and therefore unlikely to be confounded with environmental factors. We develop and apply a method (GSMR) that performs a multi-SNP Mendelian Randomization analysis using summary-level data from large genome-wide association studies (sample sizes of up to 405,072) to test the causal associations of BMI, waist-to-hip ratio, serum cholesterols, blood pressures, height and years of schooling (EduYears) with a range of common diseases. We identify a number of causal associations including a protective effect of LDL-cholesterol against type-2 diabetes (T2D) that might explain the side effects of statins on T2D, a protective effect of EduYears against Alzheimers disease, and bidirectional associations with opposite effects (e.g. higher BMI increases the risk of T2D but the effect T2D of BMI is negative). HDL-cholesterol has a significant risk effect on age-related macular degeneration, and the effect size remains significant accounting for the other risk factors. Our study develops powerful tools to integrate summary data from large studies to infer causality, and provides important candidates to be prioritized for further studies in medical research and for drug discovery.

genetics

Identification of 55,000 Replicated DNA Methylation QTL

DNA methylation plays an important role in the regulation of transcription. Genetic control of DNA methylation is a potential candidate for explaining the many identified SNP associations with disease that are not found in coding regions. We replicated 52,916 cis and 2,025 trans DNA methylation quantitative trait loci (mQTL) using methylation measured on Illumina HumanMethylation450 arrays in the Brisbane Systems Genetics Study (n=614 from 177 families) and the Lothian Birth Cohorts of 1921 and 1936 (combined n = 1366). The trans mQTL SNPs were found to be over-represented in 1Mbp subtelomeric regions, and on chromosomes 16 and 19. There was a significant increase in trans mQTL DNA methylation sites in upstream and 5 UTR regions. No association was observed between either the SNPs or DNA methylation sites of trans mQTL and telomere length. The genetic heritability of a number of complex traits and diseases was partitioned into components due to mQTL and the remainder of the genome. Significant enrichment was observed for height (p = 2.1x10-10), ulcerative colitis (p = 2x10-5), Crohns disease (p = 6x10-8) and coronary artery disease (p = 5.5x10-6) when compared to a random sample of SNPs with matched minor allele frequency, although this enrichment is explained by the genomic location of the mQTL SNPs.

genomics

Narrow-sense heritability estimation of complex traits using identity-by-descent information.

Heritability is a fundamental parameter in genetics. Traditional estimates based on family or twin studies can be biased due to shared environmental or non-additive genetic variance. Alternatively, those based on genotyped or imputed variants typically underestimate narrow-sense heritability contributed by rare or otherwise poorly-tagged causal variants. Identical-by-descent (IBD) segments of the genome share all variants between pairs of chromosomes except new mutations that have arisen since the last common ancestor. Therefore, relating phenotypic similarity to degree of IBD sharing among classically unrelated individuals is an appealing approach to estimating the near full additive genetic variance while avoiding biases that can occur when modeling close relatives. We applied an IBD-based approach (GREML-IBD) to estimate heritability in unrelated individuals using phenotypic simulation with thousands of whole genome sequences across a range of stratification, polygenicity levels, and the minor allele frequencies of causal variants (CVs). IBD-based heritability estimates were unbiased when using unrelated individuals, even for traits with extremely rare CVs, but stratification led to strong biases in IBD-based heritability estimates with poor precision. We used data on two traits in ~120,000 people from the UK Biobank to demonstrate that, depending on the trait and possible confounding environmental effects, GREML-IBD can be applied successfully to very large genetic datasets to infer the contribution of very rare variants lost using other methods. However, we observed apparent biases in this real data that were not predicted from our simulation, suggesting that more work may be required to understand factors that influence IBD-based estimates.

genetics

GEM: A manifold learning based framework for reconstructing spatial organizations of chromosomes

Decoding the spatial organizations of chromosomes has crucial implications for studying eukaryotic gene regulation. Recently, Chromosomal conformation capture based technologies, such as Hi-C, have been widely used to uncover the interaction frequencies of genomic loci in high-throughput and genome-wide manner and provide new insights into the folding of three-dimensional (3D) genome structure. In this paper, we develop a novel manifold learning framework, called GEM (Genomic organization reconstructor based on conformational Energy and Manifold learning), to elucidate the underlying 3D spatial organizations of chromosomes from Hi-C data. Unlike previous chromatin structure reconstruction methods, which explicitly assume specific relationships between Hi-C interaction frequencies and spatial distances between distal genomic loci, GEM is able to reconstruct an ensemble of chromatin conformations by directly embedding the neigh-boring affinities from Hi-C space into 3D Euclidean space based on a manifold learning strategy that considers both the fitness of Hi-C data and the biophysical feasibility of the modeled structures, which are measured by the conformational energy derived from our current biophysical knowledge about the 3D polymer model. Extensive validation tests on both simulated interaction frequency data and experimental Hi-C data of yeast and human demonstrated that GEM not only greatly outperformed other state-of-art modeling methods but also reconstructed accurate chromatin structures that agreed well with the hold-out or independent Hi-C data and sparse geometric restraints derived from the previous fluorescence in situ hybridization (FISH) studies. In addition, as GEM can generate accurate spatial organizations of chromosomes by integrating both experimentally-derived spatial contacts and conformational energy, we for the first time extended our modeling method to recover long-range genomic interactions that are missing from the original Hi-C data. All these results indicated that GEM can provide a physically and physiologically valid 3D representations of the organizations of chromosomes and thus serve as an effective and useful genome structure reconstructor.

bioinformatics

Widespread signatures of negative selection in the genetic architecture of human complex traits

Estimation of the joint distribution of effect size and minor allele frequency (MAF) for genetic variants is important for understanding the genetic basis of complex trait variation and can be used to detect signature of natural selection. We develop a Bayesian mixed linear model that simultaneously estimates SNP-based heritability, polygenicity (i.e. the proportion of SNPs with nonzero effects) and the relationship between effect size and MAF for complex traits in conventionally unrelated individuals using genome-wide SNP data. We apply the method to 28 complex traits in the UK Biobank data (N = 126,752), and show that on average across 28 traits, 6% of SNPs have nonzero effects, which in total explain 22% of phenotypic variance. We detect significant (p < 0.05/28 =1.8x10-3) signatures of natural selection for 23 out of 28 traits including reproductive, cardiovascular, and anthropometric traits, as well as educational attainment. We further apply the method to 27,869 gene expression traits (N = 1,748), and identify 30 genes that show significant (p < 2.3x10-6) evidence of natural selection. All the significant estimates of the relationship between effect size and MAF in either complex traits or gene expression traits are consistent with a model of negative selection, as confirmed by forward simulation. We conclude that natural selection acts pervasively on human complex traits shaping genetic variation in the form of negative selection.

genetics

Partitioning Phenotypic Variance Due To Parent-Of-Origin Effects Using Genomic Relatedness Matrices

Introduction Introduction Methods Statistical Methods Results Discussion References Parent-of-origin effects (POEs) describe the phenomenon in which the effects of alleles depend upon their parental origin. POEs imply that heterozygote individuals have phenotypes which are distributed differently depending upon which of their alleles were maternally and paternally transmitted (Guilmatre and Sharp 2012; Lawson et al. 2013). The extreme case of POEs is polar overdominance, where the two heterozygotes' phenotypes differ in distribution but the two homozygotes share the same distribution (Hoggart et al. 2014). Imprinting, a phenomenon in which one parent's allele is not expressed, is probably the most widely studied example of POE (Peters 2 ...

genetics

Inoculation Of Biocontrol Bacteria Alleviated Panax ginseng Replanting Problem

Replanting problem is a common and serious issue hindering the continuous cultivation of Panax plants. Changes in soil microbial community driven by plant species of different ages and developmental stages are speculated to cause this problem. Inoculation of microbial antagonists is proposed to alleviate replanting issues efficiently.\n\nHigh-throughput sequencing revealed that bacterial diversity evidently decreased, and fungal diversity markedly increased in soils of adult ginseng plants in the root growth stage. Relatively few beneficial microbe agents, such as Luteolibacter, Cytophagaceae, Luteibacter, Sphingomonas, Sphingomonadaceae, and Zygomycota, were observed. On the contrary, the relative abundance of harmful microorganism agents, namely, Brevundimonas, Enterobacteriaceae, Pandoraea, Cantharellales, Dendryphion, Fusarium, and Chytridiomycota, increased with pant age. Furthermore, Bacillus subtilis 50-1 was isolated and served as microbial antagonists against pathogenic Fusarium oxysporum of ginseng root-rot, and its biocontrol efficacy was 67.8% using a dual culture assay. The ginseng death rate and relative abundance of Fusarium decreased by 63.3% and 46.1%, respectively, after inoculation with 50-1 in replanting soils. Data revealed that changes in the diversity and composition of rhizospheric microbial communities driven by ginseng of different ages and developmental stages could cause microecological degradation. Biocontrol using microbial antagonists was an effective method for alleviating the replanting problem.\n\nHighlightChanges in rhizospheric microbial communities driven by ginseng plants 13 of different ages and developmental stages could cause microecological degradation. 14 Biocontrol using microbial antagonists effectively alleviated the replanting problem.

microbiology

Cellular Adaptation Through Fitness-Directed Transcriptional Tuning

Cells adapt to changes in their environment through transcriptional responses that are hard-coded in their regulatory networks. Such dedicated pathways, however, may be inadequate for adaptation to novel or extreme environments. We propose the existence of a fitness optimization mechanism that tunes the global transcriptional output of a genome to match arbitrary external conditions in the absence of dedicated gene-regulatory networks. We provide evidence for the proposed tuning mechanism in the adaptation of Saccharomyces cerevisiae to laboratory-engineered environments that are foreign to its native gene-regulatory network. We show that transcriptional tuning operates locally at individual gene promoters and its efficacy is modulated by genetic perturbations to chromatin modification machinery.

systems biology

Parallel Altitudinal Clines Reveal Adaptive Evolution Of Genome Size In Zea mays

While the vast majority of genome size variation in plants is due to differences in repetitive sequence, we know little about how selection acts on repeat content in natural populations. Here we investigate parallel changes in intraspecific genome size and repeat content of domesticated maize (Zea mays) landraces and their wild relative teosinte across altitudinal gradients in Mesoamerica and South America. We combine genotyping, low coverage whole-genome sequence data, and flow cytometry to test for evidence of selection on genome size and individual repeat abundance. We find that population structure alone cannot explain the observed variation, implying that clinal patterns of genome size are maintained by natural selection. Our modeling additionally provides evidence of selection on individual heterochromatic knob repeats, likely due to their large individual contribution to genome size. To better understand the phenotypes driving selection on genome size, we conducted a growth chamber experiment using a population of highland teosinte exhibiting extensive variation in genome size. We find weak support for a positive correlation between genome size and cell size, but stronger support for a negative correlation between genome size and the rate of cell production. Reanalyzing published data of cell counts in maize shoot apical meristems, we then identify a negative correlation between cell production rate and flowering time. Together, our data suggest a model in which variation in genome size is driven by natural selection on flowering time across altitudinal clines, connecting intraspecific variation in repetitive sequence to important differences in adaptive phenotypes.\n\nAuthor summaryGenome size in plants can vary by orders of magnitude, but this variation has long been considered to be of little to no functional consequence. Studying three independent adaptations to high altitude in Zea mays, we find that genome size experiences parallel pressures from natural selection, causing a linear reduction in genome size with increasing altitude. Though reductions in repetitive content are responsible for the genome size change, we find that only those individual loci contributing most to the variation in genome size are individually targeted by selection. To identify the phenotype influenced by genome size, we study how variation in genome size within a single teosinte population impacts leaf growth and cell division. We find that genome size variation correlates negatively with the rate of cell division, suggesting that individuals with larger genomes require longer to complete a mitotic cycle. Finally, we reanalyze data from maize inbreds to show that faster cell division is correlated with earlier flowering, connecting observed variation in genome size to an important adaptive phenotype.

evolutionary biology

Synthesis-Free PET Imaging Of Brown Adipose Tissue And TSPO Via Combination Of Disulfiram And 64CuCl2

PET imaging is a widely applicable but a very expensive technology. Strategies that can significantly reduce the high cost of PET imaging are highly desirable both for research and commercialization. On-site synthesis is one important contributor to the high cost. In this report, we demonstrated the feasibility of a synthesis-free method for PET imaging of brown adipose tissue (BAT) and translocator protein 18kDa (TSPO) via a combination of Disulfiram, an FDA approved drug for alcoholism, and 64CuCl2 (termed 64Cu-Dis). Our blocking studies, Western blot, and tissue histological imaging suggested that the observed BAT contrast was due to 64Cu-Dis binding to TSPO, which was further confirmed as a specific biomarker for BAT imaging using [18F]-F-DPA, a TSPO-specific PET tracer. Our studies, for the first time, demonstrated that TSPO could serve as a potential imaging biomarker for BAT. Furthermore, since imaging contrast obtained with both 64Cu-Dis and [18F]-F-DPA was not dependent on BAT activation, these agents could be used for reliably imaging BAT mass. Additional value of our synthesis-free approach could be applied to imaging TSPO in other tissues as it is an established biomarker of neuro-inflammation in activated microglia and plays a role in immune response, steroid synthesis, and apoptosis. Although here we applied 64Cu-Dis for a synthesis-free PET imaging of BAT, we believe that our strategy could be extended to other targets while significantly reducing the cost of PET imaging.\n\nSignificanceBrown adipose tissue (BAT) has been considered as \"good fat,\" and large-scale analysis has undoubtedly validated its clinical significance. BAT tightly correlates with body-mass index (BMI), suggesting that BAT bears clear significance for metabolic disorders such as obesity and diabetes. BAT imaging with [18F]-FDG, the most used method for visualizing BAT, primarily reflects BAT activation, but not BAT mass. A convenient imaging method that can consistently reflect BAT mass is still lacking. In this report, we demonstrated that BAT mass can be reliably imaged with a synthesis-free method using the combination of Disulfiram and 64CuCl2 (64Cu-Dis) via TSPO binding. We further demonstrated for the first time that TSPO is a specific imaging biomarker for BAT.

pharmacology and toxicology

Establishment In Culture Of Expanded Potential Stem Cells

Mouse embryonic stem cells are derived from in vitro explantation of blastocyst epiblasts1,2 and contribute to both the somatic lineage and germline when returned to the blastocyst3 but are normally excluded from the trophoblast lineage and primitive endoderm4-6. Here, we report that cultures of expanded potential stem cells (EPSCs) can be established from individual blastomeres, by direct conversion of mouse embryonic stem cells (ESCs) and by genetically reprogramming somatic cells. Remarkably, a single EPSC contributes to the embryo proper and placenta trophoblasts in chimeras. Critically, culturing EPSCs in a trophoblast stem cell (TSC) culture condition permits direct establishment of TSC lines without genetic modification. Molecular analyses including single cell RNA-seq reveal that EPSCs share cardinal pluripotency features with ESCs but have an enriched blastomere transcriptomic signature and a dynamic DNA methylome. These proof-of-concept results open up the possibility of establishing cultures of similar stem cells in other mammalian species.

developmental biology

RETA: An R Package For Whole Exome And Targeted Region Sequencing Data Analysis

Whole exome and targeted sequencing have been playing a major role in diagnoses of Mendelian diseases, but analysis of these data involves using many complicated tools and comprehensive understanding of the analysis results is difficult.Here, we report RETA, an R package to provide a one-stop analysis of these data and a comprehensive, interactive and easy-to-understand report with many advanced visualization features. It facilitates clinicians and scientists alike to better analyze and interpret this type of sequencing data for disease diagnoses.\n\nAvailability and implementationhttps://github.com/reta-s/reta/releases\n\nContactyangwl@hku.hk

genetics

Comparison of methods that use whole genome data to estimate the heritability and genetic architecture of complex traits.

Heritability, h2, is a foundational concept in genetics, critical to understanding the genetic basis of complex traits. Recently-developed methods that estimate heritability from genotyped SNPs, h2 SNP, explain substantially more genetic variance than genome-wide significant loci, but less than classical estimates from twins and families. However, h2SNP estimates have yet to be comprehensively compared under a range of genetic architectures, making it difficult to draw conclusions from sometimes conflicting published estimates. Here, we used thousands of real whole genome sequences to simulate realistic phenotypes under a variety of genetic architectures, including those from very rare causal variants. We compared the performance of ten methods across different types of genotypic data (commercial SNP array positions, whole genome sequence variants, and imputed variants) and under differing causal variant frequencies, levels of stratification, and relatedness thresholds. These results provide guidance in interpreting past results and choosing optimal approaches for future studies. We then chose two methods (GREML-MS and GREML-LDMS) that best estimated overall h2SNP and the causal variant frequency spectra to six phenotypes in the UK Biobank using imputed genome-wide variants. Our results suggest that as imputation reference panels become larger and more diverse, estimates of the frequency distribution of causal variants will become increasingly unbiased and the vast majority of trait narrow-sense heritability will be accounted for.

genetics

Allopatric divergence, local adaptation, and multiple Quaternary refugia in a long-lived tree (Quercus spinosa) from subtropical China

O_LIThe complex geography and climatic changes occurring in subtropical China during the Tertiary and Quaternary might have provided substantial opportunities for allopatric speciation. To gain further insight into these processes, we reconstruct the evolutionary history of Quercus spinosa, a common evergreen tree species mainly distributed in this area.\nC_LIO_LIForty-six populations were genotyped using four chloroplast DNA regions and 12 nuclear microsatellite loci to assess genetic structure and diversity, which was supplemented by divergence time and diversification rate analyses, environmental factor analysis, and ecological niche modeling of the species distributions in the past and at present.\nC_LIO_LIThe genetic data consistently identified two lineages: the western Eastern Himalaya-Hengduan Mountains lineage and the eastern Central-Eastern China lineage, mostly maintained by populations environmental adaptation. These lineages diverged through climate/orogeny-induced vicariance during the Neogene and remained separated thereafter. Genetic data strongly supported the multiple refugia (per se, interglacial refugia) or refugia within refugia hypotheses to explain Q. spinosa phylogeography in subtropical China.\nC_LIO_LIQ. spinosa population structure highlighted the importance of complex geography and climatic changes occurring in subtropical China during the Neogene in providing substantial opportunities for allopatric divergence.\nC_LI

evolutionary biology