Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Shared activity patterns arising at genetic susceptibility loci reveal underlying genomic and cellular architecture of human disease.

Genetic variants underlying complex traits, including disease susceptibility, are enriched within the transcriptional regulatory elements, promoters and enhancers. There is emerging evidence that regulatory elements associated with particular traits or diseases share patterns of transcriptional regulation. Accordingly, shared transcriptional regulation (coexpression) may help prioritise loci associated with a given trait, and help to identify the biological processes underlying it. Using cap analysis of gene expression (CAGE) profiles of promoter and enhancer-derived RNAs across 1824 human samples, we have quantified coexpression of RNAs originating from trait-associated regulatory regions using a novel analytical method (network density analysis; NDA). For most traits studied, sequence variants in regulatory regions were linked to tightly coexpressed networks that are likely to share important functional characteristics. These networks implicate particular cell types and tissues in disease pathogenesis; for example, variants associated with ulcerative colitis are linked to expression in gut tissue, whereas Crohns disease variants are restricted to immune cells. We show that this coexpression signal provides additional independent information for fine mapping likely causative variants. This approach identifies additional genetic variants associated with specific traits, including an association between the regulation of the OCT1 cation transporter and genetic variants underlying circulating cholesterol levels. This approach enables a deeper biological understanding of the causal basis of complex traits.\n\nONE SENTENCE SUMMARYWe discover that variants associated with a specific disease share expression profiles across tissues and cell types, enabling fine mapping and identification of new disease-associated variants, illuminating key cell types involved in disease pathogenesis.

genomics

Infectious Disease Dynamics Inferred from Genetic Data via Sequential Monte Carlo

Genetic sequences from pathogens can provide information about infectious disease dynamics that may supplement or replace information from other epidemiological observations. Currently available methods first estimate phylogenetic trees from sequence data, then estimate a transmission model conditional on these phylogenies. Outside limited classes of models, existing methods are unable to enforce logical consistency between the model of transmission and that underlying the phylogenetic reconstruction. Such conflicts in assumptions can lead to bias in the resulting inferences. Here, we develop a general, statistically efficient, plug-and-play method to jointly estimate both disease transmission and phylogeny using genetic data and, if desired, other epidemiological observations. This method explicitly connects the model of transmission and the model of phylogeny so as to avoid the aforementioned inconsistency. We demonstrate the feasibility of our approach through simulation and apply it to estimate stage-specific infectiousness in a subepidemic of HIV in Detroit, Michigan. In a supplement, we prove that our approach is a valid sequential Monte Carlo algorithm. While we focus on how these methods may be applied to population-level models of infectious disease, their scope is more general. These methods may be applied in other biological systems where one seeks to infer population dynamics from genetic sequences, and they may also find application for evolutionary models with phenotypic rather than genotypic data.

epidemiology

Preserving microsatellites? Conservation genetics of the giant Galapagos tortoise.

This preprint has been reviewed and recommended by Peer Community In Evolutionary Biology (http://dx.doi.org/10.24072/pci.evolbiol.100031).\n\nConservation policy in the giant Galapagos tortoise, an iconic endangered animal, has been assisted by genetic markers for [~]15 years: a dozen loci have been used to delineate thirteen (sub)species, between which hybridization is prevented. Here, comparative reanalysis of a previously published NGS data set reveals a conflict with traditional markers. Genetic diversity and population substructure in the giant Galapagos tortoise are found to be particularly low, questioning the genetic relevance of current conservation practices. Further examination of giant Galapagos tortoise population genomics is critically needed.

evolutionary biology

Spontaneous mutations and transmission distortions of genic copy number variants shape the standing genetic variation in Picea glauca

Copy number variations (CNVs) are large genetic variations detected among the individuals of every multicellular organism examined so far. These variations are believed to play an important role in the evolution and adaptation of species. In plants, little is known about the characteristics of CNVs, particularly regarding the rates at which they are generated and the mechanics of their transmission from a generation to the next. Here, we used SNP-array raw intensity data for 55 two-generations families (3663 individuals) to scan the gene space of the conifer tree Picea glauca (Moench) Voss for CNVs. We were particularly interested in the abundance, inheritance, spontaneous mutation rate spectrum and the evolutionary consequences they may have on the standing genetic variation of white spruce. Our findings show that CNVs affect a small proportion of the gene space and are predominantly copy number losses. CNVs were either inherited or generated through de novo events. De novo CNVs present high rates of spontaneous mutations that vary for different genes and alleles and are correlated with gene expression levels. Most of the inherited CNVs (70%) are transmitted from the parents in violation of Mendelian expectations. These transmission distortions can cause considerable frequency changes between generations and be dependent on whether the heterozygote parents contribute as male or female. Transmission distortions were also influenced by the partner genotype and the parents genetic background. This study provides new insights into the effects of different evolutionary forces on copy number variations based on the analysis of a perennial plant.

genomics

Genetic equidistance at the nucleotide level

The genetic equidistance phenomenon was first discovered in 1963 by Margoliash and shows complex taxa to be all approximately equidistant to a less complex species in amino acid percentage identity. The result has been mis-interpretated by the ad hoc universal molecular clock hypothesis, and the much overlooked mystery was finally solved by the maximum genetic diversity hypothesis (MGD). Here, we studied 15 proteomes and their coding DNA sequences (CDS) to see if the equidistance phenomenon also holds at the CDS level. We performed DNA alignments for a total of 5 groups with 3 proteomes per group and found that in all cases the outgroup taxon was equidistant to the two more complex taxa species at the DNA level. Also, when two sister taxa (snake and bird) were compared to human as the outgroup, the more complex taxon bird was closer to human, confirming species complexity rather than time to be the primary determinant of MGD. Finally, we found the fraction of overlap sites where coincident substitutions occur to be inversely correlated with CDS conservation, indicating saturation to be more common in less conserved DNAs. These results establish the genetic equidistance phenomenon to be universal at the DNA level and provide additional evidence for the MGD theory.

evolutionary biology

A maternal-effect genetic incompatibility in Caenorhabditis elegans

Selfish genetic elements spread in natural populations and have an important role in genome evolution. We discovered a selfish element causing a genetic incompatibility between strains of the nematode Caenorhabditis elegans. The element is made up of sup-35, a maternal-effect toxin that kills developing embryos, and pha-1, its zygotically expressed antidote. pha-1 has long been considered essential for pharynx development based on its mutant phenotype, but this phenotype in fact arises from a loss of suppression of sup-35 toxicity. Inactive copies of the sup-35/pha-1 element show high sequence divergence from active copies, and phylogenetic reconstruction suggests that they represent ancestral stages in the evolution of the element. Our results suggest that other essential genes identified by genetic screens may turn out to be components of selfish elements.

evolutionary biology

MOSAIC: a chemical-genetic interaction data repository and web resource for exploring chemical modes of action

SummaryChemical-genomic approaches that map interactions between small molecules and genetic perturbations offer a promising strategy for functional annotation of uncharacterized bioactive compounds. We recently developed a new high-throughput platform for mapping chemical-genetic (CG) interactions in yeast that can be scaled to screen large compound collections, and we applied this system to generate CG interaction profiles for more than 13,000 compounds. When integrated with the existing global yeast genetic interaction network, CG interaction profiles can enable mode-of-action prediction for previously uncharacterized compounds as well as discover unexpected secondary effects for known drugs. To facilitate future analysis of these valuable data, we developed a public database and web interface named MOSAIC. The website provides a convenient interface for querying compounds, bioprocesses (GO terms), and genes for CG information including direct CG interactions, bioprocesses, and gene-level target predictions. MOSAIC also provides access to chemical structure information of screened molecules, chemical-genomic profiles, and the ability to search for compounds sharing structural and functional similarity. This resource will be of interest to chemical biologists for discovering new small molecule probes with specific modes-of-action as well as computational biologists interested in analyzing CG interaction networks.\n\nAvailabilityMOSAIC is available at http://mosaic.cs.umn.edu.\n\nContactchadm@umn.edu, charlie.boone@utoronto.ca, yoshidam@riken.jp, or hisyo@riken.jp

bioinformatics

Evolution and clinical impact of genetic epistasis within EGFR-mutant lung cancers

Introductory paragraph Introductory paragraph Main text Methods Author Contributions Author Information References The current understanding of tumorigenesis is largely centered on a monogenic driver oncogene model. This paradigm is incompatible with the prevailing clinical experience in most solid malignancies: monotherapy with a drug directed against an individual oncogenic driver typically results in incomplete clinical responses and eventual tumor progression1-7. By profiling the somatic genetic alterations present in over 2,000 cases of lung cancer, the leading cause of cancer mortality worldwide8,9, we show that combinations of functional genetic alterations, i.e. genetic collectives dominate the landscape of ...

cancer biology

Intra- and Inter-individual genetic variation in human ribosomal RNAs

The ribosome is an ancient RNA-protein complex essential for translating DNA to protein. At its core are the ribosomal RNAs (rRNAs), the most abundant RNA in the cell. To support high levels of transcription, repetitive arrays of ribosomal DNA (rDNA) are necessary and the long-standing hypothesis is that they undergo sequence homogenization towards rDNA uniformity.\n\nHere I present evidence of the rich genetic diversity in human rDNA, both within and between individuals. Using state-of-the-art genome sequencing data revealed an average of 192.7 intra-individual variants, including some deeply penetrating the rDNA copies, such as the bi-allelically expressed 28S.59A>G. From 104 diverse genomes, 947 high-confidence variants were identified and unmask a hidden genetic diversity of humans.\n\nThese findings support the emerging concept that ribosomes are heterogeneous within cells and extends the heterogeneity into the realm of population genetics. Fundamentally, do our ribosomal variants determine how our cells interpret the genome?

genomics

Multiplexing droplet-based single cell RNA-sequencing using natural genetic barcodes

Droplet-based single-cell RNA-sequencing (dscRNA-seq) has enabled rapid, massively parallel profiling of transcriptomes from tens of thousands of cells. Multiplexing samples for single cell capture and library preparation in dscRNA-seq would enable cost-effective designs of differential expression and genetic studies while avoiding technical batch effects, but its implementation remains challenging. Here, we introduce an in-silico algorithm demuxlet that harnesses natural genetic variation to discover the sample identity of each cell and identify droplets containing two cells. These capabilities enable multiplexed dscRNA-seq experiments where cells from unrelated individuals are pooled and captured at higher throughput than standard workflows. To demonstrate the performance of demuxlet, we sequenced 3 pools of peripheral blood mononuclear cells (PBMCs) from 8 lupus patients. Given genotyping data for each individual, demuxlet correctly recovered the sample identity of > 99% of singlets, and identified doublets at rates consistent with previous estimates. In PBMCs, we demonstrate the utility of multiplexed dscRNA-seq in two applications: characterizing cell type specificity and inter-individual variability of cytokine response from 8 lupus patients and mapping genetic variants associated with cell type specific gene expression from 23 donors. Demuxlet is fast, accurate, scalable and could be extended to other single cell datasets that incorporate natural or synthetic DNA barcodes.

bioinformatics

Investigating The Genetic Regulation Of The Expression Of 63 Lipid Metabolism Genes In The Pig Skeletal Muscle

Despite their potential involvement in the determination of fatness phenotypes, a comprehensive and systematic view about the genetic regulation of lipid metabolism genes is still lacking in pigs. Herewith, we have used a dataset of 104 pigs, with available genotypes for 62,163 single nucleotide polymorphisms and microarray gene expression measurements in the gluteus medius muscle, to investigate the genetic regulation of 63 genes with crucial roles in the uptake, transport, synthesis and catabolism of lipids. By performing an eQTL scan with the GEMMA software, we have detected 12 cis- and 18 trans-eQTL modulating the expression of 19 loci. Genes regulated by eQTL had a variety of functions such as the {beta}-oxidation of fatty acids, lipid biosynthesis and lipolysis, fatty acid activation and desaturation, lipoprotein uptake, apolipoprotein assembly and cholesterol trafficking. These data provide a first picture about the genetic regulation of loci involved in porcine lipid metabolism.

genomics

Like Sugar in Milk: Reconstructing the genetic history of the Parsi population

BackgroundThe Parsis, one of the smallest religious community in the world, reside in South Asia. Previous genetic studies on them, although based on low resolution markers, reported both Iranian and Indian ancestries. To understand the population structure and demographic history of this group in more detail, we analyzed Indian and Pakistani Parsi populations using high-resolution autosomal and uniparental (Y-chromosomal and mitochondrial DNA) markers. Additionally, we also assayed 108 mitochondrial DNA markers among 21 ancient Parsi DNA samples excavated from Sanjan, in present day Gujarat, the place of their original settlement in India.\n\nResultsOur extensive analyses indicated that among present-day populations, the Parsis are genetically closest to Middle Eastern (Iranian and the Caucasus) populations rather than their South Asian neighbors. They also share the highest number of haplotypes with present-day Iranians and we estimate that the admixture of the Parsis with Indian populations occurred [~]1,200 years ago. Enriched homozygosity in the Parsi reflects their recent isolation and inbreeding. We also observed 48% South-Asian-specific mitochondrial lineages among the ancient samples, which might have resulted from the assimilation of local females during the initial settlement.\n\nConclusionsWe show that the Parsis are genetically closest to the Neolithic Iranians, followed by present-day Middle Eastern populations rather than those in South Asia and provide evidence of sex-specific admixture from South Asians to the Parsis. Our results are consistent with the historically-recorded migration of the Parsi populations to South Asia in the 7thcentury and in agreement with their assimilation into the Indian sub-continents population and cultural milieu \"like sugar in milk\". Moreover, in a wider context our results suggest a major demographic transition in West Asia due to Islamic-conquest.

evolutionary biology

Gestational Age At Birth And Risk Of Intellectual Disability Without A Common Genetic Cause: Findings From The Stockholm Youth Cohort

BackgroundPreterm birth is linked to intellectual disability and there is evidence to suggest post-term birth may also incur risk. However, these associations have not yet been investigated in the absence of common genetic causes of intellectual disability (where risk associated with late delivery may be preventable) or with methods allowing stronger causal inference from non-experimental data. We aimed to examine risk of intellectual disability without a common genetic cause across the entire range of gestation, using a matched-sibling design to account for unmeasured confounding by shared familial factors.\n\nMethods and FindingsWe conducted a population-based retrospective study using data from the Stockholm Youth Cohort (n=499,621) and examined associations in a nested cohort of matched siblings (n=8,034). Children born at non-optimal gestational duration (before/after 40 weeks 3 days) were at greater risk of intellectual disability. Risk was greatest among those born extremely early (adjusted OR24 weeks=14.54 [95% CI 11.46-18.44]), lessening with advancing gestational age toward term (aOR32 weeks=3.59 [3.22-4.01]; aOR37 weeks=1.50 [1.38-1.63]); aOR38 weeks=1.26 [1.16-1.37]; aOR39 weeks=1.10 [1.04-1.17]) and increasing with advancing gestational age post-term (aOR42 weeks=1.16 [1.08-1.25]; aOR42 weeks=1.41 [1.21-1.64]; aOR44 weeks=1.71 [1.34-2.18]; aOR45 weeks=2.07 [1.47-2.92]). Associations persisted in a nested cohort of matched outcome-discordant siblings suggesting they were robust against confounding from shared genetic or environmental traits, although there may have been residual confounding by unobserved non-shared characteristics. Risk of intellectual disability was greatest among children showing evidence of fetal growth restriction, especially when birth occurred before or after term.\n\nConclusionsBirth at non-optimal gestational duration may be linked causally with greater risk of intellectual disability. The mechanisms underlying these associations need to be elucidated as they will be relevant to clinical practice concerning elective delivery within the term period and the mitigation of risk in children who are born post-term.

epidemiology

With Our Powers Combined: Integrating Behavioral And Genetic Data To Estimate Mating Success And Sexual Selection

The analysis of sexual selection classically relies on the regression of individual phenotypes against the marginal sums of a males x females matrix of pairwise reproductive success, assessed by genetic parentage analysis. When the matrix is binarized, the marginal sums give the individual mating success. Because such analysis treats male and female mating/reproductive success independently, it ignores that the success of a male x female sexual interaction can be attributable to the phenotype of both individuals. Also, because it is based on genetic data only, it is oblivious to costly yet unproductive matings, which may be documented by behavioral observations. To solve these problems, we propose a statistical model which combines matrices of offspring numbers and behavioral observations. It models reproduction on each mating occasion of a mating season as three stochastic and interdependent pairwise processes, each potentially affected by the phenotype of both individuals and by random individual effect: encounter (Bernoulli), concomitant gamete emission (Bernoulli), and offspring production (Poisson). Applied to data from a mating experiment on brown trout, the model yielded different results from the classical regression analysis, with only a limited effect of male body size on the probability of gamete release and a negative effect of female body size on the probability of encounter and gamete release. Because the general structure of the model can be adapted to other partitioning of the reproductive process, it can be used for a variety of biological systems where behavioral and genetic data are available.

evolutionary biology

Coessentiality And Cofunctionality: A Network Approach To Learning Genetic Vulnerabilities From Cancer Cell Line Fitness Screens

Genetic interaction networks are a powerful approach for functional genomics, and the synthetic lethal interactions that comprise these networks offer a compelling strategy for identifying candidate cancer targets. As the number of published shRNA and CRISPR perturbation screens in cancer cell lines expands, there is an opportunity for integrative analysis that goes further than pairwise synthetic lethality and discovers genetic vulnerabilities of related sets of cell lines. We re-analyze over 100 high-quality, genome-scale shRNA screens in human cancer cell lines and derive a quantitative fitness score for each gene that accurately reflects genotype-specific gene essentiality. We identify pairs of genes with correlated essentiality profiles and merge them into a cancer coessentiality network, where shared patterns of genetic vulnerability in cell lines give rise to clusters of functionally related genes in the network. Network clustering discriminates among all three defined subtypes of breast cancer cell lines (basal, luminal, and Her2-amplified), and further identifies novel subsets of Her2+ and ovarian cancer cells. We demonstrate the utility of the network as a platform for both hypothesis-driven and data-driven discovery of context-specific essential genes and their associated biomarkers.

genomics

Inferring Genetic Interactions From Comparative Fitness Data

AO_SCPLOWBSTRACTC_SCPLOWDarwinian fitness is a central concept in evolutionary biology. In practice, however, it is hardly possible to measure fitness for all genotypes in a natural population. Here, we present quantitative tools to make inferences about epistatic gene interactions when the fitness landscape is only incompletely determined due to imprecise measurements or missing observations. We demonstrate that genetic interactions can often be inferred from fitness rank orders, where all genotypes are ordered according to fitness, and even from partial fitness orders. We provide a complete characterization of rank orders that imply higher order epistasis. Our theory applies to all common types of gene interactions and facilitates comprehensive investigations of diverse genetic interactions. We analyzed various genetic systems comprising HIV-1, the malaria-causing parasite Plasmodium vivax, the fungus Aspergillus niger, and the TEM-family of {beta}-lactamase associated with antibiotic resistance. For all systems, our approach revealed higher order interactions among mutations.

evolutionary biology

Mitochondrial Genetic Effects On Reproductive Success: Signatures Of Positive Intra-Sexual, But Negative Inter-Sexual Pleiotropy

Mitochondria contain their own DNA, and numerous studies have reported that genetic variation in this (mt)DNA sequence modifies the expression of life-history phenotypes. Maternal inheritance of mitochondria adds a layer of complexity to trajectories of mtDNA evolution, because theory predicts the accumulation of mtDNA mutations that are male-biased in effect. While it is clear that mitochondrial genomes routinely harbor genetic variation that affects components of reproductive performance, the extent to which this variation is sex-biased, or even sex-specific in effect, remains elusive. This is because nearly all previous studies have failed to examine mitochondrial genetic effects on both male and female reproductive performance within the one-and-the-same study. Here, we show that variation across naturally-occurring mitochondrial haplotypes affects components of reproductive success in both sexes, in Drosophila melanogaster. However, while we uncovered evidence for positive pleiotropy, across haplotypes, in effects on separate components of reproductive success when measured within the same sex, such patterns were not evident across sexes. Rather, we found a pattern of sexual antagonism across haplotypes on some reproductive parameters. This suggests the pool of polymorphisms that delineate global mtDNA haplotypes is likely to have been partly shaped by maternal transmission of mtDNA and its evolutionary consequences.

evolutionary biology

Genetic Overlap Of Depression With Cardiometabolic Diseases And Implications For Drug Repurposing For Comorbidities

Numerous studies have suggested associations between depression and cardiometabolic abnormalities or diseases, such as coronary artery disease and type 2 diabetes. However, little is known about the mechanism underlying this comorbidity, and whether the relationship differs by depression subtypes. Using the polygenic risk score (PRS) approach and linkage disequilibrium (LD) score regression, we investigated the genetic overlap of various depression-related phenotypes with a comprehensive panel of 20 cardiometabolic traits. GWAS results for major depressive disorder (MDD) were taken from the PGC and CONVERGE studies, with the latter focusing on severe melancholic depression. GWAS results on general depressive symptoms (DS) and neuroticism were also included. We also identified the shared genetic variants and inferred enriched pathways. In addition, we looked for drugs over-represented among the top shared genes, with an aim to finding repositioning opportunities for comorbidities.\n\nWe found significant polygenic sharing between MDD, DS and neuroticism with various cardiometabolic traits. In general, positive polygenic associations with CV risks were observed for most depression phenotypes except MDD-CONVERGE. Counterintuitively, PRS representing severe melancholic depression was associated with reduced CV risks. Enrichment analyses of shared SNPs revealed many interesting pathways, such as those related to inflammation, that underlie the comorbidity of depressive and cardiometabolic traits. Using a gene-set analysis approach, we also revealed a number of repositioning candidates, some of which were supported by prior studies, such as bupropion and glutathione. Our study highlights shared genetic bases of depression with cardiometabolic traits, and suggests the associations vary by depression subtypes. To our knowledge, this is the also first study to make use of human genomic data to guide drug discovery or repositioning for comorbid disorders.

genomics