Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,045 records · Page 58Linked to original sources

XWAS: a software toolset for genetic data analysis and association studies of the X chromosome

XWAS is a new software suite for the analysis of the X chromosome in association studies and similar studies. The X chromosome plays an important role in human disease, especially those with sexually dimorphic characteristics. Special attention needs to be given to its analysis due to the unique inheritance pattern, which leads to analytical complications that have resulted in the majority of genome-wide association studies (GWAS) either not considering X or mishandling it with toolsets that had been designed for non-sex chromosomes. We hence developed XWAS to fill the need for tools that are specially designed for analysis of X. Following extensive, stringent, and X-specific quality control, XWAS offers an array of statistical tests of association, including: (1) the standard test between a SNP (single nucleotide polymorphism) and disease risk, including after first stratifying individuals by sex, (2) a test for a differential effect of a SNP on disease between males and females, (3) motivated by X-inactivation, a test for higher variance of a trait in heterozygous females as compared to homozygous females, and (4) for all tests, a version that allows for combining evidence from all SNPs across a gene. We applied the toolset analysis pipeline to 16 GWAS datasets of immune-related disorders and 7 risk factors of coronary artery disease, and discovered several new X-linked genetic associations. XWAS will provide the tools and incentive for others to incorporate the X chromosome into GWAS, hence enabling discoveries of novel loci implicated in many diseases and in their sexual dimorphism.

Genetics

Integration of experiments across diverse environments identifies the genetic determinants of variation in Sorghum bicolor seed element composition

AbstractSeedling establishment and seed nutritional quality require the sequestration of sufficient element nutrients. Identification of genes and alleles that modify element content in the grains of cereals, including Sorghum bicolor, is fundamental to developing breeding and selection methods aimed at increasing bioavailable element content and improving crop growth. We have developed a high throughput workflow for the simultaneous measurement of multiple elements in sorghum seeds. We measured seed element levels in the genotyped Sorghum Association Panel (SAP), representing all major cultivated sorghum races from diverse geographic and climatic regions, and mapped alleles contributing to seed element variation across three environments by genome-wide association. We observed significant phenotypic and genetic correlation between several elements across multiple years and diverse environments. The power of combining high-precision measurements with genome wide association was demonstrated by implementing rank transformation and a multilocus mixed model (MLMM) to map alleles controlling 20 element traits, identifying 255 loci affecting the sorghum seed ionome. Sequence similarity to genes characterized in previous studies identified likely causative genes for the accumulation of zinc (Zn) manganese (Mn), nickel (Ni), calcium (Ca) and cadmium (Cd) in sorghum seed. In addition to strong candidates for these four elements, we provide a list of candidate loci for several other elements. Our approach enabled identification of SNPs in strong LD with causative polymorphisms that can be evaluated in targeted selection strategies for plant breeding and improvement.\n\nOne sentence summaryHigh-throughput measurements of element accumulation and genome-wide association analysis across multiple environments identified novel alleles controlling seed element accumulation in Sorghum bicolor.\n\nThis project was partially funded by the iHUB Visiting Scientist Program (http://www.ionomicshub.org), Chromatin, Inc., NSF EAGER (1450341) to I.B. and BPD, NSF IOS 1126950 to IB, NSF IOS-0919739 to EC, and BMGF (OPP 1052924) to B.P.D.

Genetics

The Nature, Extent, and Consequences of Cryptic Genetic Variation in the opa Repeats of Notch in Drosophila

ABTRACTPolyglutamine (pQ) tracts are abundant in many proteins co-interacting on DNA. The lengths of these pQ tracts can modulate their interaction strengths. However, pQ tracts > 40 residues are pathologically prone to amyloidogenic self-assembly. Here, we assess the extent and consequences of variation in the pQ-encoding opa repeats of Notch (N) in Drosophila melanogaster. We use Sanger sequencing to genotype opa sequences (5-CAX repeats), which have resisted assembly using short sequence reads. While the majority of N sequences pertain to reference opa31 (Q13HQ17) and opa32 (Q13HQ18) allelic classes, several rare alleles encode tracts > 32 residues: opa33a (Q14HQ18), opa33b (Q15HQ17), opa34 (Q16HQ17), opa35a1/opa35a2 (Q13HQ21), opa36 (Q13HQ22), and opa37 (Q13HQ23). Only one rare allele encodes a tract < 31 residues: opa23 (Q13-Q10). This opa23 allele shortens the pQ tract while simultaneously eliminating the interrupting histidine. Homozygotes for the short and long opa alleles have defects in sensory bristle organ specification, abdominal patterning, and embryonic survival. Inbred stocks with wild-type opa31 alleles become more viable when outbred, while an inbred stock with the longer opa35 becomes less viable after outcrossing to different backgrounds. In contrast, an inbred stock with the short opa23 allele is semi-viable in both inbred and outbred genetic backgrounds. This opa23 Notch allele also produces notched wings when recombined out of the X chromosome. Importantly, wa-linked X balancers carry the N allele opa33b and suppress AS-C insufficiency caused by the sc8 inversion. Our results demonstrate significant cryptic variation and epistatic sensitivity for the N locus, and the need for long read genotyping of key repeat variables underlying gene regulatory networks.

Genetics

An accurate genetic clock

Our method for \"Time to most recent common ancestor\" TMRCA of genetic trees for the first time deals with natural selection by apriori mathematics and not as a random factor. Bioprocesses such as \"kin selection\" generate a few overrepresented \"singular lineages\" while almost all other lineages terminate. This non-uniform branching gives greatly exaggerated TMRCA with current methods. Thus we introduce an inhomogenous stochastic process which will detect singular lineages by asymmetries, whose \"reduction\" then gives true TMRCA. Reduction implies younger TMRCA, with smaller errors. This gives a new phylogenetic method for computing mutation rates, with results similar to \"pedigree\" (meiosis) data. Despite these low rates, reduction implies younger TMRCA, with smaller errors. We establish accuracy by a comparison across a wide range of time, indeed this is only y-clock giving consistent results for 500-15,000 ybp. In particular we show that the dominant European y-haplotypes R1a1a & R1b1a2, expand from c3700BC, not reaching Anatolia before c3300BC. This contradicts current clocks dating R1b1a2 to either the Neolithic Near East or Paleo-Europe. However our dates match R1a1a & R1b1a2 found in Yamnaya cemetaries of c3300BC by Svante Paabo et al, together proving R1a1a & R1b1a2 originates in the Russian Steppes.

Genetics

The genetic basis of cone serotiny in Pinus contorta as a function of mixed-severity and stand-replacement fire regimes

ABSTRACT\n\nWildfires and mountain pine beetle (MPB) attacks are important contributors to the development of stand structure in lodgepole pine, and major drivers of its evolution. The historical pattern of these events have been correlated with variation in cone serotiny (possessing cones that remain closed and retain seeds until opened by fire) across the Rocky Mountain region of Western North America. As climate change brings about a marked increase in the size, intensity, and severity of our wildfires, it is becoming increasingly important to study the genetic basis of serotiny as an adaptation to wildfire. Knowledge gleaned from these studies would have direct implications for forest management in the future, and for the future. In this study, we collected physical data and DNA samples from 122 trees of two different areas in the IDF-dk of British Columbia; multi-cohort stands (Cariboo-Chilcotin) with a history of mixed-severity fire and frequent MPB disturbances, and single-cohort stands (Logan Lake) with a history of stand replacing (crown) fire and infrequent MPB disturbances. We used QuantiNemo to construct simulated populations of lodgepole pine at five different growth rates, and compared the statistical outputs to physical data, then ran a random forest analysis to shed light on sources of variation in serotiny. We also sequenced 39 SNPs, of which 23 failed or were monomorphic. The 16 informative SNPs were used to calculate HO and HE, which were included alongside genotypes for a second random forest analysis. Our best random forest model explained 33% of variation in serotiny, using simulation and physical variables. Our results highlight the need for more investigation into this matter, using more extensive approaches, and also consideration of alternative methods of heredity such as epigenetics.

Genetics

An Accurate Genetic Clock

Our method for \"Time to most recent common ancestor\" TMRCA of genetic trees for the first time deals with natural selection by apriori mathematics and not as a random factor. Bioprocesses such as \"kin selection\" generate a few overrepresented \"singular lineages\" while almost all other lineages terminate. This non-uniform branching gives greatly exaggerated TMRCA with current methods. Thus we introduce an inhomogenous stochastic process which will detect singular lineages by asymmetries, whose \"reduction\" then gives true TMRCA. This gives a new phylogenetic method for computing mutation rates, with results similar to \"pedigree\" (meiosis) data. Despite these low rates, reduction implies younger TMRCA, with smaller errors. We establish accuracy by a comparison across a wide range of time, indeed this is only y-clock giving consistent results for 500-15,000 ybp. In particular we show that the dominant European Y-haplotypes R1a1a & R1b1a2, expand from c4000BC, not reaching Anatolia before c3800BC. This contradicts previous clocks dating R1b1a2 to either the Neolithic Near East or Paleo-Europe. However our dates match R1a1a & R1b1a2 found in Yamnaya cemetaries of c3300BC by Nielsen et al (2015), Paabo et al(2015), together proving R1a1a & R1b1a2 originates in the Russian Steppes.

Genetics

Analytical limits of hybrid identification using genetic markers: an empirical and simulation study in Hippolais warblers

Hybridization is known to occur in a wide range of avian species, yet the rate and persistence of hybridization on populations is often hard to assess. Genotyping using variable genetic marker sets has become a common tool to identify hybrid individuals, however assignment outputs can differ depending on the marker set used. Here, we study hybrid assignment in two sibling Hippolais warblers, where hybrid assignment has shown to differ between SSR and AFLP markers. Simulation of heterospecific individuals as well as backcrosses (typed using SSR markers) reveals a rapid loss of assignment probability in higher backcross generations.\n\nHowever, the characterization of F1 hybrids was clearly distinguished from both parental taxa. The differences in marker sets are not contradictory but complementary. The rate of hybridization is lower than previously expected with AFLP markers but introgression might be long-lasting. This could be either due to differences in power of the marker systems used or due to non-neutral variation covered by AFLP but not SSR markers. We call for more attention to be paid regarding the potential limits of classical marker systems to investigate hybridization and its persistence in natural systems.

Genetics

Reconstructing Genetic History of Siberian and Northeastern European Populations

Siberia and Western Russia are home to over 40 culturally and linguistically diverse indigenous ethnic groups. Yet, genetic variation of peoples from this region is largely uncharacterized. We present whole-genome sequencing data from 28 individuals belonging to 14 distinct indigenous populations from that region. We combine these datasets with additional 32 modern-day and 15 ancient human genomes to build and compare autosomal, Y-DNA and mtDNA trees. Our results provide new links between modern and ancient inhabitants of Eurasia. Siberians share 38% of ancestry with descendants of the 45,000-year-old Ust-Ishim people, who were previously believed to have no modern-day descendants. Western Siberians trace 57% of their ancestry to the Ancient North Eurasians, represented by the 24,000-year-old Siberian Malta boy. In addition, Siberians admixtures are present in lineages represented by Eastern European hunter-gatherers from Samara, Karelia, Hungary and Sweden (from 8,000-6,600 years ago), as well as Yamnaya culture people (5,300-4,700 years ago) and modern-day northeastern Europeans. These results provide new evidence of ancient gene flow from Siberia into Europe.

Genetics

Multiple Conformations of Gal3 Protein Drive the Galactose Induced Allosteric Activation of the GAL Genetic Switch of Saccharomyces cerevisiae

Gal3p is an allosteric monomeric protein which activates the GAL genetic switch of Saccharomyces cerevisiae in response to galactose. Expression of constitutive mutant of Gal3p or over-expression of wild-type Gal3p activates the GAL switch in the absence of galactose. These data suggest that Gal3p exists as an ensemble of active and inactive conformations. Structural data has indicated that Gal3p exists in open (inactive) and closed (active) conformations. However, mutant of Gal3p that predominantly exists in inactive conformation and yet capable of responding to galactose has not been isolated. To understand the mechanism of allosteric transition, we have isolated a triple mutant of Gal3p with V273I, T404A and N450D substitutions which upon over-expression fails to activate the GAL switch on its own, but activates the switch in response to galactose. Over-expression of Gal3p mutants with single or double mutations in any of the three combinations failed to exhibit the behavior of the triple mutant. Molecular dynamics analysis of the wild-type and the triple mutant along with two previously reported constitutive mutants suggests that the wild-type Gal3p may also exist in super-open conformation. Further, our results suggest that the dynamics of residue F237 situated in the hydrophobic pocket located in the hinge region drives the transition between different conformations. Based on our study and what is known in human glucokinase, we suggest that the above mechanism could be a general theme in causing the allosteric transition.

Genetics

CRISPR-directed mitotic recombination enables genetic mapping without crosses

Linkage and association studies have mapped thousands of genomic regions that contribute to phenotypic variation, but narrowing these regions to the underlying causal genes and variants has proven much more challenging. Resolution of genetic mapping is limited by the recombination rate. We developed a method that uses CRISPR to build mapping panels with targeted recombination events. We tested the method by generating a panel with recombination events spaced along a yeast chromosome arm, mapping trait variation, and then targeting a high density of recombination events to the region of interest. Using this approach, we fine-mapped manganese sensitivity to a single polymorphism in the transporter Pmr1. Targeting recombination events to regions of interest allows us to rapidly and systematically identify causal variants underlying trait differences.

Genetics

Personalized Risk Prediction for Type 2 Diabetes: the Potential of Genetic Risk Scores

PurposeThe study aims to develop a Genetic Risk Score (GRS) for the prediction of Type 2 Diabetes (T2D) that could be used for risk assessment in general population.\n\nMethodsUsing the results of genome-wide association studies, we develop a doubly-weighted GRS for the prediction of T2D risk, aiming to capture the effect of 1000 single nucleotide polymorphisms. The GRS is evaluated in the Estonian Biobank cohort (n=10273), analysing its effect on prevalent and incident T2D, while adjusting for other predictors. We assessed the effect of GRS on all-cause and cardiovascular mortality and its association with other T2D risk factors, and conducted the reclassification analysis.\n\nResultsThe adjusted hazard for incident T2D is 1.90 (95% CI 1.48, 2.44) times higher and for cardiovascular mortality 1.27 (95% CI 1.10, 1.46) times higher in the highest GRS quintile compared to the rest of the cohort. No significant association between BMI and GRS is found in T2D-free individuals. Adding GRS to the prediction model for 5-year T2D risks results in continuous Net Reclassification Improvement of 0.26 (95% CI 0.15, 0.38).\n\nConclusionThe proposed GRS would considerably improve the accuracy of T2D risk prediction when added to the set of predictors used so far.

Genetics

Forward genetics by sequencing EMS variation induced inbred lines.

In order to leverage novel sequencing techniques for cloning genes in eukaryotic organisms with complex genomes, the false positive rate of variant discovery must be controlled for by experimental design and informatics. We sequenced five lines from three pedigrees of EMS mutagenized Sorghum bicolor, including a pedigree segregating a recessive dwarf mutant. Comparing the sequences of the lines, we were able to identify and eliminate error prone positions. One genomic region contained EMS mutant alleles in dwarfs that were homozygous reference sequence in wild-type siblings and heterozygous in segregating families. This region contained a single non-synonymous change that cosegregated with dwarfism in a validation population and caused a premature stop codon in the sorghum ortholog encoding the giberellic acid biosynthetic enzyme ent-kaurene oxidase. Application of exogenous giberillic acid rescued the mutant phenotype. Our method for mapping did not require outcrossing and introduced no segregation variance. This enables work when line crossing is complicated by life history, permitting gene discovery outside of genetic models.This inverts the historical approach of first using recombination to define a locus and then sequencing genes. Our formally identical approach first sequences all the genes and then seeks co-segregation with the trait. Mutagenized lines lacking obvious phenotypic alterations are available for an extention of this approach: mapping with a known marker set in a line that is phenotypically identical to starting material for EMS mutant generation.

Genetics

Cosmid based mutagenesis causes genetic instability in Streptomyces coelicolor, as shown by targeting of the lipoprotein signal peptidase gene

Bacterial lipoproteins are a class of extracellular proteins tethered to cell membranes by covalently attached lipids. Deleting the lipoprotein signal peptidase (lsp) gene in Streptomyces coelicolor results in growth and developmental defects that cannot be restored by reintroducing the lsp. We report resequencing of the genomes of the wild-type M145 and the cis-complemented {Delta}lsp mutant (BJT1004), mapping and identifying secondary mutations, including an insertion into a novel putative small RNA, scr6809. Disruption of scr6809 led to a range of developmental phenotypes. However, these secondary mutations do not increase the efficiency of disrupting lsp suggesting they are not lsp specific suppressors. Instead we suggest that these were induced by introducing the cosmid St4A10{Delta}lsp as part of the Redirect mutagenesis protocol, which transiently duplicates a number of important cell division genes. Disruption of lsp using no gene duplication resulted in the previously observed phenotype. We conclude that lsp is not essential in S. coelicolor but loss of lsp does lead to developmental defects due to the loss of lipoproteins from the cell. Significantly, our results indicate the use of cosmid libraries for the genetic manipulation of bacteria can lead to unexpected phenotypes not necessarily linked to the gene or pathway of interest.

Genetics

The genetic basis of natural variation in C. elegans telomere length

Telomeres are involved in the maintenance of chromosomes and the prevention of genome instability. Despite this central importance, significant variation in telomere length has been observed in a variety of organisms. The genetic determinants of telomere-length variation and their effects on organismal fitness are largely unexplored. Here, we describe natural variation in telomere length across the Caenorhabditis elegans species. We identify a large-effect variant that contributes to differences in telomere length. The variant alters the conserved oligosaccharide/oligonucleotide-binding fold of POT-2, a homolog of a human telomere-capping shelterin complex subunit. Mutations within this domain likely reduce the ability of POT-2 to bind telomeric DNA, thereby increasing telomere length. We find that telomere-length variation does not correlate with offspring production or longevity in C. elegans wild isolates, suggesting that naturally long telomeres play a limited role in modifying fitness phenotypes in C. elegans.

Genetics

Exploring the genetic architecture of inflammatory bowel disease by whole genome sequencing identifies association at ADCY7

In order to further resolve the genetic architecture of the inflammatory bowel diseases, ulcerative colitis and Crohns disease, we sequenced the whole genomes of 4,280 patients at low coverage, and compared them to 3,652 previously sequenced population controls across 73.5 million variants. To increase power we imputed from these sequences into new and existing GWAS cohorts, and tested for association at ~12 million variants in a total of 16,432 cases and 18,843 controls. We discovered a 0.6% frequency missense variant in ADCY7 that doubles risk of ulcerative colitis, and offers insight into a new aspect of disease biology. Despite good statistical power, we did not identify any other new low-frequency risk variants, and found that such variants as a class explained little heritability. We did detect a burden of very rare, damaging missense variants in known Crohns disease risk genes, suggesting that more comprehensive sequencing studies will continue to improve our understanding of the biology of complex diseases.

Genetics

A gene for genetic background in Zea mays: fine-mapping enhancer of teosinte branched1.2 (etb1.2) to a YABBY class transcription factor

The effects of an allelic substitution at a gene often depend critically on genetic background, the genotype at other genes in the genome. During the domestication of maize from its wild ancestor (teosinte), an allelic substitution at teosinte branched (tb1) caused changes in both plant and ear architecture. The effects of tb1 on phenotype were shown to depend on multiple background loci including one called enhancer of tb1.2 (etb1.2). We mapped etb1.2 to a YABBY class transcription factor (ZmYAB2.1) and showed that the maize alleles of ZmYAB2.1 are either expressed at a lower level than teosinte alleles or disrupted by insertions in the sequences. tb1 and etb1.2 interact epistatically to control the length of internodes within the maize ear which affects how densely the kernels are packed on the ear. The interaction effect is also observed at the level of gene expression with tb1 acting as a repressor of ZmYAB2.1 expression. Curiously, ZmYAB2.1 was previously identified as a candidate gene for another domestication trait in maize, non-shattering ears. Consistent with this proposed role,ZmYAB2.1 is expressed in a narrow band of cells in immature ears that appears to represent a vestigial abscission (shattering) zone. Expression in this band of cells may also underlie the effect on internode elongation. The identification of ZmYAB2.1 as a background factor interacting with tb1 is a first step toward a gene-level understanding of how tb1 and the background within which it works evolved in concert during maize domestication.

Genetics

Genome-wide association analyses of sleep disturbance traits identify new loci and highlight shared genetics with neuropsychiatric and metabolic traits

Chronic sleep disturbances, associated with cardio-metabolic diseases, psychiatric disorders and all-cause mortality1,2, affect 25-30% of adults worldwide3. While environmental factors contribute importantly to self-reported habitual sleep duration and disruption, these traits are heritable4-9, and gene identification should improve our understanding of sleep function, mechanisms linking sleep to disease, and development of novel therapies. We report single and multi-trait genome-wide association analyses (GWAS) of self-reported sleep duration, insomnia symptoms including difficulty initiating and/or maintaining sleep, and excessive daytime sleepiness in the UK Biobank (n=112,586), with discovery of loci for insomnia symptoms (near MEIS1, TMEM132E, CYCL1, TGFBI in females and WDR27 in males), excessive daytime sleepiness (near AR/OPHN1) and a composite sleep trait (near INADL and HCRTR2), as well as replication of a locus for sleep duration (at PAX-8). Genetic correlation was observed between longer sleep duration and schizophrenia (rG=0.29, p=1.90x10-13) and between increased excessive daytime sleepiness and increased adiposity traits (BMI rG=0.20, p=3.12x10-09; waist circumference rG=0.20, p=2.12x10-07).

genetics

Connecting genetic risk to disease endpoints through the human blood plasma proteome

Genome-wide association studies (GWAS) with intermediate phenotypes, like changes in metabolite and protein levels, provide functional evidence for mapping disease associations and translating them into clinical applications. However, although hundreds of genetic risk variants have been associated with complex disorders, the underlying molecular pathways often remain elusive. Associations with intermediate traits across multiple chromosome locations are key in establishing functional links between GWAS-identified risk-variants and disease endpoints. Here, we describe a GWAS performed with a highly multiplexed aptamer-based affinity proteomics platform. We quantified associations between protein level changes and gene variants in a German cohort and replicated this GWAS in an Arab/Asian cohort. We identified many independent, SNP-protein associations, which represent novel, inter-chromosomal links, related to autoimmune disorders, Alzheimer's disease, cardiovascular disease, cancer, and many other disease endpoints. We integrated this information into a genome-proteome network, and created an interactive web-tool for interrogations. Our results provide a basis for new approaches to pharmaceutical and diagnostic applications.

genetics