Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,063 records · Page 59Linked to original sources

Statistical properties of simple random-effects models for genetic heritability

Random-effects models are a popular tool for analysing total narrow-sense heritability for simple quantitative phenotypes on the basis of large-scale SNP data. Recently, there have been disputes over the validity of conclusions that may be drawn from such analysis. We derive some of the fundamental statistical properties of heritability estimates arising from these models, showing that the bias will generally be small. We show that that the score function may be manipulated into a form that facilitates intelligible interpretations of the results. We use this score function to explore the behavior of the model when certain key assumptions of the model are not satisfied -- shared environment, measurement error, and genetic effects that are confined to a small subset of sites -- as well as to elucidate the meaning of negative heritability estimates that may arise.\n\nThe variance and bias depend crucially on the variance of certain functionals of the singular values of the genotype matrix. A useful baseline is the singular value distribution associated with genotypes that are completely independent -- that is, with no linkage and no relatedness -- for a given number of individuals and sites. We calculate the corresponding variance and bias for this setting.\n\nMSC 2010 subject classifications: Primary 92D10; secondary 62P10; 62F10; 60B20.

genetics

Genetic overlap between educational attainment, schizophrenia and autism

ImportanceThe genetic relationship between cognition, autism, and schizophrenia is complex. It is unclear how genes that contribute to cognition also contribute to risk for autism and schizophrenia.\n\nObjectiveTo investigate the interaction between genes related to cognition (measured via proxy through educational attainment, which we call edu genes) and genes/biological pathways that are atypical in autism and schizophrenia.\n\nDesignGenetic correlation and enrichment analysis were conducted to identify the interaction between edu genes and risk genes and biological pathways for autism or schizophrenia.\n\nResultsFirst, edu genes are enriched in a specific developmental co-expression module that is also enriched for high confidence autism risk genes. Second, modules enriched for genes that are dysregulated in autism and schizophrenia are also enriched for edu genes. Finally, genes that overlap between the two above modules and educational attainment are significantly enriched for genes that flank human accelerated regions, suggesting increased positive selection for the overlapping gene sets.\n\nConclusionOur results identify distinct co-expression modules where risk genes for the two psychiatric conditions interact with edu genes. This suggests specific pathways that contribute to both cognitive deficits and cognitive talents, in individuals with schizophrenia or autism.\n\nKey PointsO_ST_ABSQuestionC_ST_ABSHow do genes for educational attainment interact with risk genes for autism and schizophrenia?\n\nFindingsWe show that genes for educational attainment (edu genes) are significantly likely to be mutated in autism and intellectual disability. We further show that edu genes also interact with co-expression modules that are associated with autism or schizophrenia and are enriched for differentially expressed genes in autism or schizophrenia. Finally, we identify that the enrichment between risk genes for autism and schizophrenia and human accelerated regions are driven, in part, by their overlap with edu genes.\n\nMeaningEdu genes interact with schizophrenia and autism risk genes in specific pathways, contributing to both cognitive deficits and talents.

genetics

Estimating genetic kin relationships in prehistoric populations

Archaeogenomic research has proven to be a valuable tool to trace migrations of historic and prehistoric individuals and groups, whereas relationships within a group or burial site have not been investigated to a large extent. Knowing the genetic kinship of historic and prehistoric individuals would give important insights into social structures of ancient and historic cultures. Most archaeogenetic research concerning kinship has been restricted to uniparental markers, while studies using genome-wide information were mainly focused on comparisons between populations. Applications which infer the degree of relationship based on modern-day DNA information typically require diploid genotype data. Low concentration of endogenous DNA, fragmentation and other post-mortem damage to ancient DNA (aDNA) makes the application of such tools unfeasible for most archaeological samples. To infer family relationships for degraded samples, we developed the software READ (Relationship Estimation from Ancient DNA). We show that our heuristic approach can successfully infer up to second degree relationships with as little as 0.1x shotgun coverage per genome for pairs of individuals. We uncover previously unknown relationships among prehistoric individuals by applying READ to published aDNA data from several human remains excavated from different cultural contexts. In particular, we find a group of five closely related males from the same Corded Ware culture site in modern-day Germany, suggesting patrilocality, which highlights the possibility to uncover social structures of ancient populations by applying READ to genome-wide aDNA data.

genetics

Insights into the genetic determinism andevolution of recombination rates fromcombining multiple genome-wide datasets inSheep

Recombination is a complex biological process that results from a cascade of multiple events during meiosis. Understanding the genetic determinism of recombination can help to understand if and how these events are interacting. To tackle this question, we studied the patterns of recombination in sheep, using multiple approaches and datasets. We constructed male recombination maps in a dairy breed from the south of France (the Lacaune breed) at a fine scale by combining meiotic recombination rates from a large pedigree genotyped with a 50K SNP array and historical recombination rates from a sample of unrelated individuals genotyped with a 600K SNP array. This analysis revealed recombination patterns in sheep similar to other mammals but also genome regions that have likely been affected by directional and diversifying selection. We estimated the average recombination rate of Lacaune sheep at 1.5 cM/Mb, identified about 50,000 crossover hotspots on the genome and found a high correlation between historical and meiotic recombination rate estimates. A genome-wide association study revealed two major loci affecting inter-individual variation in recombination rate in Lacaune, including the RNF212 and HEI10 genes and possibly 2 other loci of smaller effects including the KCNJ15 and FSHR genes. Finally, we compared our results to those obtained previously in a distantly related population of domestic sheep, the Soay. This comparison revealed that Soay and Lacaune males have a very similar distribution of recombination along the genome and that the two datasets can be combined to create more precise male meiotic recombination maps in sheep. Despite their similar recombination maps, we show that Soay and Lacaune males exhibit different heritabilities and QTL effects for inter-individual variation in genome-wide recombination rates.

genetics

Genome-scale genetic interactions position the Mitotic Exit Network as a major antagonist of transient Topoisomerase II deficiency.

Topoisomerase II (Top2) is the essential protein that resolves DNA catenations. When the Top2 is inactivated, mitotic catastrophe results from massive entanglement of chromosomes. Top2 is also the target of many first-line anticancer drugs, the so-called Top2 poisons. Often, tumours become resistant to these drugs by downregulating Top2. Here, we have compared two isogenic yeast strains carrying top2 thermosensitive alleles that differ in their resistance to Top2 poisons, the broadly-used poison-sensitive top2-4 and the poison-resistant top2-5. We found that top2-5 transits through anaphase faster than top2-4. In order to define the biological importance of this difference, we performed genome-scale Synthetic Gene Array (SGA) analyses during chronic sublethal Top2 downregulation and acute, yet transient, Top2 inactivation. We find that downregulation of cell cycle progression, especially the Mitotic Exit Network (MEN), protects against Top2 deficiency. In all conditions, genetic protection was stronger in top2-5, and this correlated with destabilization of anaphase bridges by execution of MEN. We suggest that mitotic exit may be a therapeutic target to hypersensitize cancer cells carrying downregulating mutations in TOP2.

genetics

STPGA: Selection of training populations with a genetic algorithm

Optimal subset selection is an important task that has numerous algorithms designed for it and has many application areas. STPGA contains a special genetic algorithm supplemented with a tabu memory property (that keeps track of previously tried solutions and their fitness for a number of iterations), and with a regression of the fitness of the solutions on their coding that is used to form the ideal estimated solution (look ahead property) to search for solutions of generic optimal subset selection problems. I have initially developed the programs for the specific problem of selecting training populations for genomic prediction or association problems, therefore I give discussion of the theory behind optimal design of experiments to explain the default optimization criteria in STPGA, and illustrate the use of the programs in this endeavor. Nevertheless, I have picked a few other areas of application: supervised and unsupervised variable selection based on kernel alignment, supervised variable selection with design criteria, influential observation identification for regression, solving mixed integer quadratic optimization problems, balancing gains and inbreeding in a breeding population. Some of these illustrations pertain new statistical approaches.

genetics

Novel CRISPR/Cas9 gene drive constructs in Drosophila reveal insights into mechanisms of resistance allele formation and drive efficiency in genetically diverse populations

A functioning gene drive system could fundamentally change our strategies for the control of vector-borne diseases by facilitating rapid dissemination of transgenes that prevent pathogen transmission or reduce vector capacity. CRISPR/Cas9 gene drive promises such a mechanism, which works by converting cells that are heterozygous for the drive construct into homozygotes, thereby enabling super-Mendelian inheritance. Though CRISPR gene drive activity has already been demonstrated, a key obstacle for current systems is their propensity to generate resistance alleles. In this study, we developed two CRISPR gene drive constructs based on the nanos and vasa promoters that allowed us to illuminate the different mechanisms by which resistance alleles are formed in the model organism Drosophila melanogaster. We observed resistance allele formation at high rates both prior to fertilization in the germline and post-fertilization in the embryo due to maternally deposited Cas9. Assessment of drive activity in genetically diverse backgrounds further revealed substantial differences in conversion efficiency and resistance rates. Our results demonstrate that the evolution of resistance will likely impose a severe limitation to the effectiveness of current CRISPR gene drive approaches, especially when applied to diverse natural populations.

genetics

Genetic variability in the promoter region of TNF-α gene in the reservoir of Junin virus, Calomys musculinus (Rodentia, Cricetidae)

In this study, we assessed the genetic variability of the promoter region of TNF- gene in three natural populations of the cricetid rodent Calomys musculinus. This species is the natural reservoir of Junin virus, the etiological agent of Argentine Hemorrhagic fever. We found different levels of variability and varying signatures of natural selection in populations with different epidemiological histories.

genetics

Failure to Replicate a Genetic Signal for Sex Bias in the Steppe Migration into Central Europe

Goldberg et al.(1) used genome-wide ancient DNA data (2) from central European Bronze Age (BA) populations, and their three ancestral sources of steppe pastoralists (SP), Anatolian farmers (AF), and European hunter-gatherers (HG), to investigate whether the SP migration into central Europe after 5,000 years ago (3, 4) was sex biased. By estimating a lower proportion of SP ancestry on the X-chromosome (36.6%) which is primarily carried in females than on the autosomes (61.8%), they suggested that the migration involved a ratio of 5-14 SP males for every female.\n\nWe attempted to replicate this finding using qpAdm (3), which leverages allele frequency correlations between the admixed (BA) and source (SP, AF, HG) populations with distant outgroups to eliminate potential biases due to genetic drift between the true source populations and the ones used as surrogates for them. Our outgroups are ...

genetics

Crosslink: A Fast, Scriptable Genetic Mapper For Outcrossing Species

SummaryCrosslink is genetic mapping software for outcrossing species designed to run efficiently on large datasets by combining the best from existing tools with novel approaches. Tests show it runs much faster than several comparable programs whilst retaining a similar accuracy.\n\nAvailability and implementationAvailable under the GNU General Public License version 2 from https://github.com/eastmallingresearch/crosslink\n\nContactrobert.vickerstaff@emr.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online and from https://github.com/eastmallingresearch/crosslink/releases/tag/v0.5.

genetics

Accumulation And Functional Architecture Of Deleterious Genetic Variants During The Extinction Of Wrangel Island Mammoths

Woolly mammoths were among the most abundant cold adapted species during the Pleistocene. Their once large populations went extinct in two waves, an end-Pleistocene extinction of continental populations followed by the mid-Holocene extinction of relict populations on St. Paul Island ~5,600 years ago and Wrangel Island ~4,000 years ago. Wrangel Island mammoths experienced an episode of rapid demographic decline coincident with their isolation, leading to a small population, reduced genetic diversity, and the fixation of putatively deleterious alleles, but the functional consequences of these processes are unclear. Here we show that the Wrangel Island mammoth accumulated many putative deleterious mutations that are predicted to cause diverse behavioral and developmental defects. Resurrection and functional characterization of Wrangel Island mammoth genes carrying these substitutions identified both loss and gain of function mutations in genes associated with developmental defects (HYLS1), oligozoospermia and reduced male fertility (NKD1), diabetes (NEUROG3), and the ability to detect floral scents (OR5A1). These results suggest that Wrangel Island mammoths may have suffered adverse consequences from their reduced population sizes and isolation.

genetics

Genetics of the Research Domain Criteria (RDoC): genome-wide association study of delay discounting

Delay discounting (DD), which is the tendency to discount the value of delayed versus current rewards, is elevated in a constellation of diseases and behavioral conditions. We performed a genome-wide association study of DD using 23,127 research participants of European ancestry. The most significantly associated SNP was rs6528024 (P = 2.40 x 10-8), which is located in an intron of the gene GPM6B. We also showed that 12% of the variance in DD was accounted for by genotype, and that the genetic signature of DD overlapped with attention-deficit/hyperactivity disorder, schizophrenia, major depression, smoking, personality, cognition, and body weight.

genetics

Placental gene expression mediates the interaction between obstetrical history and genetic risk for schizophrenia

Defining the environmental context in which genes enhance susceptibility can provide insight into the pathogenesis of complex disorders, like schizophrenia. Here we show that the intrauterine and perinatal environment modulates the association of schizophrenia with genomic risk, as measured with polygenic risk scores (PRS) based primarily on GWAS significant variants. Genomic risk interacts with intrauterine and perinatal complications (Early Life Complications, ELCs) in each of three independent samples from USA, Italy and Germany (overall n= 1693, p= 6e-05). In each sample, the liability of schizophrenia explained by PRS is nominally more than five times greater in the presence of a history of ELCs compared with its absence. In each sample, patients with positive ELC histories have higher PRS than patients without ELCs, which is further confirmed in two additional patient samples from Germany and Japan (overall n=2038, p= 1e-04). The gene set based on the schizophrenia loci interacting with ELCs is highly expressed in multiple placental compartments and dynamically regulated in placenta from complicated in comparison with normal pregnancies. The same genes are differentially up-regulated in placentae from male compared with female offspring. The interaction between genomic risk and ELCs is mainly driven by GWAS significant loci enriched for genes highly expressed in the various placenta samples. Molecular pathway analyses based on the genes not driving this interaction reflect previous analyses about schizophrenia risk-genes, while genes highly and differentially expressed in placentae implicate an orthogonal biology involving cellular stress. These results suggest that the most significant genetic variants detected by current schizophrenia GWAS contribute to risk in part by converging on a developmental trajectory sensitive to events affecting placentation, which may underlie the male preponderance of schizophrenia and offer new insights into primary prevention.

genetics

New insights from Thailand into the maternal genetic history of Mainland Southeast Asia

Tai-Kadai (TK) is one of the major language families in Mainland Southeast Asia (MSEA), with a concentration in the area of Thailand and Laos. Our previous study of 1,234 mtDNA genome sequences supported a demic diffusion scenario in the spread of TK languages from southern China to Laos as well as northern and northeastern Thailand. Here we add an additional 560 mtDNA sequences from 22 groups, with a focus on the TK-speaking central Thai people and the Sino-Tibetan speaking Karen. We find extensive diversity, including 62 haplogroups not reported previously from this region. Demic diffusion is still a preferable scenario for central Thais, emphasizing the extension and expansion of TK people through MSEA, although there is also some support for an admixture model. We also tested competing models concerning the genetic relationships of groups from the major MSEA languages, and found support for an ancestral relationship of TK and Austronesian-speaking groups.

genetics

Genome-wide Association Studies Reveal Similar Genetic Architecture with Shared and Unique QTL for Bacterial Cold Water Disease Resistance in Two Rainbow Trout Breeding Populations

Bacterial cold water disease (BCWD) causes significant mortality and economic losses in salmonid aquaculture. In previous studies, we identified moderate-large effect QTL for BCWD resistance in rainbow trout (Oncorhynchus mykiss). However, the recent availability of a 57K SNP array and a genome physical map have enabled us to conduct genome-wide association studies (GWAS) that overcome several experimental limitations from our previous work. In the current study, we conducted GWAS for BCWD resistance in two rainbow trout breeding populations using two genotyping platforms, the 57K Affymetrix SNP array and restriction-associated DNA (RAD) sequencing. Overall, we identified 14 moderate-large effect QTL that explained up to 60.8% of the genetic variance in one of the two populations and 27.7% in the other. Four of these QTL were found in both populations explaining a substantial proportion of the variance, although major differences were also detected between the two populations. Our results confirm that BCWD resistance is controlled by the oligogenic inheritance of few moderate-large effect loci and a large-unknown number of loci each having a small effect on BCWD resistance. We detected differences in QTL number and genome location between two GWAS models (weighted single-step GBLUP and Bayes B), which highlights the utility of using different models to uncover QTL. The RAD-SNPs detected a greater number of QTL than the 57K SNP array in one population, suggesting that the RAD-SNPs may uncover polymorphisms that are more unique and informative for the specific population in which they were discovered.

genetics

Genome-wide genetic data on ~500,000 UK Biobank participants

The UK Biobank project is a large prospective cohort study of ~500,000 individuals from across the United Kingdom, aged between 40-69 at recruitment. A rich variety of phenotypic and health-related information is available on each participant, making the resource unprecedented in its size and scope. Here we describe the genome-wide genotype data (~805,000 markers) collected on all individuals in the cohort and its quality control procedures. Genotype data on this scale offers novel opportunities for assessing quality issues, although the wide range of ancestries of the individuals in the cohort also creates particular challenges. We also conducted a set of analyses that reveal properties of the genetic data - such as population structure and relatedness - that can be important for downstream analyses. In addition, we phased and imputed genotypes into the dataset, using computationally efficient methods combined with the Haplotype Reference Consortium (HRC) and UK10K haplotype resource. This increases the number of testable variants by over 100-fold to ~96 million variants. We also imputed classical allelic variation at 11 human leukocyte antigen (HLA) genes, and as a quality control check of this imputation, we replicate signals of known associations between HLA alleles and many common diseases. We describe tools that allow efficient genome-wide association studies (GWAS) of multiple traits and fast phenome-wide association studies (PheWAS), which work together with a new compressed file format that has been used to distribute the dataset. As a further check of the genotyped and imputed datasets, we performed a test-case genome-wide association scan on a well-studied human trait, standing height.

genetics

Rare genetic variants in the endocannabinoid system genes CNR1 and DAGLA are associated with neurological phenotypes in humans

Rare genetic variants in the core endocannabinoid system genes CNR1, CNR2, DAGLA, MGLL and FAAH were identified in molecular testing data from up to 6.032 patients with a broad spectrum of neurological disorders. The variants were evaluated for association with phenotypes similar to those observed in the orthologous gene knockouts in mice. Heterozygous rare coding variants in CNR1, which encodes the type 1 cannabinoid receptor (CB1), were found to be significantly associated with pain sensitivity (especially migraine), sleep and memory disorders - alone or in combination with anxiety - compared to a set of controls without such CNR1 variants. Similarly, heterozygous rare variants in DAGLA, which encodes diacylglycerol lipase alpha, were found to be significantly associated with seizures and developmental disorders, including abnormalities of brain morphology, compared to controls. Rare variants in MGLL, FAAH and CNR2 were not associated with any neurological phenotypes in the patients tested. Diacylglycerol lipase alpha synthesizes the endocannabinoid 2-AG in the brain, which interacts with CB1 receptors. The phenotypes associated with rare CNR1 variants are reminiscent of those implicated in the theory of clinical endocannabinoid deficiency syndrome. The severe phenotypes associated with rare DAGLA variants underscore the critical role of rapid 2-AG synthesis and the endocannabinoid system in regulating neurological function and development. Mapping of the variants to the 3D structure of the type 1 cannabinoid receptor, or primary structure of diacylglycerol lipase alpha, reveals clustering of variants in certain structural regions and is consistent with impacts to function.

genetics

Relationships between clans and genetic kin explain cultural similarities over vast distances: the case of Yakutia

Archaeological studies sample ancient human populations one site at a time, often limited to a fraction of the regions and periods occupied by a given group. While this bias is known and discussed in the literature, few model populations span areas as large and unforgiving as the Yakuts of Eastern Siberia. We systematically surveyed 31,000 square kilometres in the Sakha Republic (Yakutia) and completed the archaeological study of 174 frozen graves, assembled between the 15th and the 19th century. We analysed genetic data (autosomal genotypes, Y-chromosome haplotypes and mitochondrial haplotypes) for all ancient subjects and confronted these to data on 190 modern subjects from the same area and the same population. Ancient familial links were identified between graves up to 1500 km apart, as well as paternal clans. We provide new insights on the origins of the contemporary Yakut population and demonstrate that cultural similarities in the past were linked to (i) the expansion of specific paternal clans, (ii) preferential marriage among the elites and (iii) funeral choices that could constitute a bias in any ancient population study.

genetics