Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,333 records · Page 74Linked to original sources

Massively scalable genetic analysis of antibody repertoires

With technical breakthroughs in the throughput and read-length of next-generation sequencing platforms, antibody repertoire sequencing is becoming an increasingly important tool for detailed characterization of the immune response. There is a need for open, scalable software for the genetic analysis of repertoire-scale antibody sequence data. To address this gap, we have developed the ab[x] package of software tools. There are three core components of the ab[x] toolkit, all of which are freely available: abcloud (github.com/briney/abcloud) for deployment and management of computational resources on Amazons Elastic Compute Cloud; abstar (github.com/briney/abstar) for pre-processing, germline gene assignment and primary annotation of antibody sequence data; and abutils (github.com/briney/abutils), which provides utilities for interactive downstream analysis of antibody repertoire data.

bioinformatics

A Haploid Genetic Screening Method for Proteins Influencing Mammalian Nonsense-Mediated mRNA Decay Activity

Despite a long appreciation for the role of nonsense-mediated mRNA decay (NMD) in the destruction of faulty, disease-causing mRNAs, as well as its role in the maintenance of normal, endogenous transcript abundance, systematic unbiased methods for uncovering modifiers of NMD activity in mammalian cells remain scant. Here we present and validate a haploid genetic screening method for identifying proteins and processes that stimulate NMD activity involving a 3'-untranslated region exon-junction complex. This reporterbased screening method can be adapted for interrogating other pathways whose output can be measured by the intracellular production of fluorescent proteins.

molecular biology

An agent-based 3D model of non-genetic adaptation in cancer tissues under electrical, mechanical, and hypoxic stress

Non-genetic adaptation enables cancer cells to alter their phenotype under stress without requiring new mutations. However, the mechanisms by which electrical, mechanical, and hypoxic cues combine to shape this process in 3D tissues remain poorly understood. This work presents an agent-based tumor model that integrates vascular oxygen supply, a globally imposed electric field, mechanically mediated crowding and compression cues, phenotype transitions, cell growth, mitosis, death, and inheritance of adaptive memory across division. The simulated tumors exhibit a three-stage trajectory consisting of necrosis onset, transient collapse of live mass, and partial regrowth accompanied by progressive accumulation of adapted cells. Continuous electrical stimulation produces a dose-dependent reduction in live mass while markedly increasing the adapted fraction, with comparatively limited changes in final necrotic burden. This response is strongly conditioned by mechanics and reshapes (and is reshaped by) adaptive capacity. Pulsed stimulation further shows that, in the model, electric field amplitude and temporal schedule jointly determine memory phenomena, phenotypic diversification, and growth recovery. These results show that coupling local oxygen availability, mechanical constraints, electrical forcing, and history-dependent phenotype transitions can generate distinct tissue-level patterns of phenotypic heterogeneity. Both stimulus magnitude and temporal protocol influenced the resulting population structure, suggesting that the history of physical stress may be an important determinant of adaptive dynamics in spatially organized tumor models.

biophysics

Human genetics and clinical aspects of neurodevelopmental disorders

Introduction Introduction Clinical classifications and the... De novo mutations, germline... Rare and compensatory mutations Current ability / approaches Prenatal diagnosis,... Implications for acceptance,... Conclusions References \"our incomplete studies do not permit actual classification; but it is better to leave things by themselves rather than to force them into classes which have their foundation only on paper\" -- Edouard Seguin (Seguin, 1866)\n\n\"The fundamental mistake which vitiates all work based upon Mendels method is the neglect of ancestry, and the attempt to regard the whole effect upon offspring, produced by a particular parent, as due to the existence in the parent of particular structural c ...

Genetics

Genetic predictors of gene expression associated with risk of bipolar disorder

Bipolar disorder (BD) affects the quality of life of approximately 1% of the population and represents a major public health concern. It is known to be highly heritable but large-scale genome-wide association studies (GWAS) have discovered only a handful of markers associated with the disease. Furthermore, the biological mechanisms underlying these markers need to be elucidated. We recently published a gene-level association test, PrediXcan that integrates transcriptome regulation data to characterize the function of these markers in a tissue specific manner. In this study, we developed prediction models for mRNA levels in 10 brain regions using data from the GTEx project and performed PrediXcan analysis in WTCCC as well as in an independent cohort, GAIN. We replicate the association between predicted expression of PTPRE and BD risk in whole blood and recapitulate the association in brain tissues. PTPRE encodes the protein tyrosine phosphatase, receptor type E, that is known to be involved in RAS signaling and activation of voltage-gated K+ channels. We also found a new genome-wide significant association between lower predicted expression of BBX (bobby sox homolog) in the anterior cingulate cortex region of the brain and increased risk of BD (pWTCCC = 7.02 x 10-6, pGAIN = 4.68 x 10-3, pmeta = 1.11 x 10-7). In sum, we used our mechanistically informed approach, PrediXcan, to identify and replicate two novel genome-wide significant genes using existing GWAS studies.

Genetics

Integrative genetic and epigenetic analysis uncovers regulatory mechanisms of autoimmune disease

Genome-wide association studies in autoimmune and inflammatory diseases (AID) have uncovered hundreds of loci mediating risk1,2. These associations are preferentially located in non-coding DNA regions3,4 and in particular to tissue-specific Dnase I hypersensitivity sites (DHS)5,6. Whilst these analyses clearly demonstrate the overall enrichment of disease risk alleles on gene regulatory regions, they are not designed to identify individual regulatory regions mediating risk or the genes under their control, and thus uncover the specific molecular events driving disease risk. To do so we have departed from standard practice by identifying regulatory regions which replicate across samples, and connect them to the genes they control through robust re-analysis of public data. We find substantial evidence of regulatory potential in 132/301 (44%) risk loci across nine autoimmune and inflammatory diseases, and are able to prioritize a single gene in 104/132 (79%) of these. Thus, we are able to generate testable mechanistic hypotheses of the molecular changes that drive disease risk.

Genetics

Genetic Mapping by Bulk Segregant Analysis in Drosophila: Experimental Design and Simulation-Based Inference

Identifying the genomic regions that underlie complex phenotypic variation is a key challenge in modern biology. Many approaches to quantitative trait locus mapping in animal and plant species suffer from limited power and genomic resolution. Here, I investigate whether bulk segregant analysis (BSA), which has been successfully applied for yeast, may have utility in the genomic era for trait mapping in Drosophila (and other organisms that can be experimentally bred in similar numbers). I perform simulations to investigate the statistical signal of a quantitative trait locus (QTL) in a wide range of BSA and introgression mapping (IM) experiments. BSA consistently provides more accurate mapping signals than IM (in addition to allowing the mapping of multiple traits from the same experimental population). The performance of BSA and IM is maximized by having multiple independent crosses, more generations of interbreeding, larger numbers of breeding individuals, and greater genotyping effort, but is less affected by the proportion of individuals selected for phenotypic extreme pools. I also introduce a prototype analysis method for Simulation-based Inference for BSA Mapping (SIBSAM). This method identifies significant QTLs and estimates their genomic confidence intervals and relative effect sizes. Importantly, it also tests whether overlapping peaks should be considered as two distinct QTLs. This approach will facilitate improved trait mapping in Drosophila and other species for which hundreds or thousands of offspring (but not millions) can be studied.

Genetics

Codon usage is a stochastic process across genetic codes of the kingdoms of life

DNA encodes protein primary structure using 64 different codons to specify 20 different amino acids and a stop signal. To uncover rules of codon use, ranked codon frequencies have previously been analyzed in terms of empirical or statistical relations for a small number of genomes. These descriptions fail on most genomes reported in the Codon Usage Tabulated from GenBank (CUTG) database. Here we model codon usage as a random variable. This stochastic model provides accurate, one-parameter characterizations of 2210 nuclear and mitochondrial genomes represented with > 104 codons/genome in CUTG. We show that ranked codon frequencies are well characterized by a truncated normal (Gaussian) distribution. Most genomes use codons in a nearuniform manner. Lopsided usages are also widely distributed across genomes but less frequent. Our model provides a universal framework for investigating determinants of codon use.

Genetics

How cognitive genetic factors influence fertility outcomes: A mediational SEM analysis.

Utilizing a newly released cognitive Polygenic Score (PGS) from Wave IV of Add Health (n = 1,886), structural equation models (SEMs) examining the relationship between PGS and fertility (which is approximately 50% complete in the present sample), utilizing measures of verbal IQ and educational attainment as potential mediators, were estimated. The results of indirect pathway models revealed that verbal IQ mediates the positive relationship between PGS and educational attainment, and educational attainment in turn mediates the negative relationship between IQ and a latent fertility measure. The direct path from PGS to fertility was non-significant. The model was robust to controlling for age, sex and race, furthermore the results of a multi-group SEM revealed no significant differences in the estimated path coefficients across sex. These results indicate that those predisposed towards higher IQ by virtue of higher PGS values are also predisposed towards trading fertility against time spent in education, which contributes to those with higher PGS values producing fewer offspring.

Genetics

Cassava HapMap: Managing genetic load in a clonal crop species

Cassava (Manihot esculenta Crantz) is an important staple food crop in Africa and South America, however, ubiquitous deleterious mutations may severely reduce its fitness. To evaluate these deleterious mutations in the cassava genome, we constructed a cassava haplotype map using deep sequencing from 241 diverse accessions and identified over 28 million segregating variants. We found that, 1) while domestication modified starch and ketone metabolism pathways for human consumption, the concomitant bottleneck and clonal propagation resulted in a large proportion of fixed deleterious amino acid changes, raised the number of deleterious mutations by 26%, and shifted the mutational burden towards common variants; 2) deleterious mutations are ineffectively purged due to limited recombination in cassava genome; 3) recent breeding efforts maintained the yield by masking the most damaging recessive mutations in the heterozygous state, but unable to purge the mutation burden, which should be a key target for future cassava breeding.

Genetics

The population genetics of human disease: the case of recessive, lethal mutations

Do the frequencies of disease mutations in human populations reflect a simple balance between mutation and purifying selection? What other factors shape the prevalence of disease mutations? To begin to answer these questions, we focused on one of the simplest cases: recessive mutations that alone cause lethal diseases or complete sterility. To this end, we generated a hand-curated set of 417 Mendelian mutations in 32 genes, reported to cause a recessive, lethal Mendelian disease. We then considered analytic models of mutation-selection balance in infinite and finite populations of constant sizes and simulations of purifying selection in a more realistic demographic setting, and tested how well these models fit allele frequencies estimated from 33,370 individuals of European ancestry. In doing so, we distinguished between CpG transitions, which occur at a substantially elevated rate, and three other mutation types. The observed frequency for CpG transitions is slightly higher than expectation but close, whereas the frequencies observed for the three other mutation types are an order of magnitude higher than expected. This discrepancy is even larger when subtle fitness effects in heterozygotes or lethal compound heterozygotes are taken into account. In principle, higher than expected frequencies of disease mutations could be due to widespread errors in reporting causal variants, compensation by other mutations, or balancing selection. It is unclear why these factors would have a greater impact on variants with lower mutation rates, however. We argue instead that the unexpectedly high frequency of disease mutations and the relationship to the mutation rate likely reflect an ascertainment bias: of all the mutations that cause recessive lethal diseases, those that by chance have reached higher frequencies are more likely to have been identified and thus to have been included in this study. Beyond the specific application, this study highlights the parameters likely to be important in shaping the frequencies of Mendelian disease alleles.\n\nAuthor SummaryWhat determines the frequencies of disease mutations in human populations? To begin to answer this question, we focus on one of the simplest cases: mutations that cause completely recessive, lethal Mendelian diseases. We first review theory about what to expect from mutation and selection in a population of finite size and further generate predictions based on simulations using a realistic demographic scenario of human evolution. For a highly mutable type of mutations, such as transitions at CpG sites, we find that the predictions are close to the observed frequencies of recessive lethal disease mutations. For less mutable types, however, predictions substantially under-estimate the observed frequency. We discuss possible explanations for the discrepancy and point to a complication that, to our knowledge, is not widely appreciated: that there exists ascertainment bias in disease mutation discovery. Specifically, we suggest that alleles that have been identified to date are likely the ones that by chance have reached higher frequencies and are thus more likely to have been mapped. More generally, our study highlights the factors that influence the frequencies of Mendelian disease alleles.

genetics

Genetic Epidemiology And Mendelian Randomization For Informing Disease Therapeutics: Conceptual And Methodological Challenges

The past decade has been proclaimed as a hugely successful era of gene discovery through the high yields of many genome-wide association studies (GWAS). However, much of the perceived benefit of such discoveries lies in the promise that the identification of genes that influence disease would directly translate into the identification of potential therapeutic targets (1-4), but this has yet to be realised at a level reflecting expectation. One reason for this, we suggest, is that GWAS to date have generally not focused on phenotypes that directly relate to the progression of disease, and thus speak to disease treatment.

genetics

A Versatile Genetic Tool For Post-Translational Control Of Gene Expression With A Small Molecule In Drosophila melanogaster

Several techniques have been developed in Drosophila to control gene expression temporally. While some of these techniques are incompatible with existing GAL4 lines, others suffer from side effects on physiology or behavior. Here, we describe a method of post-translational temporal control of gene expression which is compatible with the current library of transgenic reagents. We adopted a strategy to regulate protein degradation by fusing a protein of interest to a destabilizing domain (DD) derived from the Escherichia coli dihydrofolate reductase (ecDHFR). Trimethoprim (TMP), a stabilizing small molecule, binds to DD and blocks degradation of the chimeric protein. With a GFP-DD reporter, we show that this system is effective across different tissues and developmental stages in the fly. Notably, feeding flies with TMP can increase the expression level of GFP-DD up to 34 times in a dosage-dependent and reversible manner without altering the lifespan or behavior of the animal. To broaden the utility of our method, we engineered GAL80-DD flies that can be crossed to the available GAL4 lines to control the temporal pattern of gene expression with TMP. We also developed an inducible recombinase, FLP-DD, for high-efficiency sparse labeling and intersectional lineage analysis. Finally, we demonstrated the utility of the DD system in manipulating neuronal activity of sensory neurons. In summary, we have developed a system to control in vivo gene expression levels with negligible background, large dynamic range, and in a reversible manner, all by feeding a small molecule to Drosophila melanogaster.

genetics

Meta-analysis of exome array data identifies six novel genetic loci for lung function

Over 90 regions of the genome have been associated with lung function to date, many of which have also been implicated in chronic obstructive pulmonary disease (COPD). We carried out meta-analyses of exome array data and three lung function measures: forced expiratory volume in one second (FEV1), forced vital capacity (FVC) and the ratio of FEV1 to FVC (FEV1/FVC). These analyses by the SpiroMeta and CHARGE consortia included 60,749 individuals of European ancestry from 23 studies, and 7,721 individuals of African Ancestry from 5 studies in the discovery stage, with follow-up in up to 111,556 independent individuals. We identified significant (P<2{middle dot}8x10-7) associations with six SNPs: a nonsynonymous variant in RPAP1, which is predicted to be damaging, three intronic SNPs (SEC24C, CASC17 and UQCC1) and two intergenic SNPs near to LY86 and FGF10. eQTL analyses found evidence for regulation of gene expression at three signals and implicated several genes including TYRO3 and PLAU. Further interrogation of these loci could provide greater understanding of the determinants of lung function and pulmonary disease.

genetics

Mechanistic view and genetic control of DNA recombination during meiosis

Meiotic recombination is essential for fertility and allelic shuffling. Canonical recombination models fail to capture the observed complexity of meiotic recombinants. Here we revisit these models by analyzing meiotic heteroduplex DNA tracts genome-wide in combination with meiotic DNA double-strand break (DSB) locations. We provide unprecedented support to the synthesis-dependent strand annealing model and establish estimates of its associated template switching frequency and polymerase processivity. We show that resolution of double Holliday junctions (dHJs) is biased toward cleavage of the pair of strands containing newly synthesized DNA near the junctions. The suspected dHJ resolvase Mlh1-3 as well as Mlh1-2, Exo1 and Sgs1 promote asymmetric positioning of crossover intermediates relative to the initiating DSB and bidirectional conversions. Finally, we show that crossover-biased dHJ resolution depends on Mlh1-3, Exo1, Msh5 and to a lesser extent on Sgs1. These properties are likely conserved in eukaryotes containing the ZMM proteins, which includes mammals.

genetics

The genetic architecture of osteoarthritis: insights from UK Biobank

Osteoarthritis is a common complex disease with huge public health burden. Here we perform a genome-wide association study for osteoarthritis using data across 16.5 million variants from the UK Biobank resource. Following replication and meta-analysis in up to 30,727 cases and 297,191 controls, we report 9 new osteoarthritis loci, in all of which the most likely causal variant is non-coding. For three loci, we detect association with biologically-relevant radiographic endophenotypes, and in five signals we identify genes that are differentially expressed in degraded compared to intact articular cartilage from osteoarthritis patients. We establish causal effects for higher body mass index, but not for triglyceride levels or type 2 diabetes liability, on osteoarthritis.

genetics

Evolutionary Genetics of a Disease Susceptibility Locus in CDHR3

Selective pressures imposed by pathogens have varied among human populations throughout their evolution, leading to marked inter-population differences at some genes mediating susceptibility to infectious and immune-related diseases. A common polymorphism resulting in a C529 versus T529 change in the Cadherin-Related Family Member 3 (CDHR3) receptor is associated with rhinovirus-C (RV-C) susceptibility and severe childhood asthma. Given the morbidity and mortality associated with RV-C dependent respiratory infections and asthma, we hypothesized that the protective variant has been under selection in the human population. Supporting this idea, a recent cross-species outbreak of RV-C among chimpanzees in Uganda, which carry the ancestral risk allele at this position, resulted in a mortality rate of 8.9%. Using publicly available genomic data, we sought to determine the evolutionary history and role of selection acting on this infectious disease susceptibility locus. The protective variant is the derived allele and is found at high frequency worldwide, with the lowest relative frequency in African populations and highest in East Asian populations. There is minimal population structure among haplotypes, and we detect genomic signatures consistent with a rapid increase in frequency of the protective allele across all human populations. However, given strong evidence that the protective allele arose in anatomically modern humans prior to their migrations out of Africa and that the allele has not fixed in any population, the patterns observed here are not consistent with a classical selective sweep. We hypothesize that patterns may indicate frequency-dependent selection worldwide. Irrespective of the mode of selection, our analyses show the derived allele has been subject to selection in recent human evolution.

genetics

Genetic Identification of Novel Separase regulators in Caenorhabditis elegans

Separase is a highly conserved protease required for chromosome segregation. Although observations that separase also regulates membrane trafficking events have been made, it is still not clear how separase achieves this function. Here we present an extensive ENU mutagenesis suppressor screen aimed at identifying suppressors of sep-1(e2406), a temperature sensitive maternal effect embryonic lethal separase mutant. We screened nearly a million haploid genomes, and isolated sixty-eight suppressed lines. We identified fourteen independent intragenic sep-1(e2406) suppressed lines. These intragenic alleles map to seven SEP-1 residues within the N-terminus, compensating for the original mutation within the poorly conserved N-terminal domain. Interestingly, 47 of the suppressed lines have novel mutations throughout the entire coding region of the pph-5 phosphatase, indicating that this is an important regulator of separase. We also found that a mutation near the MEEVD motif of HSP-90, which binds and activates PPH-5, also rescues sep-1(e2406) mutants. Finally, we identified six potentially novel suppressor lines that fall into five complementation groups. These new alleles provide the opportunity to more exhaustively investigate the regulation and function of separase.

genetics