Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71Linked to original sources

Genetic variant pathogenicity prediction trained using large-scale disease specific clinical sequencing datasets

Recent advances in DNA sequencing technologies have expanded our understanding of the molecular underpinnings for several genetic disorders, and increased the utilization of genomic tests by clinicians. Given the paucity of evidence to assess each variant, and the difficulty of experimentally evaluating a variants clinical significance, many of the thousand variants that can be generated by clinical tests are reported as variants of unknown clinical significance. However, the creation of population-scale variant databases can significantly improve clinical variant interpretation. Specifically, pathogenicity prediction for novel missense variants can now utilize features describing regional variant constraint. Constrained genomic regions are those that have an unusually low variant count in the general population. Several computational methods have been introduced to capture these regions and incorporate them into pathogenicity classifiers, but these methods have yet to be compared on an independent clinical variant dataset. Here we introduce one variant dataset derived from clinical sequencing panels, and use it to compare the ability of different genomic constraint metrics to determine missense variant pathogenicity. This dataset is compiled from 17,071 patients surveyed with clinical genomic sequencing for cardiomyopathy, epilepsy, or RASopathies. We further utilize this dataset to demonstrate the necessity of disease-specific classifiers, and to train PathoPredictor, a disease-specific ensemble classifier of pathogenicity based on regional constraint and variant level features. PathoPredictor achieves an average precision greater than 90% for variants from all 99 tested disease genes while approaching 100% accuracy for some genes. Accumulation of larger clinical variant datasets and their utilization to train existing pathogenicity metrics can significantly enhance their performance in a disease and gene-specific manner.

genomics

Disentangling the effects of genetic architecture, mutational bias and selection on evolutionary forecasting

Predicting evolutionary change poses numerous challenges. Here we take advantage of the model bacterium Pseudomonas fluorescens in which the genotype-to-phenotype map determining evolution of the adaptive \"wrinkly spreader\" (WS) type is known. We present mathematical descriptions of three necessary regulatory pathways and use these to predict both the rate at which each mutational route is used and the expected mutational targets. To test predictions, mutation rates and targets were determined for each pathway. Unanticipated mutational hotspots caused experimental observations to depart from predictions but additional data led to refined models. A mismatch was observed between the spectra of WS-causing mutations obtained with and without selection due to low fitness of previously undetected WS-causing mutations. Our findings contribute toward the development of mechanistic models for forecasting evolution, highlight current limitations, and draw attention to challenges in predicting locus-specific mutational biases and fitness effects.\n\nImpact statementA combination of genetics, experimental evolution and mathematical modelling defines information necessary to predict the outcome of short-term adaptive evolution.

evolutionary biology

Agonistic character displacement of genetically based male color patterns across darters

A growing number of studies are demonstrating that interspecific male-male competitive interactions can promote trait evolution and contribute to speciation. Agonistic character displacement (ACD) occurs when selection to avoid maladaptive interspecific aggression leads to the evolution of agonistic signals and/or associated behavioral biases in sympatry. Here we test for a pattern consistent with ACD in male color pattern in darters (Percidae: Etheostoma). Despite the presence of traditional sex roles and sexual dimorphism, male color pattern has been shown to function in male-male competition in several darter species, rather than female mating preferences. A pattern consistent with divergent ACD in male behavioral biases has also been documented in darters. Males bias their aggression towards conspecific over heterospecific males in sympatry but not in allopatry. Here we use a common garden approach to show that differences in male color pattern among four closely related darter species are genetically based. We also demonstrate that male color pattern exhibits enhanced differences in sympatric compared to allopatric populations of two darter species. This study provides evidence that interspecific male-male aggressive interactions alone can promote elaborate male signal evolution both between and within species. We discuss the implications this has for male-driven ACD and cascade ACD.

evolutionary biology

Pancreatic adenocarcinoma human organoids share structural and genetic features with primary tumors

Patient-derived pancreatic ductal adenocarcinoma (PDAC) organoid systems show great promise for understanding the biological underpinnings of disease and advancing therapeutic precision medicine. Despite the increased use of organoids, the fidelity of molecular features, genetic heterogeneity, and drug response to the tumor of origin remain important unanswered questions limiting their utility. To address this gap in knowledge, we created primary tumor- and PDX-derived organoids, and 2D cultures for in-depth genomic and histopathological comparisons to the primary tumor. Histopathological features and PDAC representative protein markers showed strong concordance. DNA and RNA sequencing of single organoids revealed patient-specific genomic and transcriptomic consistency. Single-cell RNAseq demonstrated that organoids are primarily a clonal population. In drug response assays, organoids displayed patient-specific sensitivities. Additionally, we examined the in vivo PDX response to FOLFIRINOX and Gemcitabine/Abraxane treatments, which was recapitulated in vitro by organoids. The patient-specific molecular and histopathological fidelity of organoids indicate that they can be used to understand the etiology of the patients tumor and the differential response to therapies and suggests utility for predicting drug responses.

cancer biology

Chromatin interactions and expression quantitative trait loci reveal genetic drivers of multimorbidities

Clinical studies of non-communicable diseases identify multimorbidities that reflect our relatively limited fixed metabolic capacity. Despite the fact that we have [~]24000 genes, we do not understand the genetic pathways that contribute to the development of multimorbid non-communicable disease. We created a \"multimorbidity atlas\" of traits based on pleiotropy of spatially regulated genes using convex biclustering. Using chromatin interaction and expression Quantitative Trait Loci (eQTL) data, we analysed 20,782 variants (p < 5 x 10-6) associated with 1,351 phenotypes, to identify 16,248 putative eQTL-eGene pairs that are involved in 76,013 short- and long-range regulatory interactions (FDR < 0.05) in different human tissues. Convex biclustering of eGenes that are shared between phenotypes identified complex inter-relationships between nominally different phenotype associated SNPs. Notably, the loci at the centre of these inter-relationships were subject to complex tissue and disease specific regulatory effects. The largest cluster, 40 phenotypes that are related to fat and lipid metabolism, inflammatory disorders, and cancers, is centred on the FADS1-FADS3 locus (chromosome 11). Our novel approach enables the simultaneous elucidation of variant interactions with genes that are drivers of multimorbidity and those that contribute to unique phenotype associated characteristics.

bioinformatics

Individual based genetic analyses support asexual hydrochory dispersal in Zostera noltei

Dispersal beyond the local patch in clonal plants was typically thought to result from sexual reproduction via seed dispersal. However, evidence for the separation, transport by water, and re-establishment of asexual propagules (asexual hydrochory) is mounting suggesting other important means of dispersal in aquatic plants. Using an unprecedented sampling size and microsatellite genetic identification, we describe the distribution of seagrass clones along tens of km within a coastal lagoon in Southern Portugal. Our spatially explicit individual-based sampling design covered 84 km2 and collected 3 185 Zostera noltei ramets from 803 sites. We estimated clone age, assuming rhizome elongation as the only mechanism of clone spread, and contrasted it with paleo-oceanographic sea level change. We also studied the association between a source of disturbance and the location of large clones. A total of 16 clones were sampled more than 10 times and the most abundant one was sampled 59 times. The largest distance between two samples from the same clone was 26.4 km and a total of 58 and 10 clones were sampled across more than 2 and 10 km, respectively. The number of extremely large clone sizes, and their old ages when assuming the rhizome elongation as the single causal mechanism, suggests other processes are behind the span of these clones. We discuss how the dispersal of vegetative fragments in a stepping-stone manner might have produced this pattern. We found higher probabilities to sample large clones away from the lagoon inlet, considered a source of disturbance. This study corroborates previous experiments on the success of transport and re-establishment of asexual fragments and supports the hypothesis that asexual hydrochory is responsible for the extent of these clones.

ecology

Inferring time-dependent migration and coalescence patterns from genetic sequence and predictor data in structured populations

Population dynamics can be inferred from genetic sequence data using phylodynamic methods. These methods typically quantify the dynamics in unstructured populations or assume the parameters describing the dynamics to be constant through time in structured populations. Inference methods allowing for structured populations and parameters to vary through time involve many parameters which have to be inferred. Each of these parameters might be however only weakly informed by data. Here we introduce an approach that uses so-called predictors, such as geographic distance between locations, within a generalized linear model to inform the population dynamic parameters, namely the time-varying migration rates and effective population sizes under the marginal approximation of the structured coalescent. By using simulations, we show that we are able to reliably infer the parameters from phylogenetic trees. We then apply this framework to a previously described Ebola virus dataset. We infer incidence to be the strongest predictor for effective population size and geographic distance the strongest predictor for migration. This allows us to show not only on simulated data, but also on real data, that we are able to identify reasonable predictors. Overall, we provide a novel method that allows to identify predictors for migration rates and effective population sizes and to use these predictors to quantify migration rates and effective population sizes. Its implementation as part of the BEAST2 software package MASCOT allows to jointly infer population dynamics within structured populations, the phylogenetic tree, and evolutionary parameters.

evolutionary biology

New genetic signals for lung function highlight pathways and pleiotropy, and chronic obstructive pulmonary disease associations across multiple ancestries

Reduced lung function predicts mortality and is key to the diagnosis of COPD. In a genome-wide association study in 400,102 individuals of European ancestry, we define 279 lung function signals, one-half of which are new. In combination these variants strongly predict COPD in deeply-phenotyped patient populations. Furthermore, the combined effect of these variants showed generalisability across smokers and never-smokers, and across ancestral groups. We highlight biological pathways, known and potential drug targets for COPD and, in phenome-wide association studies, autoimmune-related and other pleiotropic effects of lung function associated variants. This new genetic evidence has potential to improve future preventive and therapeutic strategies for COPD.

genomics

Switchable genome editing via genetic code expansion

Multiple applications of genome editing by CRISPR-Cas9 necessitate stringent regulation and Cas9 variants have accordingly been generated whose activity responds to small ligands, temperature or light. However, these approaches are often impracticable, for example in clinical therapeutic genome editing in situ or gene drives in which environmentally-compatible control is paramount. With this in mind, we have developed heritable Cas9-mediated mammalian genome editing that is acutely controlled by the cheap lysine derivative, Lys(Boc) (BOC). Genetic code expansion permitted non-physiological BOC incorporation such that Cas9 (Cas9BOC) was expressed in a full-length, active form in cultured somatic cells only after BOC exposure. Stringently BOC-dependent, heritable editing of transgenic and native genomic loci occurred when Cas9BOC was expressed at the onset of mouse embryonic development from cRNA or Cas9BOC transgenic females. The tightly controlled Cas9 editing system reported here promises to have broad applications and is a first step towards purposed, spatiotemporal gene drive regulation over large geographical ranges.

developmental biology

SNP Selection and Concordance in Consumer Genetics Testing

The use of Direct To Consumer (DTC) genetic testing for predicting health risks and a variety of other phenotypes has been extensively discussed. Additionally, there have been wide ranging discourses on privacy and ethical concerns. Much less attention has been paid to what most people actually use DTC testing for: ancestry determination. Furthermore, comparison of the platforms used by different companies and how they have chosen SNPs to address the questions of health and ancestry have not been broadly reported. When SNPs across three genotyping platforms are compared, only 16-18% of SNPs with reported genotypes are shared across all platforms. Only 110,051 of the more than 600,000 SNPs are called on all three panels examined (Ancestry, 23andMe and MyHeritage). SNPs genotyped on all platforms are highly concordant with only two SNPs having discordant calls. When the SNPs unique to a single panel are examined, it is apparent that each company has its own strategy for choosing SNPs. When each platform is examined, the unique SNPs have different frequencies, ethnic selectivities, and chromosomal locations. Because each company separates the world into different, overlapping geographical regions, it is impossible to do an exact comparison of ancestry results. Factoring in the ways the regions overlap, congruent results are generated for the major contributors to ancestry.

genomics

Development of reduced gluten wheat enabled by determination of the genetic basis of the lys3a low hordein barley mutant

Celiac disease is the most common food-induced enteropathy in humans with a prevalence of approximately 1% world-wide [1]. It is induced by digestion-resistant, proline- and glutamine-rich seed storage proteins, collectively referred to as \"gluten,\" found in wheat. Related prolamins are present in barley and rye. Both celiac disease and a related condition called non-celiac gluten sensitivity (NCGS) are increasing in incidence [2] [3]. This has prompted efforts to identify methods of lowering gluten in wheat, one of the most important cereal crops. Here we used BSR-seq (Bulked Segregant RNA-seq) and map-based cloning to identify the genetic lesion underlying a recessive, low prolamin mutation (lys3a) in diploid barley. We confirmed the mutant identity by complementing the lys3a mutant with a transgenic copy of the wild type barley gene and then used TILLING (Targeting Induced Local Lesions in Genomes) [4] to identify induced SNPs (Single Nucleotide Polymorphisms) in the three homoeologs of the corresponding wheat gene. Combining inactivating mutations in the three sub-genomes of hexaploid bread wheat in a single wheat line lowered gliadin and low molecular weight glutenin accumulation by 50-60% and increased free and protein-bound lysine by 33%. This is the first report of the combination of mutations in homoeologs of a single gene that reduces gluten in wheat.

plant biology

Cell-intrinsic genetic regulation of peripheral memory-phenotype T cell frequencies

Memory T and B lymphocyte numbers are thought to be regulated by recent and cumulative microbial exposures. We report here that memory-phenotype lymphocyte frequencies in B, CD4 and CD8 T-cells in 3-monthly serial bleeds from healthy young adult humans were relatively stable over a 1-year period, while recently activated -B and -CD4 T cell frequencies were not, suggesting that recent environmental exposures affected steady state levels of recently activated but not of memory lymphocyte subsets. Frequencies of memory B and CD4 T cells were not correlated, suggesting that variation in them was unlikely to be determined by cumulative antigenic exposures. Immunophenotyping of adult siblings showed high concordance in memory, but not of recently activated lymphocyte subsets, suggesting genetic regulation of memory lymphocyte frequencies. To explore this possibility further, we screened effector memory (EM)-phenotype T cell frequencies in common independent inbred mice strains. Using two pairs from these strains that differed predominantly in either CD4EM and/or CD8EM frequencies, we constructed bi-parental bone marrow chimeras in F1 recipient mice, and found that memory T cell frequencies in recipient mice were determined by donor genotypes. Together, these data suggest cell-autonomous determination of memory T niche size, and suggest mechanisms maintaining immune variability.

immunology

MODE for detecting and estimating genetic causal variants

Determining the genetic causal variants and estimating their effect sizes are considered to be correlated but independent problems. Fine-mapping studies often rely on the ability to integrate useful functional annotation information into genome wide association univariate/multivariate analysis. In the present study, by modeling the probability of a SNP being causal and its effect size as a set of correlated Gaussian/non-Gaussian random variables, we design an optimization routine for simultaneous fine-mapping and effect size estimation. The algorithm is released as an open source C package MODE.\n\nAvailability and Implementation: http://sites.google.com/site/sundarvelkur/mode\n\nContact: amdale@ucsd.edu, svelkur@ucsd.edu

bioinformatics

Functional and genetic characterization of an incF-type multidrug resistance plasmid isolated from fresh spinach.

The presence of antibiotic-resistant bacteria and clinically-relevant antibiotic resistance genes within raw foods is an on-going food safety concern. It is particularly important to be aware of the microbial quality of fresh produce because foods such as leafy greens including lettuce and spinach are minimally processed and often consumed raw therefore they often lack a microbial inactivation step. This study characterizes the genetic and functional aspects of a mobile, multidrug resistance plasmid, pLGP4, isolated from fresh spinach bought from a farmers market. pLGP4 was isolated using a bacterial conjugation approach. The functional characteristics of the plasmid were determined using multidrug resistance profiling and plasmid stability assays. pLGP4 was resistant to six of the eight antibiotics tested and included ciprofloxacin and meropenem. The plasmid was stably maintained within host strains in the absence of an antibiotic selection. The plasmid DNA was sequenced using an Illumina MiSeq high throughput sequencing approach and assembled into contigs using SPAdes. PCR mapping and Sanger DNA sequencing of PCR amplicons was used to complete the plasmid DNA sequence. Comparative sequence analysis determined that the plasmid was similar to plasmids that have been frequently associated with multidrug resistant clinical isolates of Klebsiella spp. DNA sequence analysis showed pLGP4 harboured qnrB1 and several other antibiotic resistance genes including three {beta}-lactamases: blaTEM-1, blaCTX-M-15 and blaOXA-1. The detection of a multidrug-resistant, clinically-relevant plasmid on fresh spinach emphasizes the importance for vegetable producers to implement evidence-based food safety approaches into their production practises to ensure the food safety of leafy greens.

microbiology

Efficient Synthesis of Mutants Using Genetic Crosses

The genetic cross is a fundamental, flexible, and widely-used experimental technique to create new mutant strains from existing ones. Surprisingly, the problem of how to efficiently compute a sequence of crosses that can make a desired target mutant from a set of source mutants has received scarce attention. In this paper, we make three contributions to this question.\n\nFirst, we formulate several natural problems related to efficient synthesis of a target mutant from source mutants. Our formulations capture experimentally-useful notions of verifiability (e.g the need to confirm that a mutant contains mutations in the desired genes) and permissibility (e.g., the requirement that no intermediate mutants in the synthesis be inviable).\n\nSecond, we develop combinatorial techniques to solve these problems. We prove that checking the existence of a verifiable, permissible synthesis is NP-complete in general. We complement this result with three polynomial time or fixed-parameter tractable algorithms for optimal synthesis of a target mutant for special cases of the problem that arise in practice.\n\nThird, we apply these algorithms to simulated data and to synthetic data. We use results from simulations of a mathematical model of the cell cycle to replicate realistic experimental scenarios where a biologist may be interested in creating several mutants in order to verify model predictions. Our results show that the consideration of permissible mutants can affect the existence of a synthesis or the number of crosses in an optimal one. Our algorithms gracefully handle the restrictions that permissible mutants impose. Results on synthetic data show that our algorithms scale well with increases in the size of the input and the fixed parameters.

systems biology

Development of a genetically encoded sensor for endogenous CaMKII activity

CaMKII is a crucial oligomeric enzyme in neuronal and cardiac signaling, fertilization and immunity. Here, we report the construction of a novel, substrate-based, genetically-encoded sensor for CaMKII activity, FRESCA (FRET-based Sensor for CaMKII Activity). Currently, there is one biosensor for CaMKII activity, Camui, which contains CaMKII. FRESCA allows us to measure all endogenous CaMKII variants, while Camui can track a single variant. Since there are ~40 CaMKII variants, using FRESCA to measure aggregate activity allows a fresh perspective on CaMKII activity. We show, using live-cell imaging, FRESCA response is concurrent with Ca2+ rises in HEK293T cells and mouse eggs. In eggs, we stimulate oscillatory patterns of Ca2+ and observe the differential responses of FRESCA and Camui. Our results implicate an important role for the variable linker region in CaMKII, which tunes its activation. FRESCA will be a transformative tool for studies in neurons, cardiomyocytes and other CaMKII-containing cells.

biochemistry

Genetic, inflammatory, and tissue-specific factors control expression of human calpain-14

Eosinophilic esophagitis (EoE) is a chronic, food-driven allergic disease resulting in eosinophilic esophageal inflammation. We recently found that EoE susceptibility is associated with genetic variants in the promoter of CAPN14, a gene with reported esophagus-specific expression. CAPN14 is dynamically up-regulated as a function of EoE disease activity and after exposure of epithelial cells to interleukin-13 (IL-13). Herein, we aimed to explore molecular modulation of CAPN14 expression. We identified three putative binding sites for the IL-13-activated transcription factor STAT6 in the promoter and first intron of CAPN14. Luciferase reporter assays revealed that the two most distal STAT6 elements were required for the ~10-fold increase in promoter activity subsequent to stimulation with IL-13 or IL-4, and also for the genotype-dependent reduction in IL-13-induced promoter activity. One of the STAT6 elements in the promoter was necessary for IL-13-mediated induction of CAPN14 promoter activity while the other STAT6 promoter element was necessary for full induction. Chromatin immunoprecipitation in IL-13 stimulated esophageal epithelial cells was used to further support STAT6 binding to the promoter of CAPN14 at these STAT6 binding sites. The highest CAPN14 and calpain-14 expression occurred with IL-13 or IL-4 stimulation of esophageal epithelial cells under culture conditions that allow the cells to differentiate into a stratified epithelium. This work corroborates a candidate molecular mechanism for EoE disease etiology in which the risk variant at 2p23 dampens mediated CAPN14 expression in differentiated esophageal epithelial cells following IL-13/STAT6 induction of CAPN14 promoter activity.

genomics

Probabilistic ancestry maps: a method to assess and visualize population substructures in genetics

Principal component analysis (PCA) is a standard method to correct for population stratification in ancestry-specific genome-wide association studies (GWASs) and is used to cluster individuals by ancestry. Using the 1000 genomes project data, we examine how non-linear dimensionality reduction methods such as t-distributed stochastic neighbor embedding (t-SNE) or generative topographic mapping (GTM) can be used to provide improved ancestry maps by accounting for a higher percentage of explained variance in ancestry, and how they can help to estimate the number of principal components necessary to account for population stratification. GTM also generates posterior probabilities of class membership which can be used to assess the probability of an individual to belong to a given population - as opposed to t-SNE, GTM can be used for both clustering and classification. This paper is a first application of GTM for ancestry classification models. Our maps and software are available online.\n\nAuthor summaryWith this paper, we seek to encourage researchers working in genetics to use other methods than PCA to visualize ancestry and identify substructures in populations. We propose to use methods which do not only allow visualization of ancestry, but also the estimation of probabilities of belonging to different ancestry groups.

bioinformatics