Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Recombination of ecologically and evolutionarily significant loci maintains genetic cohesion in the Pseudomonas syringae species complex

Pseudomonas syringae is a highly diverse bacterial species complex capable of causing a wide range of serious diseases on numerous agronomically important crop species. Here, we examine the evolutionary relationships of 391 agricultural and environmental strains from the P. syringae species complex using whole-genome sequencing and evolutionary genomic analyses. Our collection includes strains from 11 of the 13 previously described phylogroups isolated off of over 90 hosts. We describe the phylogenetic distribution of all orthologous gene families in the P. syringae pan-genome, reconstruct the phylogeny of P. syringae using a core genome alignment and a hierarchical clustering analysis of pan-genome content, predict ecologically and evolutionary relevant loci, and establish the forces of molecular evolution operating on each gene family. We find that the common ancestor of the species complex likely carried a Rhizobium-like type III secretion system (TTSS) and later acquired the canonical TTSS. The phylogenetic analysis also showed that the species complex is subdivided into primary and secondary phylogroups based on genetic diversity and rates of genetic exchange. The primary phylogroups, which largely consist of agricultural isolates, are no more divergent than a number of other bacterial species, while the secondary phylogroups, which largely consists of environmental isolates, have levels of diversity more in line with multiple distinct species within a genus. An analysis of rates of recombination within and between phylogroups revealed a higher rate of recombination within primary phylogroups than between primary and secondary phylogroups. We also found that \"ecologically significant\" virulence-associated loci and \"evolutionary significant\" loci under positive selection are over-represented among loci that undergo inter-phylogroup genetic exchange. These results indicate that while inter-phylogroup recombination occurs relatively rarely in the species complex, it is an important force of genetic cohesion, particularly among the strains in the primary phylogroups. This level of genetic cohesion and the shared plant-associated niche argues for considering the primary phylogroups as a true biological species.

genomics

Copy number variation in fungi and its implications for wine yeast genetic diversity and adaptation

In recent years, copy number (CN) variation has emerged as a new and significant source of genetic polymorphisms contributing to the phenotypic diversity of populations. CN variants are defined as genetic loci that, due to duplication and deletion, vary in their number of copies across individuals in a population. CN variants range in size from 50 base pairs to whole chromosomes, can influence gene activity, and are associated with a wide range of phenotypes in diverse organisms, including the budding yeast Saccharomyces cerevisiae. In this review, we introduce CN variation, discuss the genetic and molecular mechanisms implicated in its generation, how they can contribute to genetic and phenotypic diversity in fungal populations, and consider how CN variants may influence wine yeast adaptation in fermentation-related processes. In particular, we focus on reviewing recent work investigating the contribution of changes in CN of fermentation-related genes associated with the adaptation and domestication of yeast wine strains and offer notable illustrations of such changes, including the high levels of CN variation among the CUP genes, which confer resistance to copper, and the preferential deletion and duplication of the MALI and MAL3 loci, respectively, which are responsible for metabolizing maltose and sucrose. Based on the available data, we propose that CN variation is a substantial dimension of yeast genetic diversity that occurs largely independent of single nucleotide polymorphisms. As such, CN variation harbors considerable potential for understanding and manipulating yeast strains in the wine fermentation environment and beyond.

microbiology

A structural equation model for imaging genetics using spatial transcriptomics

Alzheimers disease is a neurodegenerative disorder that causes changes in the structure of the brain, observable with MRI scans, and that has a strong heritable component, reflected in the DNA. Imaging genetics deals with such relationships between genetic variation and imaging variables, often in a disease context. The complex relationships between brain volumes and genetic variants have been explored both with dimension reduction methods and model based approaches. However, these models usually do not make use of the extensive knowledge of the spatio-anatomical patterns of gene activity. We present a method for integrating genetic markers (single nucleotide polymorphisms) and imaging features, which is based on a causal model and, at the same time, uses the power of dimension reduction. We use structural equation models to find latent variables that explain brain volume changes in a disease context, and which are in turn affected by genetic variants. We make use of publicly available spatial transcriptome data from the Allen Human Brain Atlas to specify the model structure, which reduces noise and improves interpretability. The model is tested in a simulation setting, and applied on a case study of the Alzheimers Disease Neuroimaging Initiative.

bioinformatics

Phylogeny Recapitulates Learning: Self-Optimization of Genetic Code

Learning algorithms have been proposed as a non-selective mechanism capable of creating complex adaptive systems in life. Evolutionary learning however has not been demonstrated to be a plausible cause for the origin of a specific molecular system. Here we show that genetic codes as optimal as the Standard Genetic Code (SGC) emerge readily by following a molecular analog of the Hebbs rule (\"neurons fire together, wire together\"). Specifically, error-minimizing genetic codes are obtained by maximizing the number of physio-chemically similar amino acids assigned to evolutionarily similar codons. Formulating genetic code as a Traveling Salesman Problem (TSP) with amino acids as \"cities\" and codons as \"tour positions\" and implemented with a Hopfield neural network, the unsupervised learning algorithm efficiently finds an abundance of genetic codes that are more error-minimizing than SGC. Drawing evidence from molecular phylogenies of contemporary tRNAs and aminoacyl-tRNA synthetases, we show that co-diversification between gene sequences and gene functions, which cumulatively captures functional differences with sequence differences and creates a genomic \"memory\" of the living environment, provides the biological basis for the Hebbian learning algorithm. Like the Hebbs rule, the locally acting phylogenetic learning rule, which may simply be stated as increasing phylogenetic divergence for increasing functional difference, could lead to complex and robust life systems. Natural selection, while essential for maintaining gene function, is not necessary to act at system levels. For molecular systems that are self-organizing through phylogenetic learning, the TSP model and its Hopfield network solution offer a promising framework for simulating emerging behavior, forecasting evolutionary trajectories, and designing optimal synthetic systems.

evolutionary biology

Sex-specific additive genetic variances and correlations for fitness in a song sparrow (Melospiza melodia) population subject to natural immigration and inbreeding

Quantifying sex-specific additive genetic variance (VA) in fitness, and the cross-sex genetic correlation (rA), is pre-requisite to predicting evolutionary dynamics and the magnitude of sexual conflict. Quantifying VA and rA in underlying fitness components, and multiple genetic consequences of immigration and resulting gene flow, is required to identify mechanisms that maintain VA in fitness. However, these key parameters have rarely been estimated in wild populations experiencing natural environmental variation and immigration. We used comprehensive pedigree and life-history data from song sparrows (Melospiza melodia) to estimate VA and rA in sex-specific fitness and underlying fitness components, and to estimate additive genetic effects of immigrants as well as inbreeding depression. We found substantial VA in female and male fitness, with a moderate positive cross-sex rA. There was also substantial VA in adult reproductive success in males but not females, and moderate VA in juvenile survival but not adult survival. Immigrants introduced alleles for which additive genetic effects on local fitness were negative, potentially reducing population mean fitness through migration load, yet alleviating expression of inbreeding depression. Substantial VA for fitness can consequently be maintained in the wild, and be concordant between the sexes despite marked sex-specific VA in reproductive success.

evolutionary biology

Evaluation of Whole Exome Sequencing as an Alternative of BeadChip and Whole Genome Sequencing in Human Population Genetic Analysis

Understanding the underlying genetic structure of human populations is of fundamental interest to both biological and social sciences. Advances in high-throughput genotyping technology have markedly improved our understanding of global patterns of human genetic variation. The most widely used methods for collecting variant information at the DNA-level include whole genome sequencing, which continues to remain costly, and the more economical solution of array-based techniques, as these are capable of simultaneously genotyping a pre-selected set of variable DNA sites in the human genome. The largest publicly accessible set of human genomic sequence data available today originates from exome sequencing that comprises around 1.2% of the whole genome (approximately 30 million base pairs). In this study, we compared the application of the exome dataset to the array-based dataset and to the gold standard whole genome dataset using the same population genetic analysis methods. Our results draw attention to some of the inherent problems that arise from using pre-selected SNP sets for population genetic analysis. Additionally, we demonstrate that exome sequencing provides a better alternative to the array-based methods for population genetic analysis. In this study, we propose a strategy for unbiased variant collection from exome data and offer a bioinformatics protocol for proper data processing.

genomics

Genome wide association analysis identifies genetic variants associated with reproductive variation across domestic dog breeds and uncovers links to domestication

The diversity of eutherian reproductive strategies has led to variation in many traits, such as number of offspring, age of reproductive maturity, and gestation length. While reproductive trait variation has been extensively investigated and is well established in mammals, the genetic loci contributing to this variation remain largely unknown. The domestic dog, Canis lupus familiaris is a powerful model for studies of the genetics of inherited disease due to its unique history of domestication. To gain insight into the genetic basis of reproductive traits across domestic dog breeds, we collected phenotypic data for four traits - cesarean section rate (n = 97 breeds), litter size (n = 60), stillbirth rate (n = 57), and gestation length (n = 23) - from primary literature and breeders handbooks. By matching our phenotypic data to genomic data from the Cornell Veterinary Biobank, we performed genome wide association analyses for these four reproductive traits, using body mass and kinship among breeds as co-variates. We identified 14 genome-wide significant associations between these traits and genetic loci, including variants near CACNA2D3 with gestation length, MSRB3 with litter size, SMOC2 with cesarean section rate, MITF with litter size and still birth rate, KRT71 with cesarean section rate, litter size, and stillbirth rate, and HTR2C with stillbirth rate. Some of these loci, such as CACNA2D3 and MSRB3, have been previously implicated in human reproductive pathologies. Many of the variants that we identified have been previously associated with domestication-related traits, including brachycephaly (SMOC2), coat color (MITF), coat curl (KRT71), and tameness (HTR2C). These results raise the hypothesis that the artificial selection that gave rise to dog breeds also shaped the observed variation in their reproductive traits. Overall, our work establishes the domestic dog as a system for studying the genetics of reproductive biology and disease.

evolutionary biology

Systematic genetic mapping of necroptosis identifies SLC39A7 as modulator of death receptor trafficking

Regulation of cell and tissue homeostasis by programmed cell death is a fundamental process with wide physiological and pathological implications. The advent of scalable somatic cell genetic technologies creates the opportunity to functionally map these essential pathways, thereby identifying potential disease-relevant components. We investigated the genetic basis underlying necroptotic cell death by performing a complementary set of loss- and gain-of-function genetic screens. To this end, we established FADD-deficient haploid human KBM7 cells, which specifically and efficiently undergo necroptosis after a single treatment with either TNF or the SMAC mimetic compound birinapant. A series of unbiased gene-trap screens identified key signaling mediators, such as TNFR1, RIPK1, RIPK3, and MLKL. Among the novel components, we focused on the zinc transporter SLC39A7, whose knock-out led to necroptosis resistance by affecting TNF receptor trafficking and ER homeostasis. Orthogonal, solute carrier (SLC)-focused CRISPR/Cas9-based genetic screens revealed the exquisite specificity of SLC39A7, among ~ 400 SLC genes, for TNFR1- and FAS-but not TRAIL-R1-mediated responses. The newly established cellular model also allowed genome-wide gain-of-function screening for genes conferring resistance to necroptosis via the CRISPR/Cas9 synergistic activation mediator approach. Among these, we found cIAP1 and cIAP2, and characterized the role of TNIP1 (TNFAIP3-interacting protein 1), which prevented pathway activation in a ubiquitin-binding dependent manner. Altogether, the gain- and loss-of-function screens described here provide a global genetic chart of the molecular factors involved in necroptosis and death receptor signaling, prompting investigation of their individual contribution and potential role in pathological conditions.

cell biology

A fast and agnostic method for bacterial genome-wide association studies: bridging the gap between kmers and genetic events

MotivationGenome-wide association study (GWAS) methods applied to bacterial genomes have shown promising results for genetic marker discovery or fine-assessment of marker effect. Recently, alignment-free methods based on kmer composition have proven their ability to explore the accessory genome. However, they lead to redundant descriptions and results which are hard to interpret.\n\nMethodsHere, we introduce DBGWAS, an extended kmer-based GWAS method producing interpretable genetic variants associated with pheno-types. Relying on compacted De Bruijn graphs (cDBG), our method gathers cDBG nodes identified by the association model into subgraphs defined from their neighbourhood in the initial cDBG. DBGWAS is fast, alignment-free and only requires a set of contigs and phenotypes. It produces annotated subgraphs representing local polymorphisms as well as mobile genetic elements (MGE) and offers a graphical framework to interpret GWAS results.\n\nResultsWe validated our method using antibiotic resistance phenotypes for three bacterial species. DBGWAS recovered known resistance determinants such as mutations in core genes in Mycobacterium tuberculosis and genes acquired by horizontal transfer in Staphylococcus aureus and Pseudomonas aeruginosa - along with their MGE context. It also enabled us to formulate new hypotheses involving genetic variants not yet described in the antibiotic resistance literature.\n\nConclusionOur novel method proved its efficiency to retrieve any type of phenotype-associated genetic variant without prior knowledge. All experiments were computed in less than two hours and produced a compact set of meaningful subgraphs, thereby outperforming other GWAS approaches and facilitating the interpretation of the results.\n\nAvailabilityOpen-source tool available at https://gitlab.com/leoisl/dbgwas

bioinformatics

Ecological genetic conflict between specialism and plasticity through genomic islands of divergence

There can be genetic conflict between genome elements differing in transmission patterns, and thus in evolutionary interests. We show here that the concept of genetic conflict provides new insight into local adaptation and phenotypic plasticity. Local adaptation to heterogeneous habitats sometimes occurs as tightly linked clusters of genes with among-habitat polymorphism, referred to as genomic islands of divergence, and our work sheds light on their evolution. Phenotypic plasticity can also influence the divergence between ecotypes, through developmental responses to habitat-specificcues. We show that clustered genes coding for ecological specialism and unlinked generalist genes coding for phenotypic plasticity differ in their evolutionary interest. This is an ecological genetic conflict, operating between habitat specialism and phenotypically plastic generalism. The phenomenon occurs both for single traits and for syndromes of co-adapted traits. Using individual-based simulations and numerical analysis, we investigate how among-habitat genetic polymorphism and phenotypic plasticity depend on genetic architecture. We show that for plasticity genes that are unlinked to a genomic island of divergence, the slope of a reaction norm will be steeper in comparison with the slope favored by plasticity genes that are tightly linked to genes for local adaptation.

evolutionary biology

Identification of 12 genetic loci associated with human healthspan

The mounting challenge of preserving the quality of life in an aging population directs the focus of longevity science to the regulatory pathways controlling healthspan. To understand the nature of the relationship between the healthspan and lifespan and uncover the genetic architecture of the two phenotypes, we studied the incidence of major age-related diseases in the UK Biobank (UKB) cohort. We observed that the incidence rates of major chronic diseases increase exponentially. The risk of disease acquisition doubled approximately every eight years, i.e., at a rate compatible with the doubling time of the Gompertz mortality law. Assuming that aging is the single underlying factor behind the morbidity rates dynamics, we built a proportional hazards model to predict the risks of the diseases and therefore the age corresponding to the end of healthspan of an individual depending on their age, gender, and the genetic background. We suggested a computationally efficient procedure for the determination of the effect size and statistical significance of individual gene variants associations with healthspan in a form suitable for a Genome-Wide Association Studies (GWAS). Using the UKB sub-population of 300,447 genetically Caucasian, British individuals as a discovery cohort, we identified 12 loci associated with healthspan and reaching the whole-genome level of significance. We observed strong (|{rho}g| > 0.3) genetic correlations between healthspan and the incidence of specific age-related disease present in our healthspan definition (with the notable exception of dementia). Other examples included all-cause mortality (as derived from parental survival, with{rho} g = -0.76), life-history traits (metrics of obesity, age at first birth), levels of different metabolites (lipids, amino acids, glycemic traits), and psychological traits (smoking behaviour, cognitive performance, depressive symptoms, insomnia). We conclude by noting that the healthspan phenotype, suggested and characterized here, offers a promising new way to investigate human longevity by exploiting the data from genetic and clinical data on living individuals.

epidemiology

The Genetic Insulator RiboJ Increases Expression of Insulated Genes

The self-cleaving ribozyme RiboJ is an insulator commonly used in genetic circuits to prevent unexpected interactions between neighboring parts. These interactions can compromise the modularity of the circuit, impeding the implementation of predictable genetic constructs. Despite its utility as an insulator, a quantitative assessment of the effect of RiboJ on the properties of downstream genetic parts is lacking. Here, we characterized the impact of insulation with RiboJ on expression of a reporter gene driven by a promoter from a library of 24 frequently employed constitutive promoters. We show that depending on the strength of the promoters, insulation with RiboJ increased protein abundance between twofold and tenfold and increased transcript abundance by an average of twofold. This result is the first to demonstrate that genetic insulators can impact the expression of downstream genes, potentially hindering the design of predictable genetic circuits and constructs.

synthetic biology

Mapping the genetics of neuropsychological traits to the molecular network of the human brain using a data integrative approach

MotivationComplex neuropsychiatric conditions including autism spectrum disorders are among the most heritable neurodevelopmental disorders with distinct profiles of neuropsychological traits. A variety of genetic factors modulate these traits (phenotypes) underlying clinical diagnoses. To explore the associations between genetic factors and phenotypes, genome-wide association studies are broadly applied. Stringent quality checks and thorough downstream analyses for in-depth interpretation of the associations are an indispensable prerequisite. However, in the area of neuropsychology there is no framework existing, which besides performing association studies also affiliates genetic variants at the brain and gene network level within a single framework.\n\nResultsWe present a novel bioinformatics approach in the field of neuropsychology that integrates current state-of-the-art tools, algorithms and brain transcriptome data to elaborate the association of phenotype and genotype data. The integration of transcriptome data gives an advantage over the existing pipelines by directly translating genetic associations to brain regions and developmental patterns. Based on our data integrative approach, we identify genetic variants associated with Intelligence Quotient (IQ) in an autism cohort and found their respective genes to be expressed in specific brain areas.\n\nConclusionOur data integrative approach revealed that IQ is related to early down-regulated and late up-regulated gene modules implicated in frontal cortex and striatum, respectively. Besides identifying new gene associations with IQ we also provide a proof of concept, as several of the identified genes in our analysis are candidate genes related to intelligence in autism, intellectual disability, and Alzheimers disease. The framework provides a complete extensive analysis starting from a phenotypic trait data to its association at specific brain areas at vulnerable time points within a timespan of four days.\n\nAvailability and ImplementationOur framework is implemented in R and Python. It is available as an in-house script, which can be provided on demand.\n\nContactafsheen.yousaf@kgu.de

bioinformatics

circuitSNPs: Predicting genetic effects using a Neural Network to model regulatory modules of DNase-seq footprints

MotivationIdentifying and characterizing the function of non coding regions in the genome, and the genetic variants disrupting gene regulation, is a challenging question in genetics. Through the use of high throughput experimental assays that provide information about the chromatin state within a cell, coupled with modern computational approaches, much progress has been made towards this goal, yet we still lack a comprehensive characterization of the regulatory grammar. We propose a new method that combines sequence and chromatin accessibility information through a neural network framework with the goal of determining and annotating the effect of genetic variants on regulation of chromatin accessibility and gene transcription. Importantly, our new approach can consider multiple combinations of transcription factors binding at the same location when assessing the functional impact of non-coding genetic variation.\n\nResultsOur method, circuitSNPs, generates predictions describing the functional effect of genetic variants on local chromatin accessibility. Further, we demonstrate that circuitSNPs not only performs better than other variant annotation tools, but also retains the causal motifs / transcription factors that drive the predicted regulatory effect.\n\nContactfluca@wayne.edu, rpique@wayne.edu\n\nAvailabilityhttp://github.com/piquelab/circuitSNPs

bioinformatics

Maternal-Fetal Genetic Interactions, Imprinting, and Risk of Placental Abruption

Maternal genetic variations, including variations in mitochondrial biogenesis (MB) and oxidative phosphorylation (OP), have been associated with placental abruption (PA). However, the role of maternal-fetal genetic interactions (MFGI) and parent-of-origin (imprinting) effects in PA remain unknown. We investigated MFGI in MB-OP, and imprinting effects in relation to risk of PA. Among Peruvian mother-infant pairs (503 PA cases and 1,052 controls), independent single nucleotide polymorphisms (SNPs), with linkage-disequilibrium coefficient <0.80, were selected to characterize genetic variations in MB-OP (78 SNPs in 24 genes) and imprinted genes (2713 SNPs in 73 genes). For each MB-OP SNP, four multinomial models corresponding to fetal allele effect, maternal allele effect, maternal and fetal allele additive effect, and maternal-fetal allele interaction effect were fit under Hardy-Weinberg equilibrium, random mating, and rare disease assumptions. The Bayesian information criterion (BIC) was used for model selection. For each SNP in imprinted genes, imprinting effect was tested using a likelihood ratio test.\n\nBonferroni corrections were used to determine statistical significance (p-value<6.4e-4 for MFGI and p-value<1.8e-5 for imprinting). Abruption cases were more likely to experience preeclampsia, have shorter gestational age, and deliver infants with lower birthweight compared with controls. Models with MFGI effects provided improved fit than models with only maternal and fetal genotype main effects for SNP rs12530904 (log-likelihood ratio=18.2; p-value=1.2e-04) in CAMK2B, and, SNP rs73136795 (log-likelihood ratio=21.7; p-value=1.9e-04) in PPARG, both MB genes. We identified 311 SNPs in 35 maternally-imprinted genes (including KCNQ1, NPM, and, ATP10A) associated with abruption. Top hits included rs8036892 (p-value=2.3e-15) in ATP10A, rs80203467 (p-value=6.7e-15) and rs12589854 (p-value=1.4e-14) in MEG8, and rs138281088 in SLC22A2 (p-value=1.7e-13). We identified novel PA-related maternal-fetal MB gene interactions and imprinting effects that highlight the role of the fetus in PA risk development. Findings can inform mechanistic investigations to understand the pathogenesis of PA.\n\nAuthor summaryPlacental Abruption (PA) is a complex multifactorial and heritable disease characterized by premature separation of the placenta from the wall of the uterus. PA is a consequence of complex interplay of maternal and fetal genetics, epigenetics, and metabolic factors. Previous studies have identified common maternal single nucleotide polymorphisms (SNPs) in several mitochondrial biogenesis (MB) and oxidative phosphorylation (OP) genes that are associated with PA risk, although findings were inconsistent. Using the largest assembled mother-infant dyad of PA cases and controls, that includes participants from a previous report, we identified novel PA-related maternal-fetal MB gene interactions and imprinting effects that highlight the role of the fetus in PA risk development. Our findings have the potential for enhancing our understanding of genetic variations in maternal and fetal genome that contribute to PA.

genomics

Spatiotemporally controlled genetic perturbation for efficient large-scale studies of cell non-autonomous effects

Studies in genetic model organisms have revealed much about the development and pathology of complex tissues. Most have focused on cell-intrinsic gene functions and mechanisms. Much less is known about how transformed, or otherwise functionally disrupted, cells interact with healthy ones towards a favorable or pathological outcome. This is largely due to technical limitations. We developed new genetic tools in Drosophila melanogaster that permit efficient multiplexed gain- and loss-of-function genetic perturbations with separable spatial and temporal control. Importantly, our novel tool-set is independent of the commonly used GAL4/UAS system, freeing the latter for additional, non-autonomous, genetic manipulations; and is built into a single strain, allowing one-generation interrogation of non-autonomous effects. Altogether, our design opens up efficient genome-wide screens on any deleterious phenotype. Specifically, we developed tools to study extrinsic effects on neural tumor growth but the strategy presented has endless applications within and beyond neurobiology, and in other model organisms.\n\nImpact statementA novel genetic strategy for induction of reproducible neural tumors (or any other deleterious phenotype) in a single Drosophila stock (applicable to other organisms).

developmental biology

C. elegans genetic background modifies the core transcriptional response in an α-synuclein model of Parkinson’s disease

Accumulation of protein aggregates is a major cause of Parkinsons disease (PD), a progressive neurodegenerative condition that is one of the most common causes of dementia. Transgenic Caenorhabditis elegans worms expressing the human synaptic protein -synuclein show inclusions of aggregated protein and replicate the defining pathological hallmarks of PD. It is however not known how PD progression and pathology differs among individual genetic backgrounds. Here, we compared gene expression patterns, and investigated the phenotypic consequences of transgenic -synuclein expression in five different C. elegans genetic backgrounds. Transcriptome analysis indicates that the effects of -synuclein expression on pathways associated with nutrient storage, lipid transportation and ion exchange depend on the genetic background. The gene expression changes we observe suggest that a range of phenotypes will be affected by -synuclein expression. We experimentally confirm this, showing that the transgenic lines generally show delayed development, reduced lifespan, and an increased rate of matricidal hatching. These phenotypic effects coincide with the core changes in gene expression, linking developmental arrest, mobility, metabolic and cellular repair mechanisms to -synuclein expression. Together, our results show both genotype-specific effects and core alterations in global gene expression and in phenotype in response to -synuclein. We conclude that the PD effects are substantially modified by the genetic background, illustrating that genetic background mechanisms should be elucidated to understand individual variation in PD.

genomics

A Framework for Predicting Design Failures in Engineered Genetic Codes

Extreme engineering of an organisms genetic code could impart true genetic incompatibility, even blocking effects of horizontal gene transfer and viral infection. Recent experiments exploring this possibility demonstrate that such radical genome engineering achievements are plausible. However, it is unclear when the modifications will compromise the fitness of an organism. Efforts to reformat an entire genome are difficult and expensive; computational methods predicting fruitful experimental trajectories could play a pivotal role in advancing such efforts. We present a framework for building in silico models to assist genome-scale engineering. Genetic code engineering requires choosing from many possible codon-usage schemes, to find a design that is viable and effective. We use machine learning to identify which alternative codon-usage schemes are likely to result in no observed viable cells. Our data-driven approach employs observations of how modifying codon usage in individual genes impacted observed viability in E. coli, revealing salient features for early identification of problematic genetic code designs. We achieved an average area under the receiver operating characteristic of 0.72 on out-ofsample data.\n\nAuthor SummaryAs machine learning and artificial intelligence play an increasingly central role in science and engineering, it will be important to establish standardized techniques that facilitate the dialogue between experimentation and modeling. Biological experimental techniques are concurrently evolving at a rapid pace, providing unique opportunities to collect high-quality, novel information that was previously unobtainable. This work navigates the landscape of this vast, new territory, identifies interesting landmarks for exploration and posits new approaches towards advancing our research efforts in these areas. In this work, we show that, using a small dataset of 47 observations and rigorous nested cross validation techniques, we can build a model that makes better-than-random predictions of how codon usage changes in essential genes influence viability in E. coli. These predictions can be used to inform experimental trajectories in both genetic code and codon optimization experiments. We discuss ways to improve this model, iteratively, by performing high value experiments that decrease uncertainty in predictions and extrapolation error. Finally, we present novel visualization methods to aid in developing intuitions for how re-coding impacts groups of genes. These methods are also useful tools in building important insights into how well machine learning algorithms can generalize to new data.

bioengineering