Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Accurate genetic profiling of anthropometric traits using a big data approach

Genome-wide association studies (GWAS) promised to translate their findings into clinically beneficial improvements of patient management by tailoring disease management to the individual through the prediction of disease risk1,2. However, the ability to translate genetic findings from GWAS into predictive tools that are of clinical utility and which may inform clinical practice has, so far, been encouraging but limited1,2. Here we propose to use a more powerful statistical approach that enables the prediction of multiple medically relevant phenotypes without the costs associated with developing a genetic test for each of them. As a proof of principle, we used a common panel of 319,038 SNPs to train the prediction models in 114,264 unrelated White-British for height and four obesity related traits (body mass index, basal metabolic rate, body fat percentage, and waist-to-hip ratio). We obtained prediction accuracies that ranged between 46% and 75% of the maximum achievable given their explained heritable component. This represents an improvement of up to 75% over the phenotypic variance explained by the predictors developed through large collaborations3, which used more than twice as many training samples. Across-population predictions in White nonBritish individuals were similar to those of White-British whilst those in Asian and Black individuals were informative but less accurate. The genotyping of circa 500,000 UK Biobank4 participants will yield predictions ranging between 66% and 83% of the maximum. We anticipate that our models and a common panel of genetic markers, which can be used across multiple traits and diseases, will be the starting point to tailor disease management to the individual. Ultimately, we will be able to capitalise on whole-genome sequence and environmental risk factors to realise the full potential of genomic medicine.

Genomics

A novel nuclear genetic code alteration in yeasts and the evolution of codon reassignment in eukaryotes

The genetic code is the universal cellular translation table to convert nucleotide into amino acid sequences. Changes to sense codons are expected to be highly detrimental. However, reassignments of single or multiple codons in mitochondria and nuclear genomes demonstrated that the code can evolve. Still, alterations of nuclear genetic codes are extremely rare leaving hypotheses to explain these variations, such as the codon capture, the genome streamlining and the ambiguous intermediate theory, in strong debate. Here, we report on a novel sense codon reassignment in Pachysolen tannophilus, a yeast related to the Pichiaceae. By generating proteomics data and using tRNA sequence comparisons we show that in Pachysolen CUG codons are translated as alanine and not as the universal leucine. The polyphyly of the CUG-decoding tRNAs in yeasts is best explained by a tRNA loss driven codon reassignment mechanism. Loss of the CUG-tRNA in the ancient yeast is followed by gradual decrease of respective codons and subsequent codon capture by tRNAs whose anticodon is outside the aminoacyl-tRNA synthetase recognition region. Our hypothesis applies to all nuclear genetic code alterations and provides several testable predictions. We anticipate more codon reassignments to be uncovered in existing and upcoming genome projects.

Evolutionary Biology

Craniofacial shape transition across the house mouse hybrid zone: implications for the genetic architecture and evolution of between-species differences

Craniofacial shape differences between taxa have often being linked to environmental adaptation, e.g. to new food sources, or have been studied in the context of domestication. Evidence for the genetic basis of such phenotypic differences to date suggests that within- as well as between-species variation has an oligogenic basis, i.e. few loci of large effect explain most of the variation. In mice, it has been shown that within-population craniofacial variation has a highly polygenic basis, but there are no data regarding the genetic basis of between-species differences. Here, we address this question using a phenotype-focused approach. Using 3D geometric morphometrics, we phenotyped a panel of mice derived from a natural hybrid zone between M. m. domesticus and M. m. musculus, and quantify the transition of craniofacial shape along the hybridization gradient. We find a continuous shape transition along the hybridization gradient, and unaltered developmental stability associated with hybridization. This suggests that the morphospace between the two subspecies is continuous despite reproductive isolation and strong barriers to gene flow. We show that quantitative changes in genome composition generate quantitative changes in craniofacial shape; this supports a highly polygenic basis for between-species craniofacial differences in the house mouse. We discuss our findings in the context of oligogenic versus polygenic models of the genetic architecture of morphological traits.

Evolutionary Biology

Multiple genetic changes underlie the evolution of long-tailed forest deer mice

Understanding both the role of selection in driving phenotypic change and its underlying genetic basis remain major challenges in evolutionary biology. Here we focus on a classic system of local adaptation in the North American deer mouse, Peromyscus maniculatus, which occupies two main habitat types, prairie and forest. Using historical collections we demonstrate that forest-dwelling mice have longer tails than those from non-forested habitats, even when we account for individual and population relatedness. Based on genome-wide SNP capture data, we find that mice from forested habitats in the eastern and western parts of their range form separate clades, suggesting that increased tail length evolved independently from a short-tailed ancestor. Two major changes in skeletal morphology can give rise to longer tails--increased number and increased length of vertebrae--and we find that forest mice in the east and west have both more and longer caudal vertebrae, but not trunk vertebrae, than nearby prairie forms. Using a second-generation intercross between a prairie and forest pair, we show that the number and length of caudal vertebrae are not correlated in this recombinant population, suggesting that variation in these traits is controlled by separate genetic loci. Together, these results demonstrate convergent evolution of the long-tailed forest phenotype through multiple, distinct genetic mechanisms (controlling vertebral length and vertebral number), thus suggesting that these morphological changes--either independently or together--are adaptive.

Evolutionary Biology

Vcfanno: fast, flexible annotation of genetic variants

BackgroundThe integration of genome annotations and reference databases is critical to the identification of genetic variants that may be of interest in studies of disease or other traits. However, comprehensive variant annotation with diverse file formats is difficult with existing methods.\n\nResultsWe have developed vcfanno as a flexible toolset that simplifies the annotation of genetic variants in VCF format. Vcfanno can extract and summarize multiple attributes from one or more annotation files and append the resulting annotations to the INFO field of the original VCF file. Vcfanno also integrates the lua scripting language so that users can easily develop custom annotations and metrics. By leveraging a new parallel \"chromosome sweeping\" algorithm, it enables rapid annotation of both whole-exome and whole-genome datasets. We demonstrate this performance by annotating over 85.3 million variants in less than 17 minutes (>85,000 variants per second) with 50 attributes from 17 commonly used genome annotation resources.\n\nConclusionsVcfanno is a flexible software package that provides researchers with the ability to annotate genetic variation with a wide range of datasets and reference databases in diverse genomic formats.\n\nAvailabilityThe vcfanno source code is available at https://github.com/brentp/vcfanno under the MIT license, and platform-specific binaries are available at https://github.com/brentp/vcfanno/releases. Detailed documentation is available at http://brentp.github.io/vcfanno/, and the code underlying the analyses presented can be found at https://github.com/brentp/vcfanno/tree/master/scripts/paper.

Bioinformatics

Common genetic variation drives molecular heterogeneity in human iPSCs

Induced pluripotent stem cell (iPSC) technology has enormous potential to provide improved cellular models of human disease. However, variable genetic and phenotypic characterisation of many existing iPSC lines limits their potential use for research and therapy. Here, we describe the systematic generation, genotyping and phenotyping of 522 open access human iPSCs derived from 189 healthy male and female individuals as part of the Human Induced Pluripotent Stem Cells Initiative (HipSci:http://www.hipsci.org). Our study provides a comprehensive picture of the major sources of genetic and phenotypic variation in iPSCs and establishes their suitability for use in genetic studies of complex human traits and cancer. Using a combination of genomewide analyses we find that 5-25% of the variation in different iPSC phenotypes, including differentiation capacity and cellular morphology, arises from differences betweenindividuals. We also assess the phenotypic effects of rare, genomic copy number mutations that are recurrently seen following iPSC reprogramming and present an initial map of common regulatory variants affecting the transcriptome of pluripotent cells in humans.

Genomics

Single cell transcriptomics, mega-phylogeny and the genetic basis of morphological innovations in Rhizaria

The innovation of the eukaryote cytoskeleton enabled phagocytosis, intracellular transport and cytokinesis, and is responsible for diverse eukaryotic morphologies. Still, the relationship between phenotypic innovations in the cytoskeleton and their underlying genotype is poorly understood. To explore the genetic mechanism of morphological evolution of the eukaryotic cytoskeleton we provide the first single cell transcriptomes from uncultivable, free-living unicellular eukaryotes: the radiolarian species Lithomelissa setosa and Sticholonche zanclea. Analysis of the genetic components of the cytoskeleton and mapping of the evolution of these to a revised phylogeny of Rhizaria reveals lineage-specific gene duplications and neo-functionalization of and {beta} tubulin in Retaria, actin in Retaria and Endomyxa, and Arp2/3 complex genes in Chlorarachniophyta. We show how genetic innovations have shaped cytoskeletal structures in Rhizaria, and how single cell transcriptomics can be applied for resolving deep phylogenies and studying gene evolution of uncultivable protist species.

Evolutionary Biology

The domesticated brain: genetics of brain mass and brain structure in an avian species

As brain size usually increases with body size it has been assumed that the two are tightly constrained and evolutionary studies have therefore often been based on relative brain size (i.e. brain size proportional to body size) instead of absolute brain size. The process of domestication offers an excellent opportunity to disentangle the linkage between body and brain mass due to the extreme selection for increased body mass that has occurred. By breeding an intercross between domestic chicken and their wild progenitor, we address this relationship by simultaneously mapping the genes that control inter-population variation in brain mass and body mass. Loci controlling variation in brain mass and body mass have separate genetic architectures and are therefore not directly constrained. Genetic mapping of brain regions in the intercross indicates that domestication has led to a larger body mass and to a lesser extent a larger absolute brain mass in chickens, mainly due to enlargement of the cerebellum. Domestication has traditionally been linked to brain mass regression, based on measurements of relative brain mass, which confounds the large body mass augmentation due to domestication. Our results refute this concept in chicken and confirm recent studies that show that different genetic architectures underlie these traits.

Evolutionary Biology

Genetic Diversity, Population Structure and Species Delimitation of Trialeurodes vaporariorum (Greenhouse whitefly)

Genetic diversity within Trialeurodes vaporariorum (Westwood, 1856) remains largely unexplored, particularly within regions of Sub-Saharan Africa. In this study, T. vaporariorum samples were obtained from three locations in Kenya: Katumani, Kiambu and Kajiado counties. DNA extraction, PCR and Sanger sequencing were carried out on ~750 bp fragment of the mitochondria cytochrome c oxidase I (COI) gene from individual whiteflies. In addition, global populations were assessed and 19 haplotypes were identified, with three main haplotypes (Hp_19, Hp_10, Hp_011) circulating within Kenya. Measures of genetic diversity among T. vaporariorum populations resulted in haplotype diversity of 0.411, nucleotide diversity 0.00096, and Tajimas D -0. 30315, (P>0.10). Analysis of population structure across global sequences using Structurama indicated one population globally, with posterior probability of 0.72. Bayesian and maximum likelihood phylogenetic analysis gave support for two clades (Clade I = an admixed global population and Clade II = subset of Kenyan and 1 Greek sequence). Species delimitation between the two clades was assessed by four parameters; posterior probability, Kimuras two parameter (K2P), Rodrigos P (Randomly distinct) and Rosenbergs reciprocal monophyly (P(AB). The two clades within the phylogenetic tree showed evidence of distinctness based on; Kimura two parameters (K2P) (p = -1.21E-01), Rodrigos P (RD) (p =0.05) and Rosenbergs P(AB) (p = 2.3E -13). Overall, low genetic diversity within the Kenyan samples is a likely indicator of recent population expansion and colonization with this region and plausible signs of species complex formation in Sub-Saharan Africa.

Evolutionary Biology

Predicting functional neuroanatomical maps from fusing brain networks with genetic information

A central aim, from basic neuroscience to psychiatry, is to resolve how genes control brain circuitry and behavior. This is experimentally hard, since most brain functions and behaviors are controlled by multiple genes. In low throughput, one gene at a time, experiments, it is therefore difficult to delineate the neural circuitry through which these sets of genes express their behavioral effects. The increasing amount of publicly available brain and genetic data offers a rich source that could be mined to address this problem computationally. However, most computational approaches are not tailored to reflect functional synergies in brain circuitry accumulating within sets of genes. Here, we developed an algorithm that fuses gene expression and connectivity data with functional genetic meta data and exploits such cumulative effects to predict neuroanatomical maps for multigenic functions. These maps recapture known functional anatomical annotations from literature and functional MRI data. When applied to meta data from mouse QTLs and human neuropsychiatric databases, our method predicts functional maps underlying behavioral or psychiatric traits. We show that it is possible to predict functional neuroanatomy from mouse and human genetic meta data and provide a discovery tool for high throughput functional exploration of brain anatomy in silico.

Neuroscience

Population genetic history and polygenic risk biases in 1000 Genomes populations

The vast majority of genome-wide association studies are performed in Europeans, and their transferability to other populations is dependent on many factors (e.g. linkage disequilibrium, allele frequencies, genetic architecture). As medical genomics studies become increasingly large and diverse, gaining insights into population history and consequently the transferability of disease risk measurement is critical. Here, we disentangle recent population history in the widely-used 1000 Genomes Project reference panel, with an emphasis on populations underrepresented in medical studies. To examine the transferability of single-ancestry GWAS, we used published summary statistics to calculate polygenic risk scores for six well-studied traits and diseases. We identified directional inconsistencies in all scores; for example, height is predicted to decrease with genetic distance from Europeans, despite robust anthropological evidence that West Africans are as tall as Europeans on average. To gain deeper quantitative insights into GWAS transferability, we developed a complex trait coalescent-based simulation framework considering effects of polygenicity, causal allele frequency divergence, and heritability. As expected, correlations between true and inferred risk were typically highest in the population from which summary statistics were derived. We demonstrated that scores inferred from European GWAS were biased by genetic drift in other populations even when choosing the same causal variants, and that biases in any direction were possible and unpredictable. This work cautions that summarizing findings from large-scale GWAS may have limited portability to other populations using standard approaches, and highlights the need for generalized risk prediction methods and the inclusion of more diverse individuals in medical genomics.

Genomics

Evolutionary Genetics of Insecticide Resistance and the Effects of Chemical Rotation

Repeated use of the same class of pesticides to control a target pest is a form of artificial selection that leads to pesticide resistance. We studied insecticide resistance and cross-resistance to five commercial insecticides in each of six populations of the red flour beetle, Tribolium castaneum. We estimated the dosage response curves for lethality in each parent population for each insecticide and found an 800-fold difference among populations in resistance to insecticides. As expected, a naive laboratory population was among the most sensitive of populations to most insecticides. We then used inbred lines derived from five of these populations to estimate the heritability (h2) of resistance for each pesticide and the genetic correlation (rG) of resistance among pesticides in each population. These quantitative genetic parameters allow insight into the adaptive potential of populations to further evolve insecticide resistance. Lastly, we use our estimates of the genetic variance and covariance of resistance and stochastic simulations to evaluate the efficacy of \"windowing\" as an insecticide resistance management strategy, where the application of several insecticides is rotated on a periodic basis.

Evolutionary Biology

Stochastic dynamics of genetic broadcasting networks

The complex genetic programs of eukaryotic cells are often regulated by key transcription factors occupying or clearing out of a large number of genomic locations. Orchestrating the residence times of these factors is therefore important for the well organized functioning of a large network. The classic models of genetic switches sidestep this timing issue by assuming the binding of transcription factors to be governed entirely by thermodynamic protein-DNA affinities. Here we show that relying on passive thermodynamics and random release times can lead to a \"time-scale crisis\" of master genes that broadcast their signals to large number of binding sites. We demonstrate that this \"time-scale crisis\" can be resolved by actively regulating residence times through molecular stripping. We illustrate these ideas by studying the stochastic dynamics of the genetic network of the central eukaryotic master regulator NF{kappa}B which broadcasts its signals to many downstream genes that regulate immune response, apoptosis etc.

Systems Biology

Variation in olfactory neuron repertoires is genetically controlled and environmentally modulated

The mouse olfactory sensory neuron (OSN) repertoire is composed of 10 million cells and each expresses one olfactory receptor (OR) gene from a pool of over 1000. Thus, the nose is sub-stratified into more than a thousand OSN subtypes. Here, we employ and validate an RNA-sequencing based method to quantify the abundance of all OSN subtypes in parallel, and investigate the genetic and environmental factors that contribute to neuronal diversity. We find that the OSN subtype distribution is stereotyped in genetically identical mice, but varies extensively between different strains. Further, we identify cis-acting genetic variation as the greatest component influencing OSN composition and demonstrate independence from OR function. However, we show that olfactory stimulation with particular odorants results in modulation of dozens of OSN subtypes in a subtle but reproducible, specific and time-dependent manner. Together, these mechanisms generate a highly individualized olfactory sensory system by promoting neuronal diversity.

Neuroscience

The genetic basis and fitness consequences of sperm midpiece size in deer mice

An extraordinary array of reproductive traits vary among species, yet the genetic mechanisms that enable divergence, often over short evolutionary timescales, remain elusive. Here we examine two sister-species of Peromyscus mice with divergent mating systems. We find that the promiscuous species produces sperm with longer midpiece than the monogamous species, and midpiece size correlates positively with competitive ability and swimming performance. Using forward genetics, we identify a gene associated with midpiece length: Prkar1a, which encodes the R1 regulatory subunit of PKA. R1 localizes to midpiece in Peromyscus and is differentially expressed in mature sperm of the two species yet is similarly abundant in the testis. We also show that genetic variation at this locus accurately predicts male reproductive success. Our findings suggest that rapid evolution of reproductive traits can occur through cell type-specific changes to ubiquitously expressed genes and have an important effect on fitness.

Evolutionary Biology

An Optimized Approach for Annotation of Large Eukaryotic Genomic Sequences using Genetic Algorithm

Detection of important functional and/or structural elements and identifying their positions in a large eukaryotic genome is an active research area. Gene is an important functional and structural unit of DNA. The computation of gene prediction is essential for detailed genome annotation. In this paper, we propose a new gene prediction technique based on Genetic Algorithm (GA) for determining the optimal positions of exons of a gene in a chromosome or genome. The correct identification of the coding and non-coding regions are difficult and computationally demanding. The proposed genetic-based method, named Gene Prediction with Genetic Algorithm (GPGA), reduces this problem by searching only one exon at a time instead of all exons along with its introns. The advantage of this representation is that it can break the entire gene-finding problem into a number of smaller subspaces and thereby reducing the computational complexity. We tested the performance of the GPGA with some benchmark datasets and compared the results with the well-known and relevant techniques. The comparison shows the better or comparable performance of the proposed method (GPGA). We also used GPGA for annotating the human chromosome 21 (HS21) using cross species comparison with the mouse orthologs.

bioinformatics

Deciphering the genic basis of environmental adaptations by simultaneous forward and reverse genetics in Saccharomyces cerevisiae

The budding yeast Saccharomyces cerevisiae is the best studied eukaryote in molecular and cell biology, but its utility for understanding the genetic basis of natural phenotypic variation is limited by the inefficiency of association mapping owing to strong and complex population structure. To facilitate association mapping, we analyzed 190 high-quality genomes of diverse strains, including 85 newly sequenced ones, to uncover yeasts population structure that varies substantially among genomic regions. We identified 181 yeast genes that are absent from the reference genome and demonstrated their expression and role in important functions such as drug resistance. We then simultaneously measured the growth rates of over 4500 lab strains each deficient of a nonessential gene and 81 natural strains across multiple environments using unique DNA barcode present in each strain. We combined the genome-wide reverse genetic information with genome-wide association analysis to determine potential genomic regions of importance to environmental adaptations, and for a subset experimentally validated their role by reciprocal hemizygosity tests. The resources provided permit efficient and reliable association mapping in yeast and significantly enhances its value as a model for understanding the genetic mechanisms of phenotypic polymorphism and evolution.

genomics

Understanding genetic changes underlying the molybdate resistance and the glutathione production in Saccharomyces cerevisiae wine strains using an evolution-based strategy

In this work we have investigated the genetic changes underlying the high glutathione (GSH) production showed by the evolved Saccharomyces cerevisiae strain UMCC 2581, selected in a molybdate-enriched environment after sexual recombination of the parental wine strain UMCC 855. To reach our goal, we first generated strains with the desired phenotype, and then we mapped changes underlying adaptation to molybdate by using a whole-genome sequencing. Moreover, we carried out the RNA-seq that allowed an accurate measurement of gene expression and an effective comparison between the transcriptional profiles of parental and evolved strains, in order to investigate the relationship between genotype and high GSH production phenotype.\n\nAmong all genes evaluated only two genes, MED2 and RIM15 both related to oxidative stress response, presented new mutations in the UMCC 2581 strain sequence and were potentially related to the evolved phenotype.\n\nRegarding the expression of high GSH production phenotype, it included over-expression of amino acids permeases and precursor biosynthetic enzymes rather than the two GSH metabolic enzymes, whereas GSH production and metabolism, transporter activity, vacuolar detoxification and oxidative stress response enzymes were probably added resulting in the molybdate resistance phenotype. This work provides an example of a combination of an evolution-based strategy to successful obtain yeast strain with desired phenotype and inverse engineering approach to genetic characterize the evolved strain. The obtained genetic information could be useful for further optimization of the evolved strains and for providing an even more rapid approach to identify new strains, with a high GSH production, through a marked-assisted selection strategy.

genomics