Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,135 records · Page 63Linked to original sources

Genetic transformation of micropropagated shoots of Pinus radiata D.Don

Key Message Agrobacterium tumefaciens was used to transform radiata pine shoots and to efficiently produce stable genetically modified pine plants.\n\nAbstract Micropropagated shoot explants from Pinus radiata D. Don were used to produce stable transgenic plants by Agrobacterium tumefaciens-mediated transformation. Using this method any genotype that can be micropropagated could produce stable transgenic lines. As over 80% of P. radiata genotypes tested can be micropropagated, this effectively means that any line chosen for superior characteristics could be transformed. There are well-established protocols for progressing such germplasm to field deployment. Here we used open and control pollinated seed lines and embryogenic clones. The method developed was faster than other methods previously developed using mature cotyledons. PCR positive shoots could be obtain within 6 months of Agrobacterium co-cultivation compared with 12 months for cotyledon methods. Transformed shoots were obtained using either kanamycin or geneticin as the selectable marker gene. Shoots were recovered from selection, were tested and were not chimeric, indicating that the selection pressure was optimal for this explant type. GFP was used as a vital marker, and the bar gene, (for resistance to the herbicide Buster(R)) was used to produce lines that could potentially be used in commercial application. As expected, a range of expression phenotypes were identified for both these reporter genes and the analyses for expression were relatively easy.

Plant Biology

The Genetic Cost of Neanderthal Introgression

Approximately 2-4% of genetic material in human populations outside Africa is derived from Neanderthals who interbred with anatomically modern humans. Recent studies have shown that this Neanderthal DNA is depleted around functional genomic regions; this has been suggested to be a consequence of harmful epistatic interactions between human and Neanderthal alleles. However, using published estimates of Neanderthal inbreeding and the distribution of mutational fitness effects, we infer that Neanderthals had at least 40% lower fitness than humans on average; this increased load predicts the reduction in Neanderthal introgression around genes without the need to invoke epistasis. We also predict a residual Neanderthal mutational load in non-Africans, leading to a fitness reduction of at least 0.5%. This effect of Neanderthal admixture has been left out of previous debate on mutation load differences between Africans and non-Africans. We also show that if many deleterious mutations are recessive, the Neanderthal admixture fraction could increase over time due to the protective effect of Neanderthal haplotypes against deleterious alleles that arose recently in the human population. This might partially explain why so many organisms retain gene flow from other species and appear to derive adaptive benefits from introgression.

Evolutionary Biology

pong: fast analysis and visualization of latent clusters in population genetic data

1 MotivationA series of methods in population genetics use multilocus genotype data to assign individuals membership in latent clusters. These methods belong to a broad class of mixed-membership models, such as latent Dirichlet allocation used to analyze text corpora. Inference from mixed-membership models can produce different output matrices when repeatedly applied to the same inputs, and the number of latent clusters is a parameter that is often varied in the analysis pipeline. For these reasons, quantifying, visualizing, and annotating the output from mixed-membership models are bottlenecks for investigators across multiple disciplines from ecology to text data mining.\n\n2 ResultsWe introduce pong, a network-graphical approach for analyzing and visualizing membership in latent clusters with a native D3.js interactive visualization. pong leverages efficient algorithms for solving the Assignment Problem to dramatically reduce runtime while increasing accuracy compared to other methods that process output from mixed-membership models. We apply pong to 225,705 unlinked genome-wide single-nucleotide variants from 2,426 unrelated individuals in the 1000 Genomes Project, and identify previously overlooked aspects of global human population structure. We show that pong outpaces current solutions by more than an order of magnitude in runtime while providing a customizable and interactive visualization of population structure that is more accurate than those produced by current tools.\n\n3 Availabilitypong is freely available and can be installed using the Python package management system pip. pongs source code is available at https://github.com/abehr/pong.\n\n4 Contactaaron_behr@alumni.brown.edu,\n\nsramachandran@brown.edu

Genomics

DisGeNET-RDF: harnessing the innovative power of the Semantic Web to explore the genetic basis of diseases

MotivationDisGeNET-RDF makes available knowledge on the genetic basis of human diseases in the Semantic Web (SW). Gene-disease associations (GDAs) and their provenance metadata are published as human-readable and machine-processable web resources. The information on GDAs included in DisGeNET-RDF is interlinked to other biomedical databases to support the development of bioinformatics approaches for translational research through evidence-based exploitation of a rich and fully interconnected Linked Open Data (LOD).\n\nAvailabilityhttp://rdf.disgenet.org/\n\nContactsupport@disgenet.org

Bioinformatics

Contrasting the genetic architecture of 30 complex traits from summary association data

Variance components methods that estimate the aggregate contribution of large sets of variants to the heritability of complex traits have yielded important insights into the disease architecture of common diseases. Here, we introduce new methods that estimate the total variance in trait explained by a single locus in the genome (local heritability) from summary GWAS data while accounting for linkage disequilibrium (LD) among variants. We apply our new estimator to ultra large-scale GWAS summary data of 30 common traits and diseases to gain insights into their local genetic architecture. First, we find that common SNPs have a high contribution to the heritability of all studied traits. Second, we identify traits for which the majority of the SNP heritability can be confined to a small percentage of the genome. Third, we identify GWAS risk loci where the entire locus explains significantly more variance in the trait than the GWAS reported variants. Finally, we identify 55 loci that explain a large proportion of heritability across multiple traits.

Bioinformatics

Sequence element enrichment analysis to determine the genetic basis of bacterial phenotypes

Bacterial genomes vary extensively in terms of both gene content and gene sequence - this plasticity hampers the use of traditional SNP-based methods for identifying all genetic associations with phenotypic variation. Here we introduce a computationally scalable and widely applicable statistical method (SEER) for the identification of sequence elements that are significantly enriched in a phenotype of interest. SEER is applicable to even tens of thousands of genomes by counting variable-length k-mers using a distributed string-mining algorithm. Robust options are provided for association analysis that also correct for the clonal population structure of bacteria. Using large collections of genomes of the major human pathogens Streptococcus pneumoniae and Streptococcus pyogenes, SEER identifies relevant previously characterised resistance determinants for several antibiotics and discovers potential novel factors related to the invasiveness of S. pyogenes. We thus demonstrate that our method can answer important biologically and medically relevant questions.

Genomics

Genetic redundancies enhance information transfer in noisy regulatory circuits

Cellular decision making is based on regulatory circuits that associate signal thresholds to specific physiological actions. This transmission of information is subjected to molecular noise what can decrease its fidelity. Here, we show instead how such intrinsic noise enhances information transfer in the presence of multiple circuit copies. The result is due to the contribution of noise to the generation of autonomous responses by each copy, which are altogether associated with a common decision. Moreover, factors that correlate the responses of the redundant units (extrinsic noise or regulatory cross-talk) contribute to reduce fidelity, while those that further uncouple them (heterogeneity within the copies) can lead to stronger information gain. Overall, our study emphasizes how the interplay of signal thresholding, redundancy, and noise influences the accuracy of cellular decision making. Understanding this interplay provides a basis to explain collective cell signaling mechanisms, and to engineer robust decisions with noisy genetic circuits.

Systems Biology

Functional Genetic Screen to Identify Interneurons Governing Behaviorally Distinct Aspects of Drosophila Larval Motor Programs

Drosophila larval crawling is an attractive system to study patterned motor output at the level of animal behavior. Larval crawling consists of waves of muscle contractions generating forward or reverse locomotion. In addition, larvae undergo additional behaviors including head casts, turning, and feeding. It is likely that some neurons are used in all these behaviors (e.g. motor neurons), but the identity (or even existence) of neurons dedicated to specific aspects of behavior is unclear. To identify neurons that regulate specific aspects of larval locomotion, we performed a genetic screen to identify neurons that, when activated, could elicit distinct motor programs. We used 165 Janelia CRM-Gal4 lines - chosen for sparse neuronal expression - to express the warmth-inducible neuronal activator TrpA1 and screened for locomotor defects. The primary screen measured forward locomotion velocity, and we identified 63 lines that had locomotion velocities significantly slower than controls following TrpA1 activation (28{degrees}C). A secondary screen was performed on these lines, revealing multiple discrete behavioral phenotypes including slow forward locomotion, excessive reverse locomotion, excessive turning, excessive feeding, immobile, rigid paralysis, and delayed paralysis. While many of the Gal4 lines had motor, sensory, or muscle expression that may account for some or all of the phenotype, some lines showed specific expression in a sparse pattern of interneurons. Our results show that distinct motor programs utilize distinct subsets of interneurons, and provide an entry point for characterizing interneurons governing different elements of the larval motor program.

Neuroscience

Genetics of cortico-cerebellar expansion in anthropoid primates: a comparative approach

What adaptive changes in brain structure and function underpin the evolution of increased cognitive performance in humans and our close relatives? Identifying the genetic basis of brain evolution has become a major tool in answering this question. Numerous cases of positive selection, altered gene expression or gene duplication have been identified that may contribute to the evolution of the neocortex, which is widely assumed to play a predominant role in cognitive evolution. However, the neocortex co-evolves with other, functionally inter-dependent, regions of the brain, most notably the cerebellum. The cerebellum is linked to a range of cognitive tasks and expanded rapidly during hominoid evolution, independently of neocortex size. Here we demonstrate that, across primates, genes with known roles in cerebellum development are just as likely to be targeted by selection as genes linked to cortical development. In fact, cerebellum genes are more likely to have evolved adaptively during hominoid evolution, consistent with phenotypic data suggesting an accelerated rate of cerebellar expansion in apes. Finally, we present evidence that selection targeted genes with specific effects on either the neocortex or cerebellum, not both. This suggests cortico-cerebellar co-evolution is maintained by selection acting on independent developmental programs.

Evolutionary Biology

Estimating dispersal kernels using genetic parentage data

Dispersal kernels are the standard method for describing and predicting the relationship between dispersal strength and distance. Statistically-fitted dispersal kernels allow observations of a limited number of dispersal events to be extrapolated across a wider landscape, and form the basis of a wide range of theories and methods in ecology, evolution and conservation. Genetic parentage data are an increasingly common source of dispersal information, particularly for species where dispersal is difficult to observe directly. It is now routinely applied to coral reef fish, whose larvae disperse over many kilometers and are too small to follow directly. However, it is not straightforward to estimate dispersal kernels from parentage data, and existing methods each have substantial limitations. Here we develop and proof a new statistical estimator for fitting dispersal kernels to parentage data, applying it to simulated and empirical datasets of reef fish parentage. The method incorporates a series of factors omitted in previous methods: the partial sampling of adults and juveniles on sampled reefs; the existence of unassigned dispersers from unsampled reefs; and post-settlement processes (e.g., density dependent mortality) that follow dispersal but precede parentage sampling. Power analyses indicate that the highest levels of sampling currently used for reef fishes is sufficient to fit accurate dispersal kernels. Sampling is best distributed equally between adults and juveniles, and over more than ten populations. Importantly, we show that accounting for unsampled or unassigned individuals - including adult individuals on partially-sampled and unsampled patches - is essential for a precise and unbiased estimate of dispersal.

Ecology

The effects of population size histories on estimates of selection coefficients from time-series genetic data

AO_SCPCAPBSTRACTC_SCPCAPMany approaches have been developed for inferring selection coefficients from time series data while accounting for genetic drift. However, the improvement in inference accuracy that can be attained by modeling drift is unknown. Here, by comparing maximum likelihood estimates of selection coefficients that account for the true population size history with estimates that ignore drift, we address the following questions: how much can modeling the population size history improve estimates of selection coefficients? How much can mis-inferred population sizes hurt inferences of selection coefficients? We conduct our analysis under the discrete Wright-Fisher model by deriving the exact probability of an allele frequency trajectory in a population of time-varying size and we replicate our results under the diffusion model by extending the exact probability of a frequency trajectory derived by Steinrucken et al. (2014) to the case of a piecewise constant population. For both the discrete Wright-Fisher and diffusion models, we find that ignoring drift leads to estimates of selection coefficients that are nearly as accurate as estimates that account for the true population history, even when population sizes are small and drift is high. In populations of time-varying size, estimates of selection coefficients that ignore drift are similar in accuracy to estimates that rely on crude, yet reasonable, estimates of the population history. These results are of interest because inference methods that ignore drift are widely used in evolutionary studies and can be many orders of magnitude faster than methods that account for population sizes.

Evolutionary Biology

The C. elegans NF2/Merlin Molecule NFM-1 Non-Autonomously Regulates Neuroblast Migration and Interacts Genetically with the Guidance Cue SLT-1/Slit

During nervous system development, neurons and their progenitors often migrate to their final destinations. In Caenorhabditis elegans, the bilateral Q neuroblasts and their descendants migrate long distances in opposite directions, despite being born in the same posterior region. QR on the right migrates anteriorly and generates the AQR neuron positioned near the head, and QL on the left migrates posteriorly, giving rise to the PQR neuron positioned near the tail. In a screen for genes required for AQR and PQR migration, we identified an allele of nfm-1, which encodes a molecule similar to vertebrate NF2/Merlin, an important tumor suppressor in humans. Mutations in NF2 lead to Neurofibromatosis Type II, characterized by benign tumors of glial tissues. These molecules contain Four-point-one Ezrin Radixin Moesin (FERM) domains characteristic of cytoskeletal-membrane linkers, and vertebrate NF2 is required for epidermal integrity. Vertebrate NF2 can also regulate several transcriptional pathways including the Hippo pathway. Here we demonstrate that in C. elegans, nfm-1 is required for complete migration of AQR and PQR, and that it likely acts outside of the Q cells themselves in a non-autonomous fashion. We also show a genetic interaction between nfm-1 and the C. elegans Slit homolog slt-1, which encodes a conserved secreted guidance cue. In vertebrates, NF2 can control Slit2 mRNA levels through the hippo pathway in axon pathfinding, suggesting a conserved interaction of NF2 and Slit2 in regulating migration.

Developmental Biology

Massively parallel whole-organism lineage tracing using CRISPR/Cas9 induced genetic scars

A key goal of developmental biology is to understand how a single cell transforms into a full-grown organism consisting of many cells. Although impressive progress has been made in lineage tracing using imaging approaches, analysis of vertebrate lineage trees has mostly been limited to relatively small subsets of cells. Here we present scartrace, a strategy for massively parallel clonal analysis based on Cas9 induced genetic scars in the zebrafish.

Systems Biology

Lazarus effects: the frequency and genetic causes of Escherichia coli population recovery under lethal heat stress

Sometimes populations crash and yet recover before being lost completely. Such recoveries have been observed incidentally in evolution experiments using Escherichia coli, and this phenomenon has been termed the \"Lazarus effect.\" To investigate how often recovery occurs and the genetic changes that drive it, we evolved ~300 populations of E. coli at lethally high temperatures (43.0{degrees}) for five days and sequenced the genomes of recovered populations. Our results revealed that the Lazarus effect is uncommon, but frequent enough, at ~9% of populations, to be a potent source of evolutionary innovation. Population sequencing uncovered a set of mutations adaptive to lethal 43.0{degrees}C that were mostly distinct from those that were beneficial at a high but nonlethal temperature (42.2{degrees}). Mutations within two operons--the heat shock hslUV operon and the RNA polymerase rpoBC operon--drove adaptation to lethal temperature. Mutations in hslUV exhibited little antagonistic pleiotropy at 37.0{degrees}C and may have arisen neutrally prior to subjection to lethal temperature. In contrast, rpoBC mutations provided greater fitness benefits than hslUV mutants, but were less prevalent and caused stronger fitness tradeoffs at lower temperatures. Recovered populations fixed mutations in only one operon or the other, but not both, indicating that epistatic interactions between beneficial mutations were important even at the earliest stages of adaptation.

Evolutionary Biology

Leveraging Functional Annotations in Genetic Risk Prediction for Human Complex Diseases

Genome wide association studies have identified numerous regions in the genome associated with hundreds of human diseases. Building accurate genetic risk prediction models from these data will have great impacts on disease prevention and treatment strategies. However, prediction accuracy remains moderate for most diseases, which is largely due to the challenges in identifying all the disease-associated variants and accurately estimating their effect sizes. We introduce AnnoPred, a principled framework that incorporates diverse functional annotation data to improve risk prediction accuracy, and demonstrate its performance on multiple human complex diseases.

Bioinformatics

Genetic and transcriptional analysis of human host response to healthy gut microbiome

Many studies have demonstrated the importance of the gut microbiome in healthy and disease states. However, establishing the causality of host-microbiome interactions in humans is still challenging. Here, we describe a novel experimental system to define the transcriptional response induced by the microbiome in human cells and to shed light on the molecular mechanisms underlying host-gut microbiome interactions. In primary human colonic epithelial cells, we identified over 6,000 genes that change expression at various time points following co-culturing with the gut microbiome of a healthy individual. The differentially expressed genes are enriched for genes associated with several microbiome-related diseases, such as obesity and colorectal cancer. In addition, our experimental system allowed us to identify 87 host SNPs that show allele-specific expression in 69 genes. Furthermore, for 12 SNPs in 12 different genes, allele-specific expression is conditional on the exposure to the microbiome. Of these 12 genes, eight have been associated with diseases linked to the gut microbiome, specifically colorectal cancer, obesity and type 2 diabetes. Our study demonstrates a scalable approach to study host-gut microbiome interactions and can be used to identify putative mechanisms for the interplay between host genetics and microbiome in health and disease.

Genomics

Genetic variability in both the adaptive and innate immune systems contribute to Alzheimer’s and Parkinson’s disease risk

Neurodegenerative disorders are devastating diseases with a worldwide health-care burden. Studies have demonstrated enrichment of disease-associated genetic variants with functional genomic annotations. Determining associated cell-types is important to understand pathogenicity.\n\nWe obtained GWAS summary statistics from Parkinsons disease (PD), Alzheimers disease (AD), amyotrophic lateral sclerosis (ALS), multiple sclerosis (MS), and frontotemporal dementia (FTD). We applied stratified LD score regression to determine if functional categories are enriched for heritability.\n\nThere was little enrichment of brain annotations, but annotations from both the innate and adaptive immune systems were enriched for MS (as expected), AD, and PD, in decreasing order of statistical significance.

Genomics

Forward genetic screen of human transposase genomic rearrangements

Background. Numerous human genes encode potentially active DNA transposases or recombinases, but our understanding of their functions remains limited due to shortage of methods to profile their activities on endogenous genomic substrates. Results. To enable functional analysis of human transposase-derived genes, we combined forward chemical genetic hypoxanthine-guanine phosphoribosyltransferase 1 (HPRT1) screening with massively parallel paired-end DNA sequencing and structural variant genome assembly and analysis. Here, we report the HPRT1 mutational spectrum induced by the human transposase PGBD5, including PGBD5-specific signal sequences (PSS) that serve as potential genomic rearrangement substrates. Conclusions. The discovered PSS motifs and high-throughput forward chemical genomic screening approach should prove useful for the elucidation of endogenous genome remodeling activities of PGBD5 and other domesticated human DNA transposases and recombinases.

Genomics