Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,495 records · Page 83Linked to original sources

A genomic view of the peopling of the Americas

Whole-genome studies have documented that most Native American ancestry stems from a single population that diversified within the continent more than twelve thousand years ago. However, this shared ancestry hides a more complex history whereby at least four distinct streams of Eurasian migration have contributed to present-day and prehistoric Native American populations. Whole genome studies enhanced by technological breakthroughs in ancient DNA now provide evidence of a sequence of events involving initial migration from a structured Northeast Asian source population, followed by a divergence into northern and southern Native American lineages. During the Holocene, new migrations from Asia introduced the Saqqaq/Dorset Paleoeskimo population to the North American Arctic ~4,500 years ago, ancestry that is potentially connected with ancestry found in Athabaskan-speakers today. This was then followed by a major new population turnover in the high Arctic involving Thule-related peoples who are the ancestors of present-day Inuit. We highlight several open questions that could be addressed through future genomic research.

Evolutionary Biology

A natural encoding of genetic variation in a Burrows-Wheeler Transform to enable mapping and genome inference

We show how positional markers can be used to encode genetic variation within aBurrows-Wheeler Transform (BWT), and use this to construct a generalisation ofthe traditional \"reference genome\", incorporating known variation within aspecies. Our goal is to support the inference of the closest mosaic of previouslyknown sequences to the genome(s) under analysis.\n\nOur scheme results in an increased alphabet size, and by using a wavelet tree encoding of the BWT we reduce the performance impact on rank operations. We give a specialised form of the backward search that allows variation-aware exact matching. We implement this, and demonstrate the cost of constructing an index of the whole human genome with 8 million genetic variants is 25GB of RAM. We also show that inferring a closer reference can close large kilobase-scale coverage gaps in P. falciparum.

Bioinformatics

Protocol: Genome-scale CRISPR-Cas9 Knockout and Transcriptional Activation Screening

Forward genetic screens are powerful tools for the unbiased discovery and functional characterization of specific genetic elements associated with a phenotype of interest. Recently, the RNA-guided endonuclease Cas9 from the microbial immune system CRISPR (clustered regularly interspaced short palindromic repeats) has been adapted for genome-scale screening by combining Cas9 with guide RNA libraries. Here we describe a protocol for genome-scale knockout and transcriptional activation screening using the CRISPR-Cas9 system. Custom-or ready-made guide RNA libraries are constructed and packaged into lentivirus for delivery into cells for screening. As each screen is unique, we provide guidelines for determining screening parameters and maintaining sufficient coverage. To validate candidate genes identified from the screen, we further describe strategies for confirming the screening phenotype as well as genetic perturbation through analysis of indel rate and transcriptional activation. Beginning with library design, a genome-scale screen can be completed in 6-10 weeks followed by 3-4 weeks of validation.

Molecular Biology

Comparative analysis highlights variable genome content of wheat rusts and divergence of the mating loci.

Three members of the Puccini genus, P. triticina (Pt), P. striiformis f.sp. tritici(Pst), and P. graminis f.sp. tritici (Pgt), cause the most common and often most significant foliar diseases of wheat. While similar in biology and life cycle, each species is uniquely adapted and specialized. The genomes of Pt and Pst were sequenced and compared to that of Pgt to identify common and distinguishing gene content, to determine gene variation among wheat rust pathogens, other rust fungi and basidiomycetes, and to identify genes of significance for infection. Pt had the largest genome of the three, estimated at 135 Mb with expansion due to mobile elements and repeats encompassing 50.9% of contig bases; by comparison repeats occupy 31.5% for Pst and 36.5% for Pgt. We find all three genomes are highly heterozygous, with Pst (5.97 SNPs/kb) nearly twice the level detected in Pt (2.57 SNPs/kb) and that previously reported for Pgt. Of 1,358 predicted effectors in Pt, 784 were found expressed across diverse life cycle stages including the sexual stage. Comparison to related fungi highlighted the expansion of gene families involved in transcriptional regulation and nucleotide binding, protein modification, and carbohydrate enzyme degradation. Two allelic homeodomain, HD1 and HD2, pairs and three pheromone receptor (STE3) mating-type genes were identified in each dikaryotic Puccinia species. The HD proteins were active in a heterologous Ustilago maydis mating assay and host induced gene silencing of the HD and STE3 alleles reduced wheat host infection.

Microbiology

Cross-species genome-wide identification of evolutionary conserved microProteins

MicroProteins are small single domain proteins that act by engaging their targets into non-productive protein complexes. In order to identify novel microProteins in any sequenced genome of interest, we have developed miPFinder, a program that identifies and classifies potential microProteins. In the past years, several microProteins have been discovered in plants where they are mainly involved in the regulation of development. The miPFinder algorithm identifies all up to date known plant microProteins and extends the microProtein concept to other protein families. Here, we reveal potential microProtein candidates in several plant and animal reference genomes. A large number of these microProteins are species-specific while others evolved early and are evolutionary highly conserved. Most known microProtein genes originated from large ancestral genes by gene duplication, mutation and subsequent degradation. Gene ontology analysis shows that putative microProtein ancestors are often located in the nucleus, and involved in DNA binding and formation of protein complexes. Additionally, microProtein candidates act in plant transcriptional regulation, signal transduction and anatomical structure development. MiPFinder is freely available to find microProteins in any genome and will aid in the identification of novel microProteins in plants and animals

Bioinformatics

Modelling the transcription factor DNA-binding affinity using genome-wide ChIP-based data

Understanding protein-DNA binding affinity is still a mystery for many transcription factors (TFs). Although several approaches have been proposed in the literature to model the DNA-binding specificity of TFs, they still have some limitations. Most of the methods require a cut-off threshold in order to classify a K-mer as a binding site (BS) and finding such a threshold is usually done by handcraft rather than a science. Some other approaches use a prior knowledge on the biological context of regulatory elements in the genome along with machine learning algorithms to build classifier models for TFBSs. Noticeably, these methods deliberately select the training and testing datasets so that they are very separable. Hence, the current methods do not actually capture the TF-DNA binding relationship. In this paper, we present a threshold-free framework based on a novel ensemble learning algorithm in order to locate TFBSs in DNA sequences. Our proposed approach creates TF-specific classifier models using genome-wide DNA-binding experiments and a prior biological knowledge on DNA sequences and TF binding preferences. Systematic background filtering algorithms are utilized to remove non-functional K-mers from training and testing datasets. To reduce the complexity of classifier models, a fast feature selection algorithm is employed. Finally, the created classifier models are used to scan new DNA sequences and identify potential binding sites. The analysis results show that our proposed approach is able to identify novel binding sites in the Saccharomyces cerevisiae genome.\n\nContactmonther.alhamdoosh@unimelb.edu.au, dh.wang@latrobe.edu.au\n\nAvailabilityhttp://homepage.cs.latrobe.edu.au/dwang/DNNESCANweb

Bioinformatics

KAT: A K-mer Analysis Toolkit to quality control NGS datasets and genome assemblies

MotivationDe novo assembly of whole genome shotgun (WGS) next-generation sequencing (NGS) data bene[fi]ts from high-quality input with high coverage. However, in practice, determining the quality and quantity of useful reads quickly and in a reference-free manner is not trivial. Gaining a better understanding of the WGS data, and how that data is utilised by assemblers, provides useful insights that can inform the assembly process and result in better assemblies.\n\nResultsWe present the K-mer Analysis Toolkit (KAT): a multi-purpose software toolkit for reference-free quality control (QC) of WGS reads and de novo genome assemblies, primarily via their k-mer frequencies and GC composition. KAT enables users to assess levels of errors, bias and contamination at various stages of the assembly process. In this paper we highlight KATs ability to provide valuable insights into assembly composition and quality of genome assemblies through pairwise comparison of k-mers present in both input reads and the assemblies.\n\nAvailabilityKAT is available under the GPLv3 license at: https://github.com/TGAC/KAT.\n\nContactbernardo.clavijo@earlham.ac.uk\n\nSupplementary InformationSupplementary Information (SI) is available at Bioinformatics online. In addition, the software documentation is available online at: http://kat.readthedocs.io/en/latest/.

Bioinformatics

Discovery of Cancer Driver Long Noncoding RNAs across 1112 Tumour Genomes: New Candidates and Distinguishing Features.

Long noncoding RNAs (lncRNAs) represent a vast unexplored genetic space that may hold missing drivers of tumourigenesis, but few such \"driver lncRNAs\" are known. Until now, they have been discovered through changes in expression, leading to problems in distinguishing between causative roles and passenger effects. We here present a different approach for driver lncRNA discovery using mutational patterns in tumour DNA. Our pipeline, ExInAtor, identifies genes with excess load of somatic single nucleotide variants (SNVs) across panels of tumour genomes. Heterogeneity in mutational signatures between cancer types and individuals is accounted for using a simple local trinucleotide background model, which yields high precision and low computational demands. We use ExInAtor to predict drivers from the GENCODE annotation across 1112 entire genomes from 23 cancer types. Using a stratified approach, we identify 15 high-confidence candidates: 9 novel and 6 known cancer-related genes, including MALAT1, NEAT1 and SAMMSON. Both known and novel driver lncRNAs are distinguished by elevated gene length, evolutionary conservation and expression. We have presented a first catalogue of mutated lncRNA genes driving cancer, which will grow and improve with the application of ExInAtor to future tumour genome projects.

Cancer Biology

The druggable genome and support for target identification and validation in drug development

Target identification (identifying the correct drug targets for each disease) and target validation (demonstrating the effect of target perturbation on disease biomarkers and disease end-points) are essential steps in drug development. We showed previously that biomarker and disease endpoint associations of single nucleotide polymorphisms (SNPs) in a gene encoding a drug target accurately depict the effect of modifying the same target with a pharmacological agent; others have shown that genomic support for a target is associated with a higher rate of drug development success. To delineate drug development (including repurposing) opportunities arising from this paradigm, we connected complex disease- and biomarker-associated loci from genome wide association studies (GWAS) to an updated set of genes encoding druggable human proteins, to compounds with bioactivity against these targets and, where these were licensed drugs, to clinical indications. We used this set of genes to inform the design of a new genotyping array, to enable druggable genome-wide association studies for drug target selection and validation in human disease.

Genetics

Genome-wide methylation data mirror ancestry information

Genetic data are known to harbor information about human demographics, and genotyping data are commonly used for capturing ancestry information by leveraging genome-wide differences between populations. In contrast, it is not clear to what extent population structure is captured by whole-genome DNA methylation data. We demonstrate, using three large cohort 450K methylation array data sets, that ancestry information signal is mirrored in genome-wide DNA methylation data, and that it can be further isolated more effectively by leveraging the correlation structure of CpGs with cis-located SNPs. Based on these insights, we propose a method, EPISTRUCTURE, for the inference of ancestry from methylation data, without the need for genotype data. EPISTRUCTURE can be used to infer ancestry information of individuals based on their methylation data in the absence of corresponding genetic data. Although genetic data are often collected in epigenetic studies of large cohorts, these are typically not made publicly available, making the application of EPISTRUCTURE especially useful for anyone working on public data. Implementation of EPISTRUCTURE is available in GLINT, our recently released toolset for DNA methylation analysis at: http://glint-epigenetics.readthedocs.io.

Genetics

Recombination-driven genome evolution and stability of bacterial species

While bacteria divide clonally, horizontal gene transfer followed by homologous recombination is now recognized as an important and sometimes even dominant contributor to their evolution. However, the details of how the competition between clonal inheritance and recombination shapes genome diversity, population structure, and species stability remains poorly understood. Using a computational model, we find two principal regimes in bacterial evolution and identify two composite parameters that dictate the evolutionary fate of bacterial species. In the divergent regime, characterized by either a low recombination frequency or strict barriers to recombination, cohesion due to recombination is not sufficient to overcome the mutational drift. As a consequence, the divergence between any pair of genomes in the population steadily increases in the course of their evolution. The species as a whole lacks genetic coherence with sexually isolated clonal sub-populations continuously formed and dissolved. In contrast, in the metastable regime, characterized by a high recombination frequency combined with low barriers to recombination, genomes continuously recombine with the rest of the population. The population remains genetically cohesive and stable over time. The transition between these two regimes can be affected by relatively small changes in evolutionary parameters. Using the Multi Locus Sequence Typing (MLST) data we classify a number of well-studied bacterial species to be either the divergent or the metastable type. Generalizations of our framework to include fitness and selection, ecologically structured populations, and horizontal gene transfer of non-homologous regions are discussed.

Evolutionary Biology

Exploring evolutionary relationships across the genome using topology weighting

We introduce the concept of topology weighting, a method for quantifying relationships between taxa that are not necessarily monophyletic, and visualising how these relationships change across the genome. A given set of taxa can be related in a limited number of ways, but if each taxon is represented by multiple sequences, the number of possible topologies becomes very large. Topology weighting reduces this complexity by quantifying the contribution of each 'taxon topology' to the full tree. We describe our method for topology weighting by iterative sampling of sub-trees (Twisst), and test it on both simulated and real genomic data. Overall, we show that this is an informative and versatile approach, suitable for exploring relationships in almost any genomic dataset.\n\nScripts to implement the method described are available at github.com/simonhmartin/twisst.

Evolutionary Biology

The evolution of the natural killer complex; a comparison between mammals using new high-quality genome assemblies and targeted annotation

Natural killer (NK) cells are a diverse population of lymphocytes with a range of biological roles including essential immune functions. NK cell diversity is created by the differential expression of cell surface receptors which modulate activation and function, including multiple subfamilies of C-type lectin receptors encoded within the NK gene complex (NKC). Little is known about the gene content of the NKC beyond rodent and primate lineages, other than it appears to be extremely variable between mammalian groups. We compared the NKC structure between mammalian species using new high quality draft genome assemblies for cattle and goat, re-annotated sheep, pig and horse genome assemblies and the published human, rat and mouse lemur NKC. The major NKC genes are largely in syntenic positions in all eight species, with significant independent expansions and deletions between species, allowing us to present a model for NKC evolution during mammalian radiation. The ruminant species, cattle and goats, have independently evolved a second KLRC locus flanked by KLRA and KLRJ and a novel KLRH-like gene has acquired an activating tail. This novel gene has duplicated several times within cattle, while other activating receptor genes have been selectively disrupted. Targeted genome enrichment in cattle identified varying levels of allelic polymorphism between these NKC genes concentrated in the predicted extracellular ligand binding domains. This novel recombination and allelic polymorphism is consistent with NKC evolution under balancing selection, suggesting this diversity influences individual immune responses and may impact on differential outcomes of pathogen infection and vaccination.

Genetics

The impact of chromatin dynamics on Cas9-mediated genome editing in human cells

In order to efficiently edit eukaryotic genomes, it is critical to test the impact of chromatin dynamics on CRISPR/Cas9 function and develop strategies to adapt the system to eukaryotic contexts. So far, research has extensively characterized the relationship between the CRISPR endonuclease Cas9 and the composition of the RNADNA duplex that mediates the systems precision. Evidence suggests that chromatin modifications and DNA packaging can block eukaryotic genome editing by custom-built DNA endonucleases like Cas9; however, the underlying mechanism of Cas9 inhibition is unclear. Here, we demonstrate that closed, gene-silencing-associated chromatin is a mechanism for the interference of Cas9-mediated DNA editing. Our assays use a transgenic cell line with a drug-inducible switch to control chromatin states (open and closed) at a single genomic locus. We show that closed chromatin inhibits editing at specific target sites, and that artificial reversal of the silenced state restores editing efficiency. These results provide new insights to improve Cas9-mediated editing in human and other mammalian cells.\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=95 SRC=\"FIGDIR/small/071464_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (26K):\norg.highwire.dtl.DTLVardef@1f2c52org.highwire.dtl.DTLVardef@96df75org.highwire.dtl.DTLVardef@1288f24org.highwire.dtl.DTLVardef@1cdaa02_HPS_FORMAT_FIGEXP M_FIG C_FIG

Synthetic Biology

Genomic signature of kin selection in an ant with obligately sterile workers

Kin selection is thought to drive the evolution of cooperation and conflict, but the specific genes and genome-wide patterns shaped by kin selection are unknown. We identified thousands of genes associated with the sterile ant worker caste, the archetype of an altruistic phenotype shaped by kin selection, and then used population and comparative genomic approaches to study patterns of molecular evolution at these genes. Consistent with population genetic theoretical predictions, worker-upregulated genes showed relaxed adaptive evolution compared to genes upregulated in reproductive castes. Worker-upregulated genes included more taxonomically-restricted genes, indicating that the worker caste has recruited more novel genes, yet these genes also showed relaxed selection. Our study identifies a putative genomic signature of kin selection and helps to integrate emerging sociogenomic data with longstanding social evolution theory.

Evolutionary Biology

Resurrected protein interaction networks reveal the strong rewiring that leads to network organisation after whole genome duplication

The evolution of plants is characterized by several rounds of ancient whole genome duplication, sometimes closely associated with the origin of large groups of species. A good example is the {gamma} triplication at the origin of core eudicots. Core eudicots comprise about 75% of flowering plants and are characterized by the canalization of reproductive development. To better understand the impact of this genomic event, we studied the protein interaction network of MADS-domain transcription factors, which are key regulators of reproductive development. We accurately inferred, resurrected and tested the interactions of ancestral proteins before and after the triplication and directly compared these ancestral networks to the networks of Arabidopsis and tomato. We find that the {gamma} triplication generated a dramatically innovated network that strongly rewired through the addition of many new interactions. Many of these interactions were established between paralogous proteins and a new interaction partner, establishing new redundancy. Simulations show that both node and edge addition through the triplication were important to maintain modularity in the network. In addition to generating insights into the impact of whole genome duplication and elementary processes involved in network evolution, our data provide a resource for comparative developmental biology in flowering plants.

Evolutionary Biology

PCR artifact in testing for homologous recombination in genomic editing in zebrafish

We report a PCR-induced artifact in testing for homologous recombination in zebrafish. We attempted to replace the lnx2a gene with a donor cassette, mediated by a TALEN induced double stranded cut. The donor construct was flanked with homology arms of about 1 kb at the 5 and 3 ends. Injected embryos (G0) were raised and outcrossed to wild type fish. A fraction of the progeny appeared to have undergone the desired homologous recombination, as tested by PCR using primer pairs extending from genomic DNA outside the homology region to a site within the donor cassette. However, Southern blots revealed that no recombination had taken place. We conclude that recombination happened during PCR in vitro between the donor integrated elsewhere in the genome and the lnx2a locus, as suggested by earlier work [1]. We conclude that PCR alone may be insufficient to verify homologous recombination in genome editing experiments in zebrafish.

Developmental Biology

Variant Set Enrichment: An R package to Identify Dis-ease-Associated Functional Genomic Regions

SummaryGenetic predispositions to diseases populate the noncoding regions of the human genome. Delineating their functional basis can inform on the mechanisms contributing to disease development. However, this remains a challenge due to the poor characterization of the noncoding genome. Variant Set Enrichment (VSE) is a fast method to calculate the enrichment of a set of disease-associated variants across functionally annotated genomic regions, consequently highlighting the mechanisms important in the etiology of the disease studied.\n\nAvailability and ImplementationVSE is implemented as an R package and can easily be implemented in any system with R. See supplementary information for details.\n\nContacthansenhe@uhnresearch.ca; mlupien@uhnresearch.ca

Bioinformatics