Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Reference Quality Assembly of the 3.5 Gb genome of Capsicum annuum from a Single Linked-Read Library

BackgroundLinked-Read sequencing technology has recently been employed successfully for de novo assembly of multiple human genomes, however the utility of this technology for complex plant genomes is unproven. We evaluated the technology for this purpose by sequencing the 3.5 gigabase (Gb) diploid pepper (Capsicum annuum) genome with a single Linked-Read library. Plant genomes, including pepper, are characterized by long, highly similar repetitive sequences. Accordingly, significant effort is used to ensure the sequenced plant is highly homozygous and the resulting assembly is a haploid consensus. With a phased assembly approach, we targeted a heterozygous F1 derived from a wide cross to assess the ability to derive both haplotypes for a pungency gene characterized by a large insertion/deletion.\n\nResultsThe Supernova software generated a highly ordered, more contiguous sequence assembly than all currently available C. annuum reference genomes. Eighty-four percent of the final assembly was anchored and oriented using four de novo linkage maps. A comparison of the annotation of conserved eukaryotic genes indicated the completeness of assembly. The validity of the phased assembly is further demonstrated with the complete recovery of both 2.5 kb insertion/deletion haplotypes of the PUN1 locus in the F1 sample that represents pungent and non-pungent peppers.\n\nConclusionsThe most contiguous pepper genome assembly to date has been generated through this work which demonstrates that Linked-Read library technology provides a rapid tool to assemble de novo complex highly repetitive heterozygous plant genomes. This technology can provide an opportunity to cost-effectively develop high-quality reference genome assemblies for other complex plants and compare structural and gene differences through accurate haplotype reconstruction.

genomics

A genome resequencing-based genetic map reveals the recombination landscape of an outbred parasitic nematode in the presence of polyploidy and polyandry

The parasitic nematode Haemonchus contortus is an economically and clinically important pathogen of small ruminants, and a model system for understanding the mechanisms and evolution of traits such as anthelmintic resistance. Anthelmintic resistance is widespread and is a major threat to the sustainability of livestock agriculture globally; however, little is known about the genome architecture and parameters such as recombination that will ultimately influence the rate at which resistance may evolve and spread. Here we performed a genetic cross between two divergent strains of H. contortus, and subsequently used whole-genome re-sequencing of a female worm and her brood to identify the distribution of genome-wide variation that characterises these strains. Using a novel bioinformatic approach to identify variants that segregate as expected in a pseudo-testcross, we characterised linkage groups and estimated genetic distances between markers to generate a chromosome-scale F1 genetic map composed of 1,618 SNPs. We exploited this map to reveal the recombination landscape, the first for any parasitic helminth species, demonstrating extensive variation in recombination rate within and between chromosomes. Analyses of these data also revealed the extent of polyandry, whereby at least eight males were found to have contributed to the genetic variation of the progeny analysed. Triploid offspring were also identified, which we hypothesise are the result of nondisjunction during female meiosis or polyspermy. These results expand our knowledge of the genetics of parasitic helminths and the unusual life-history of H. contortus, and will enable more precise characterisation of the evolution and inheritance of genetic traits such as anthelmintic resistance. This study also demonstrates the feasibility of whole-genome resequencing data to directly construct a genetic map in a single generation cross from a non-inbred non-model organism with a complex lifecycle.\n\nAuthor summaryRecombination is a key genetic process, responsible for the generation of novel genotypes and subsequent phenotypic variation as a result of crossing over between homologous chromosomes. Populations of strongylid nematodes, such as the gastrointestinal parasites that infect livestock and humans, are genetically very diverse, but little is known about patterns of recombination across the genome and how this may contribute to the genetics and evolution of these pathogens. In this study, we performed a genetic cross to quantify recombination in the barbers pole worm, Haemonchus contortus, an important parasite of sheep and goats. The reproductive traits of this worm make standard genetic crosses challenging, but by generating whole-genome sequence data from a female worm and her offspring, we identified genetic variants that act as though they come from a single mating cross, allowing the use of standard statistical approaches to build a genetic map and explore the distribution and rates of recombination throughout the genome. A number of genetic signatures associated with H. contortus life history traits were revealed in this analysis: we extend our understanding of multiple paternity (polyandry) in this species, and provide evidence and explanation for sporadic increases in chromosome complements (polyploidy) among the progeny. The resulting genetic map will aid in population genomic studies in general and enhance ongoing efforts to understand the genetic basis of resistance to the drugs used to control these worms, as well as for related species that infect humans throughout the world.

genomics

Acute condensin depletion causes genome decompaction without altering the level of global gene expression in Saccharomyces cerevisiae

Condensins are broadly conserved chromosome organizers that function in chromatin compaction and transcriptional regulation, but to what extent these two functions are linked has remained unclear. Here, we analyzed the effect of condensin inactivation on genome compaction and global gene expression in the yeast Saccharomyces cerevisiae. Spike-in-controlled 3C-seq analysis revealed that acute condensin inactivation leads to a global decrease in close-range chromosomal interactions as well as more specific losses of homotypic tRNA gene clustering. In addition, a condensin-rich topologically associated domain between the ribosomal DNA and the centromere on chromosome XII is lost upon condensin inactivation. Unexpectedly, these large-scale changes in chromosome architecture are not associated with global changes in transcript levels as determined by spike-in-controlled mRNA-seq analysis. Our data suggest that the global transcriptional program of S. cerevisiae is resistant to condensin inactivation and the associated profound changes in genome organization.\n\nSignificance StatementGene expression occurs in the context of higher-order chromatin organization, which helps compact the genome within the spatial constraints of the nucleus. To what extent higher-order chromatin compaction affects gene expression remains unknown. Here, we show that gene expression and genome compaction can be uncoupled in the single-celled model eukaryote Saccharomyces cerevisiae. Inactivation of the conserved condensin complex, which also organizes the human genome, leads to broad genome decompaction in this organism. Unexpectedly, this reorganization has no immediate effect on the transcriptome. These findings indicate that the global gene expression program is robust to large-scale changes in genome architecture in yeast, shedding important new light on the evolution and function of genome organization in gene regulation.

genomics

The genome of the water strider Gerris buenoi reveals expansions of gene repertoires associated with adaptations to life on the water

The semi-aquatic bugs conquered water surfaces worldwide and occupy ponds, streams, lakes, mangroves, and even open oceans. As such, they inspired a range of scientific studies from ecology and evolution to developmental genetics and hydrodynamics of fluid locomotion. However, the lack of a representative water strider genome hinders thorough investigations of the mechanisms underlying the processes of adaptation and diversification in this group. Here we report the sequencing and manual annotation of the Gerris buenoi (G. buenoi) genome, the first water strider genome to be sequenced so far. G. buenoi genome is about 1 000Mb and the sequencing effort recovered 20 949 predicted protein-coding genes. Manual annotation uncovered a number of local (tandem and proximal) gene duplications and expansions of gene families known for their importance in a variety of processes associated with morphological and physiological adaptations to water surface lifestyle. These expansions affect key processes such as growth, vision, desiccation resistance, detoxification, olfaction and epigenetic components. Strikingly, the G. buenoi genome contains three Insulin Receptors, a unique case among metazoans, suggesting key changes in the rewiring and function of the insulin pathway. Other genomic changes include wavelength sensitivity shifts in opsin proteins likely in association with the requirements of vision in water habitats. Our findings suggest that local gene duplications might have had an important role during the evolution of water striders. These findings along with the G. buenoi genome open exciting research opportunities to understand adaptation and genome evolution of this unique hemimetabolous insect.

genomics

Chromosome evolution at the origin of the ancestral vertebrate genome

About 450 million years ago, a marine chordate was subject to two successive whole genome duplications (WGDs) before becoming the common ancestor of vertebrates and diversifying into the more than 60,000 species found today. Here, we reconstruct in details the evolution of chromosomes of this early vertebrate along successive steps of the two WGD. We first compared 61 extant animal genomes to build a highly contiguous order of genes in a 326 million years old ancestral Amniota genome. In this genome, we established a well-supported list of duplicated genes originating from the WGDs to link chromosomes in tetrads, a telltale signature of these events. This enabled us to reconstruct a scenario where a pre-vertebrate genome composed of 17 chromosomes duplicated into 34 chromosomes, and was subject to 7 chromosome fusions before duplicating again into 54 chromosomes. After the separation of Agnatha (jawless fish) and Gnathostomata, four more fusions took place to form the ancestral Euteleostomi genome of 50 chromosomes. These results firmly establish the occurrence of the two WGD, resolving in particular the ambiguity raised by the analysis of the lamprey genetic map. In addition, we provide insight into the origin of homologous micro-chromosomes found in the chicken and the gar genomes. This work provides a foundation for studying the evolution of vertebrate chromosomes from the standpoint of a common ancestor, and particularly the pattern of duplicate gene retention and loss that resulted in the gene composition of extant genomes.

genomics

Consensus assessment of the contamination level of publicly available cyanobacterial genomes

BACKGROUNDPublicly available genomes are crucial for phylogenetic and metagenomic studies, in which contaminating sequences can be the cause of major problems. This issue is expected to be especially important for Cyanobacteria because axenic strains are notoriously difficult to obtain and keep in culture. Yet, despite their great scientific interest, no data are currently available concerning the quality of publicly available cyanobacterial genomes.\n\nRESULTSAs reliably detecting contaminants is a complex task, we designed a pipeline combining six methods in a consensus strategy to assess the contamination level of 440 genome assemblies of Cyanobacteria. Two methods are based on published reference databases of ribosomal genes (SSU rRNA 16S and ribosomal proteins), one is indirectly based on a reference database of marker genes (CheckM), and three are based on complete genome analysis. Among those genome-wide methods, Kraken and DIAMOND blastx share the same reference database that we derived from Ensembl Bacteria, whereas CONCOCT does not require any reference database, instead relying on differences in DNA tetramer frequencies. Given that all the six methods appear to have their own strengths and limitations, we used the consensus of their rankings to infer that >5% of cyanobacterial genome assemblies are highly contaminated by foreign DNA (i.e., contaminants were detected by 5 or 6 methods).\n\nCONCLUSIONSOur results will help researchers to check the quality of publicly available genomic data before use in their own analyses. Moreover, we argue that journals should make mandatory the submission of raw read data along with genome assemblies in order to facilitate the detection of contaminants in sequence databases.

genomics

Deep SNP: An End-to-end Deep Neural Network with Attention-based Localization for Break-point Detection in SNP Array Genomic data

Diagnosis and risk stratification of cancer and many other diseases require the detection of genomic breakpoints as a prerequisite of calling copy number alterations (CNA). This, however, is still challenging and requires time-consuming manual curation. As deep-learning methods outperformed classical state-of-the-art algorithms in various domains and have also been successfully applied to life science problems including medicine and biology, we here propose Deep SNP, a novel Deep Neural Network to learn from genomic data. Specifically, we used a manually curated dataset from 12 genomic single nucleotide polymorphism array (SNPa) profiles as truth-set and aimed at predicting the presence or absence of genomic breakpoints, an indicator of structural chromosomal variations, in windows of 40,000 probes. We compare our results with well-known neural network models as well as Rawcopy though this tool is designed to predict breakpoints and in addition genomic segments with high sensitivity. We show, that Deep SNP is capable of successfully predicting the presence or absence of a breakpoint in large genomic windows and outperforms state-of-the-art neural network models. Qualitative examples suggest that integration of a localization unit may enable breakpoint detection and prediction of genomic segments, even if the breakpoint coordinates were not provided for network training. These results warrant further evaluation of DeepSNP for breakpoint localization and subsequent calling of genomic segments.

genomics

Comparative genomic analysis revealed rapid differentiation in the pathogenicity-related gene repertoires between Pyricularia oryzae and Pyricularia penniseti isolated from a Pennisetum grass

BackgroundsPyricularia is a multispecies complex that could infect and cause severe blast disease on diverse hosts, including rice, wheat and many other grasses. Although the genome size of this fungal complex is small [~40 Mbp for Pyricularia oryzae (syn. Magnaporthe oryzae), and ~45 Mbp for P. grisea], the genome plasticity allows the fungus to jump and adapt to new hosts. Therefore, deciphering the genome basis of individual species could facilitate the evolutionary and genetic study of this fungus. However, except for the P. oryzae subgroup, many other species isolated from diverse hosts, such as the Pennisetum grasses, remain largely uncovered genetically.\n\nResultsHere, we report the genome sequence of a pyriform-shaped fungal strain P. penniseti P1609 isolated from a Pennisetum grass (JUJUNCAO) using PacBio SMRT sequencing technology. We performed a phylogenomic analysis of 28 Magnaporthales species and 5 non-Magnaporthales species and addressed P1609 into a Pyricularia subclade that is distant from P. oryzae. Comparative genomic analysis revealed that the pathogenicity-related gene repertoires were fairly different between P1609 and the P. oryzae strain 70-15, including the cloned avirulence genes, other putative secreted proteins, as well as some other predicted Pathogen-Host Interaction (PHI) genes. Genomic sequence comparison also identified many genomic rearrangements.\n\nConclusionTaken together, our results suggested that the genomic sequence of the P. penniseti P1609 could be a useful resource for the genetic study of the Pennisetum-infecting Pyricularia species.

genomics

Fast and flexible bacterial genomic epidemiology with PopPUNK

The routine use of genomics for disease surveillance provides the opportunity for high-resolution bacterial epidemiology.\n\nHowever, current whole-genome clustering and multi-locus typing approaches do not fully exploit core and accessory genomic variation, and cannot both automatically identify, and subsequently expand, clusters of significantly-similar isolates in large datasets and across species.\n\nHere we describe PopPUNK (Population Partitioning Using Nucleotide K-mers; https://poppunk.readthedocs.io/en/latest/). software implementing scalable and expandable annotation- and alignment-free methods for population analysis and clustering.\n\nVariable-length k-mer comparisons are used to distinguish isolates divergence in shared sequence and gene content, which we demonstrate to be accurate over multiple orders of magnitude using both simulated data and real datasets from ten taxonomically-widespread species. Connections between closely-related isolates of the same strain are robustly identified, despite variation in the discontinuous pairwise distance distributions that reflects species diverse evolutionary patterns. PopPUNK can process 103-104 genomes as single batch, with minimal memory use and runtimes up to 200-fold faster than existing methods. Clusters of strains remain consistent as new batches of genomes are added, which is achieved without needing to re-analyse all genomes de novo.\n\nThis facilitates real-time surveillance with stable cluster naming and allows for outbreak detection using hundreds of genomes in minutes. Interactive visualisation and online publication is streamlined through automatic output of results to multiple platforms.\n\nPopPUNK has been designed as a flexible platform that addresses important issues with currently used whole-genome clustering and typing methods, and has potential uses across bacterial genetics and public health research.

genomics

Unbiased whole genomes from mammalian feces using fluorescence-activated cell sorting

Ecological flexibility, extended lifespans, and large brains, have long intrigued evolutionary biologists, and comparative genomics offers an efficient and effective tool for generating new insights into the evolution of such traits. Studies of capuchin monkeys are particularly well situated to shed light on the selective pressures and genetic underpinnings of local adaptation to diverse habitats, longevity, and brain development. Distributed widely across Central and South America, they are inventive and extractive foragers, known for their sensorimotor intelligence. Capuchins have the largest relative brain size of any monkey and a lifespan that exceeds 50 years, despite their small (3-5 kg) body size. We assemble a de novo reference genome for Cebus imitator and provide the first genome annotation of a capuchin monkey. Through high-depth sequencing of DNA derived from blood, various tissues and feces via fluorescence activated cell sorting (fecalFACS) to isolate monkey epithelial cells, we compared genomes of capuchin populations from tropical dry forests and lowland rainforests and identified population divergence in genes involved in water balance, kidney function, and metabolism. Through a comparative genomics approach spanning a wide diversity of mammals, we identified genes under positive selection associated with longevity and brain development. Additionally, we provide a technological advancement in the use of non-invasive genomics for studies of free-ranging mammals. Our intra- and interspecific comparative study of capuchin genomics provides new insights into processes underlying local adaptation to diverse and physiologically challenging environments, as well as the molecular basis of brain evolution and longevity. SIGNIFICANCESurviving challenging environments, living long lives, and engaging in complex cognitive processes are hallmark characteristics of human evolution. Similar traits have evolved in parallel in capuchin monkeys, but their genetic underpinnings remain unexplored. We developed and annotated a reference assembly for white-faced capuchin monkeys to explore the evolution of these phenotypes. By comparing populations of capuchins inhabiting rainforest versus dry forests with seasonal droughts, we detected selection in genes associated with kidney function, muscular wasting, and metabolism, suggesting adaptation to periodic resource scarcity. When comparing capuchins to other mammals, we identified evidence of selection in multiple genes implicated in longevity and brain development. Our research was facilitated by our new method to generate high- and low-coverage genomes from non-invasive biomaterials.

genomics

Complete Genome Sequence of the Wolbachia wAlbB Endosymbiont of Aedes albopictus

Wolbachia, an alpha-proteobacterium closely related to Rickettsia is a maternally transmitted, intracellular symbiont of arthropods and nematodes. Aedes albopictus mosquitoes are naturally infected with Wolbachia strains wAlbA and wAlbB. Cell line Aa23 established from Ae. albopictus embryos retains only wAlbB and is a key model to study host-endosymbiont interactions. We have assembled the complete circular genome of wAlbB from the Aa23 cell line using long-read PacBio sequencing at 500X median coverage. The assembled circular chromosome is 1.48 megabases in size, an increase of more than 300 kb over the published draft wAlbB genome. The annotation of the genome identified 1,205 protein coding genes, 34 tRNA, 3 rRNA, 1 tmRNA and 3 other ncRNA loci. The long reads enabled sequencing over complex repeat regions which are difficult to resolve with short-read sequencing. Thirteen percent of the genome is comprised of IS elements distributed throughout the genome, some of which cause pseudogenization. Prophage WO genes encoding some essential components of phage particle assembly are missing, while the remainder are scattered around the genome. Orthology analysis identified a core proteome of 536 orthogroups across all completed Wolbachia genomes. The majority of proteins could be annotated using Pfam and eggNOG analyses, including ankyrins and components of the T4SS. KEGG analysis revealed the absence of 5 genes in wAlbB which are present in other Wolbachia. The availability of a complete circular chromosome from wAlbB will enable further biochemical, molecular and genetic analyses on this strain and related Wolbachia.\n\nData depositionRaw data from PacBio sequencing have been deposited in the NCBI SRA database under BioProject accession number PRJNA454708, as runs SRR7784284, SRR7784285, SRR7784286, SRR7784287. The paired-end reads from Illumina library used for indel correction are available from NCBI SRA database as accession SRR7623731. The assembled genome and annotations have been submitted to the NCBI GenBank database with accession number CP031221.

genomics

The fate of deleterious variants in a barley genomic prediction population

Targeted identification and purging of deleterious genetic variants has been proposed as a novel approach to animal and plant breeding. This strategy is motivated, in part, by the observation that demographic events and strong selection associated with cultivated species pose a \"cost of domestication.\" This includes an increase in the proportion of genetic variants where a mutation is likely to reduce fitness. Recent advances in DNA resequencing and sequence constraint-based approaches to predict the functional impact of a mutation permit the identification of putatively deleterious SNPs (dSNPs) on a genome-wide scale. Using exome capture resequencing of 21 barley 6-row spring breeding lines, we identify 3,855 dSNPs among 497,754 total SNPs. In order to polarize SNPs as ancestral versus derived, we generated whole genome resequencing data of Hordeum murinum ssp. glaucum as a phylogenetic outgroup. The dSNPs occur at higher density in portions of the genome with a higher recombination rate than in pericentromeric regions with lower recombination rate and gene density. Using 5,215 progeny from a genomic prediction experiment, we examine the fate of dSNPs over three breeding cycles. Average derived allele frequency is lower for dSNPs than any other class of variants. Adjusting for initial frequency, derived alleles at dSNPs reduce in frequency or are lost more often than other classes of SNPs. The highest yielding lines in the experiment, as chosen by standard genomic prediction approaches, carry fewer homozygous dSNPs than randomly sampled lines from the same progeny cycle. In the final cycle of the experiment, progeny selected by genomic prediction have a mean of 5.6% fewer homozygous dSNPs relative to randomly chosen progeny from the same cycle.\n\nAuthor SummaryThe nature of genetic variants underlying complex trait variation has been the source of debate in evolutionary biology. Here, we provide evidence that agronomically important phenotypes are influenced by rare, putatively deleterious variants. We use exome capture resequencing and a hypothesis-based test for codon conservation to predict deleterious SNPs (dSNPS) in the parents of a multi-parent barley breeding population. We also generated whole-genome resequencing data of Hordeum murinum, a phylogenetic outgroup to barley, to polarize dSNPs by ancestral versus derived state. dSNPs occur disproportionately in the gene-rich chromosome arms, rather than in the recombination-poor pericentromeric regions. They also decrease in frequency more often than other variants at the same initial frequency during recurrent selection for grain yield and disease resistance. Finally, we identify a region on chromosome 4H that strongly associated with agronomic phenotypes in which dSNPs appear to be hitchhiking with favorable variants. Our results show that targeted identification and removal of dSNPs from breeding programs is a viable strategy for crop improvement, and that standard genomic prediction approaches may already contain some information about unobserved segregating dSNPs.

genomics

Joint single cell DNA-Seq and RNA-Seq of gastric cancer reveals subclonal signatures of genomic instability and gene expression

Sequencing the genomes of individual cancer cells provides the highest resolution of intratumoral heterogeneity. To enable high throughput single cell DNA-Seq across thousands of individual cells per sample, we developed a droplet-based, automated partitioning technology for whole genome sequencing. We applied this approach on a set of gastric cancer cell lines and a primary gastric tumor. In parallel, we conducted a separate single cell RNA-Seq analysis on these same cancers and used copy number to compare results. This joint study, covering thousands of single cell genomes and transcriptomes, revealed extensive cellular diversity based on distinct copy number changes, numerous subclonal populations and in the case of the primary tumor, subclonal gene expression signatures. We found genomic evidence of positive selection - where the percentage of replicating cells per clone is higher than expected - indicating ongoing tumor evolution. Our study demonstrates that joining single cell genomic DNA and transcriptomic features provides novel insights into cancer heterogeneity and biology. SIGNIFICANCEWe conducted a massively parallel DNA sequencing analysis on a set of gastric cancer cell lines and a primary gastric tumor in combination with a joint single cell RNA-Seq analysis. This joint study, covering thousands of single cell genomes and transcriptomes, revealed extensive cellular diversity based on distinct copy number changes, numerous subclonal populations and in the case of the primary tumor, subclonal gene expression signatures. We found genomic evidence of positive selection where the percentage of replicating cells per clone is higher than expected indicating ongoing tumor evolution. Our study demonstrates that combining single cell genomic DNA and transcriptomic features provides novel insights into cancer heterogeneity and biology.

genomics

On enhancing variation detection through pan-genome indexing

Detection of genomic variants is commonly conducted by aligning a set of reads sequenced from an individual to the reference genome of the species and analyzing the resulting read pileup. Typically, this process finds a subset of variants already reported in databases and additional novel variants characteristic to the sequenced individual. Most of the effort in the literature has been put to the alignment problem on a single reference sequence, although our gathered knowledge on species such as human is pan-genomic: We know most of the common variation in addition to the reference sequence. There have been some efforts to exploit pan-genome indexing, where the most widely adopted approach is to build an index structure on a set of reference sequences containing observed variation combinations.\n\nThe enhancement in alignment accuracy when using pan-genome indexing has been demonstrated in experiments, but so far the above multiple references pan-genome indexing approach has not been tested on its final goal, that is, in enhancing variation detection. This is the focus of this article: We study a generic approach to add variation detection support on top of the multiple references pan-genomic indexing approach. Namely, we study the read pileup on a multiple alignment of reference genomes, and propose a heaviest path algorithm to extract a new recombined reference sequence. This recombined reference sequence can then be utilized in any standard read alignment and variation detection workflow. We demonstrate that the approach enhances variation detection on realistic data sets.

Bioinformatics

Genome sequence of tsetse bracoviruses: insights into symbiotic virus evolution

Mutualism between endogenous viruses and eukaryotes is still poorly understood. Whole genome data has highlighted the diverse distribution of viral sequences in several eukaryote host genomes. A group of endogenous double-stranded polydnaviruses known as bracoviruses has been identified in parasitic braconid wasp (Hymenoptera). Bracoviruses allow wasps to reproductively co-opt other insect larvae. Bracoviruses are excised from the host genome and injected in to the larva along side the wasp eggs; where they encode proteins that lower host immunity allowing development of parasitoid wasp larvae in the host. Interestingly, putative bracoviral sequences have recently been detected in the first sequenced genome of the tsetse fly (Diptera). This is peculiar since tsetse flies do not share this reproductive lifestyle. To investigate genome rearrangements associated with these unique mutual symbiotic relationships and examine its value as a potential vector control strategy entry point. We use comparative genomics to determine the presence, prevalence and genetic diversity of bracoviruses of five tsetse fly species (G. austeni, G. brevipalpis, G. f. fuscipes, G. m. morsitans and G. pallidipes) and the housefly (Musca domestica). We identify and use four viral Maverick genes as evolutionary models for bracoviruses. This is the first record of homologous bracoviruses in multiple Dipteran genomes. Phylogenetic reconstruction of each gene revealed two major clades that represent the two types of Mavericks. We detect varying magnitudes of purifying selection across these loci except for the poxvirus A32 gene, which is under positive selection. Moreover, these genes were inserted at conserved regions and co-evolve at similar rates with the host genomes.

Evolutionary Biology

Uniparental inheritance promotes adaptive evolution in cytoplasmic genomes

1Eukaryotes carry numerous asexual cytoplasmic genomes (mitochondria and plastids). Lacking recombination, asexual genomes should theoretically suffer from impaired adaptive evolution. Yet, empirical evidence indicates that cytoplasmic genomes experience higher levels of adaptive evolution than predicted by theory. In this study, we use a computational model to show that the unique biology of cytoplasmic genomes--specifically their organization into host cells and their uniparental (maternal) inheritance--enable them to undergo effective adaptive evolution. Uniparental inheritance of cytoplasmic genomes decreases competition between different beneficial substitutions (clonal interference), promoting the accumulation of beneficial substitutions. Uniparental inheritance also facilitates selection against deleterious cytoplasmic substitutions, slowing Mullers ratchet. In addition, uniparental inheritance generally reduces genetic hitchhiking of deleterious substitutions during selective sweeps. Overall, uniparental inheritance promotes adaptive evolution by increasing the level of beneficial substitutions relative to deleterious substitutions. When we assume that cytoplasmic genome inheritance is biparental, decreasing the number of genomes transmitted during gametogenesis (bottleneck) aids adaptive evolution. Nevertheless, adaptive evolution is always more efficient when inheritance is uniparental. Our findings explain empirical observations that cytoplasmic genomes--despite their asexual mode of reproduction--can readily undergo adaptive evolution.

Evolutionary Biology

Interacting networks of resistance, virulence and core machinery genes identified by genome-wide epistasis analysis

Recent advances in the scale and diversity of population genomic datasets for bacteria now provide the potential for genome-wide patterns of co-evolution to be studied at the resolution of individual bases. The major human pathogen Streptococcus pneumoniae represents the first bacterial organism for which densely enough sampled population data became available for such an analysis. Here we describe a new statistical method, genomeDCA, which uses recent advances in computational structural biology to identify the polymorphic loci under the strongest co-evolutionary pressures. Genome data from over three thousand pneumococcal isolates identified 5,199 putative epistatic interactions between 1,936 sites. Over three-quarters of the links were between sites within the pbp2x, pbp1a and pbp2b genes, the sequences of which are critical in determining non-susceptibility to beta-lactam antibiotics. A network-based analysis found these genes were also coupled to that encoding dihydrofolate reductase, changes to which underlie trimethoprim resistance. Distinct from these resistance genes, a large network component of 384 protein coding sequences encompassed many genes critical in basic cellular functions, while another distinct component included genes associated with virulence. These results have the potential both to identify previously unsuspected protein-protein interactions, as well as genes making independent contributions to the same phenotype. This approach greatly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.\n\nAuthor SummaryEpistatic interactions between polymorphisms in DNA are recognized as important drivers of evolution in numerous organisms. Study of epistasis in bacteria has been hampered by the lack of both densely sampled population genomic data, suitable statistical models and powerful inference algorithms for extremely high-dimensional parameter spaces. We introduce the first model-based method for genome-wide epistasis analysis and use the largest available bacterial population genome data set on Streptococcus pneumoniae (the pneumococcus) to demonstrate its potential for biological discovery. Our approach reveals interacting networks of resistance, virulence and core machinery genes in the pneumococcus, which highlights putative candidates for novel drug targets. Our method significantly enhances the future potential of epistasis analysis for systems biology, and can complement genome-wide association studies as a means of formulating hypotheses for experimental work.

Genetics

Systematic tissue-specific functional annotation of the human genome highlights immune-related DNA elements for late-onset Alzheimer’s disease

Continuing efforts from large international consortia have made genome-wide epigenomic and transcriptomic annotation data publicly available for a variety of cell and tissue types. However, synthesis of these datasets into effective summary metrics to characterize the functional non-coding genome remains a challenge. Here, we present GenoSkyline-Plus, an extension of our previous work through integration of an expanded set of epigenomic and transcriptomic annotations to produce high-resolution, single tissue annotations. After validating our annotations with a catalog of tissue-specific non-coding elements previously identified in the literature, we apply our method using data from 127 different cell and tissue types to present an atlas of heritability enrichment across 45 different GWAS traits. We show that broader organ system categories (e.g. immune system) increase statistical power in identifying biologically relevant tissue types for complex diseases while annotations of individual cell types (e.g. monocytes or B-cells) provide deeper insights into disease etiology. Additionally, we use our GenoSkyline-Plus annotations in an in-depth case study of late-onset Alzheimers disease (LOAD). Our analyses suggest a strong connection between LOAD heritability and genetic variants contained in regions of the genome functional in monocytes. Furthermore, we show that LOAD shares a similar localization of SNPs to monocyte-functional regions with Parkinsons disease. Overall, we demonstrate that integrated genome annotations at the single tissue level provide a valuable tool for understanding the etiology of complex human diseases. Our GenoSkyline-Plus annotations are freely available at http://genocanyon.med.yale.edu/GenoSkyline.\n\nAuthor SummaryAfter years of community efforts, many experimental and computational approaches have been developed and applied for functional annotation of the human genome, yet proper annotation still remains challenging, especially in non-coding regions. As complex disease research rapidly advances, increasing evidence suggests that non-coding regulatory DNA elements may be the primary regions harboring risk variants in human complex diseases. In this paper, we introduce GenoSkyline-Plus, a principled annotation framework to identify tissue and cell type-specific functional regions in the human genome through integration of diverse high-throughput epigenomic and transcriptomic data. Through validation of known non-coding tissue-specific regulatory regions, enrichment analyses on 45 complex traits, and an in-depth case study of neurodegenerative diseases, we demonstrate the ability of GenoSkyline-Plus to accurately identify tissue-specific functionality in the human genome and provide unbiased, genome-wide insights into the genetic basis of human complex diseases.

Genetics