Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Using reference-free compressed data structures to analyse sequencing reads from thousands of human genomes

We are rapidly approaching the point where we have sequenced millions of human genomes. There is a pressing need for new data structures to store raw sequencing data and efficient algorithms for population scale analysis. Current reference based data formats do not fully exploit the redundancy in population sequencing nor take advantage of shared genetic variation. In recent years, the Burrows-Wheeler transform (BWT) and FM-index have been widely employed as a full text searchable index for read alignment and de novo assembly. We introduce the concept of a population BWT and use it to store and index the sequencing reads of 2,705 samples from the 1000 Genomes Project. A key feature is that as more genomes are added, identical read sequences are increasingly observed and compression becomes more efficient. We assess the support in the 1000 Genomes read data for every base position of two human reference assembly versions, identifying that 3.2 Mbp with population support was lost in the transition from GRCh37 with 13.7 Mbp added to GRCh38. We show that the vast majority of variant alleles can be uniquely described by overlapping 31-mers and show how rapid and accurate SNP and indel genotyping can be carried out across the genomes in the population BWT. We use the population BWT to carry out non-reference queries to search for the presence of all known viral genomes, and discover human T-lymphotropic virus 1 integrations in six samples in a recognised epidemiological distribution.

Genomics

The Detailed 3D Multi-Loop Aggregate/Rosette Chromatin Architecture and Functional Dynamic Organization of the Human and Mouse Genomes

The dynamic three-dimensional chromatin architecture of genomes and its co-evolutionary connection to its function - the storage, expression, and replication of genetic information - is still one of the central issues in biology. Here, we describe the much debated 3D-architecture of the human and mouse genomes from the nucleosomal to the megabase pair level by a novel approach combining selective high-throughput high-resolution chromosomal interaction capture (T2C), polymer simulations, and scaling analysis of the 3D-architecture and the DNA sequence: The genome is compacted into a chromatin quasi-fibre with [~]5{+/-}1 nucleosomes/11nm, folded into stable [~]30-100 kbp loops forming stable loop aggregates/rosettes connected by similar sized linkers. Minor but significant variations in the architecture are seen between cell types/functional states. The architecture and the DNA sequence show very similar fine-structured multi-scaling behaviour confirming their co-evolution and the above. This architecture, its dynamics, and accessibility balance stability and flexibility ensuring genome integrity and variation enabling gene expression/regulation by self-organization of (in)active units already in proximity. Our results agree with the heuristics of the field and allow \"architectural sequencing\" at a genome mechanics level to understand the inseparable systems genomic properties.

Genomics

Identification of Klebsiella capsule synthesis loci from whole genome data

BackgroundKlebsiella pneumoniae and close relatives are a growing cause of healthcare-associated infections for which increasing rates of multi-drug resistance are a major concern. The Klebsiella polysaccharide capsule is a major virulence determinant and epidemiological marker. However, little is known about capsule epidemiology since serological typing is not widely accessible, and many isolates are serologically non-typeable. Molecular methods for capsular typing are needed, but existing methods lack sensitivity and specificity and fail to take advantage of the information available in whole-genome sequence data, which is increasingly being generated for surveillance and investigation of Klebsiella.\n\nMethodsWe investigated the diversity of capsule synthesis loci (K loci) among a large, diverse collection of 2503 genome sequences of K. pneumoniae and closely related species. We incorporated analyses of both full-length K locus DNA sequences and clustered protein coding sequences to identify, annotate and compare K locus structures, and we propose a novel method for identifying K loci based on full locus information extracted from whole genome sequences.\n\nResultsA total of 134 distinct K loci were identified, including 31 novel types. Comparative analysis of K locus gene content detected 508 unique protein coding gene clusters that appear to reassort via homologous recombination, generating novel K locus types. Extensive nucleotide diversity was detected among the wzi and wzc genes, both within and between K loci, indicating that current typing schemes based on these genes are inadequate. As a solution, we introduce Kaptive, a novel software tool that automates the process of identifying K loci from large sets of Klebsiella genomes based on full locus information.\n\nConclusionsThis work highlights the extensive diversity of Klebsiella K loci and the proteins that they encode. We propose a standardised K locus nomenclature for Klebsiella, present a curated reference database of all known K loci, and introduce a tool for identifying K loci from genome data (https://github.com/katholt/Kaptive). These developments constitute important new resources for the Klebsiella community for use in genomic surveillance and epidemiology.

Genomics

Identifying Conserved Genomic Elements and Designing Universal Probe Sets To Enrich Them

Targeted enrichment of conserved genomic regions is a popular method for collecting large amounts of sequence data from non-model taxa for phylogenetic, phylogeographic, and population genetic studies. Yet, few open-source workflows are available to identify conserved genomic elements shared among divergent taxa and to design enrichment baits targeting these regions. These shortcomings limit the application of targeted enrichment methods to many organismal groups. Here, I describe a universal workflow for identifying conserved genomic regions in available genomic data and for designing targeted enrichment baits to collect data from these conserved regions. I demonstrate how this computational approach can be applied to diverse organismal groups by identifying sets of conserved loci and designing enrichment baits targeting thousands of these loci in the understudied arthropod groups Arachnida, Coleoptera, Diptera, Hemiptera, or Lepidoptera. I then use in silico analyses to demonstrate that these conserved loci reconstruct the accepted relationships among genome sequences from the focal arthropod orders, and we perform in vitro validation of the Arachnid probe set as part of a separate manuscript (Starrett et al. Submitted). All of the documentation, design steps, software code, and probe sets developed here are available under an open-source license for restriction-free testing and use by any research group, and although the examples in this manuscript focus on understudied and exceptionally diverse arthropod groups, the software workflow is applicable to all organismal groups having some form of pre-existing genomic information.

Genomics

Single-molecule sequencing of the Drosophila serrata genome

Long read sequencing technology promises to greatly enhance de novo assembly of genomes for non-model species. While error rates have been a large stumbling block, sequencing at high coverage allows reads to be self-corrected. Here we sequence and de novo assemble the genome of Drosophila serrata, a non-model species from the montium subgroup that has been well studied for clines and sexual selection. Using 11 PacBio SMRT cells, we generated 12 Gbp of raw sequence data comprising approximately 65x whole genome coverage. Read lengths averaged 8,940 bp (NRead50 12,200) with the longest read at 53 Kbp. We self-corrected reads using the PBDagCon algorithm and assembled the genome using the MHAP algorithm within the PBcR assembler. Total genome length was 198 Mbp with an N50 just under 1 Mbp. Contigs displayed a high degree of arm-level conservation with D. melanogaster. We also provide an initial annotation for this genome using in silico gene predictions that were supported by RNA-seq data.

genomics

It’s okay to be green: Draft genome of the North American Bullfrog (Rana [Lithobates] catesbeiana)

Frogs play important ecological roles as sentinels, insect control and food sources. Several species are important model organisms for scientific research to study embryogenesis, development, immune function, and endocrine signaling. The globally-distributed Ranidae (true frogs) are the largest frog family, and have substantial evolutionary distance from the model laboratory Xenopus frog species. Consequently, the extensive Xenopus genomic resources are of limited utility for Ranids and related frog species. More widely applicable amphibian genomic data is urgently needed as more than two-thirds of known species are currently threatened or are undergoing population declines.\n\nHerein, we report on the first genome sequence of a Ranid species, an adult male North American bullfrog (Rana [Lithobates] catesbeiana). We assembled high-depth Illumina reads (66-fold coverage), into a 5.8 Gbp (NG50 = 57.7 kbp) draft genome using ABySS v1.9.0. The assembly was scaffolded with LINKS and RAILS using pseudo-long-reads from targeted denovo assembler Kollector and Illumina Synthetic Long-Reads, as well as reads from long fragment (MPET) libraries. We predicted over 22,000 protein-coding genes using the MAKER2 pipeline and identified the genomic loci of 6,227 candidate long noncoding RNAs (IncRNAs) from a composite reference bullfrog transcriptome. Mitochondrial sequence analysis supported Lithobates as a subgenus of Rana. RNA-Seq experiments identified ~6,000 thyroid hormone- responsive transcripts in the back skin of premetamorphic tadpoles; the majority of which regulate DNA/RNA processing. Moreover, 1/6th of differentially-expressed transcripts were putative lncRNAs. Our draft bullfrog genome will serve as a useful resource for the amphibian research community.

genomics

Genome-wide detection of structural variants and indels by local assembly

Structural variants (SVs), including small insertion and deletion variants (indels), are challenging to detect through standard alignment-based variant calling methods. Sequence assembly offers a powerful approach to identifying SVs, but is difficult to apply at-scale genome-wide for SV detection due to its computational complexity and the difficulty of extracting SVs from assembly contigs. We describe SvABA, an efficient and accurate method for detecting SVs from short-read sequencing data using genome-wide local assembly with low memory and computing requirements. We evaluated SvABAs performance on the NA12878 human genome and in simulated and real cancer genomes. SvABA demonstrates superior sensitivity and specificity across a large spectrum of SVs, and substantially improved detection performance for variants in the 20-300 bp range, compared with existing methods. SvABA also identifies complex somatic rearrangements with chains of short (< 1,000 bp) templated-sequence insertions copied from distant genomic regions. We applied SvABA to 344 cancer genomes from 11 cancer types, and found that templated-sequence insertions occur in ~4% of all somatic rearrangements. Finally, we demonstrate that SvABA can identify sites of viral integration and cancer driver alterations containing medium-sized SVs.

genomics

Genome-wide identification of lineage and locus specific variation associated with pneumococcal carriage duration

Streptococcus pneumoniae is a leading cause of invasive disease in infants, especially in low-income settings. Asymptomatic carriage in the nasopharynx is a prerequisite for disease, and the duration of carriage is an important consideration in modelling transmission dynamics and vaccine response. Existing studies of carriage duration variability are based at the serotype level only, and do not probe variation within lineages or fully quantify interactions with other environmental factors.\n\nHere we developed a model to calculate the duration of carriage episodes from longitudinal swab data. By combining these results with whole genome sequence data we estimate that pneumococcal genomic variation accounted for 63% of the phenotype variation, whereas host traits accounted for less than 5%. We further partitioned this heritability into both lineage and locus effects, and quantified the amount attributable to the largest sources of variation in carriage duration: serotype (17%), drug-resistance (9%) and other significant locus effects (7%). For the locus effects, a genome-wide association study identified 16 loci which may have an effect on carriage duration independent of serotype. Hits at a genome-wide level of significance were to prophage sequences, suggesting infection by such viruses substantially affects carriage duration.\n\nThese results show that both serotype and non-serotype specific effects alter carriage duration in infants and young children and are more important than other environmental factors such as host genetics. This has implications for models of pneumococcal competition and antibiotic resistance, and leads the way for the analysis of heritability of complex bacterial traits.\n\nSignificance statementOther than serotype, the genetic determinants of pneumococcal carriage duration are unknown. In this study we used longitudinal sampling to measure the duration of carriage in infants, and searched for any associated variation in the pan-genome. While we found that the pathogen genome explains most of the variability in duration, serotype did not fully account for this. Recent theoretical work has proposed the existence of alleles which alter carriage duration to explain the puzzle of continued coexistence of antibiotic-resistant and sensitive strains. Here we have shown that these alleles do exist in a natural population, and also identified candidates for the loci which fulfil this role. Together these findings have implications for future modelling of pneumococcal epidemiology and resistance.

genomics

An Integrative Framework For Detecting Structural Variations In Cancer Genomes

Structural variants can contribute to oncogenesis through a variety of mechanisms, yet, despite their importance, the identification of structural variants in cancer genomes remains challenging. Here, we present an integrative framework for comprehensively identifying structural variation in cancer genomes. For the first time, we apply next-generation optical mapping, high-throughput chromosome conformation capture (Hi-C), and whole genome sequencing to systematically detect SVs in a variety of cancer cells.\n\nUsing this approach, we identify and characterize structural variants in up to 29 commonly used normal and cancer cell lines. We find that each method has unique strengths in identifying different classes of structural variants and at different scales, suggesting that integrative approaches are likely the only way to comprehensively identify structural variants in the genome. Studying the impact of the structural variants in cancer cell lines, we identify widespread structural variation events affecting the functions of non-coding sequences in the genome, including the deletion of distal regulatory sequences, alteration of DNA replication timing, and the creation of novel 3D chromatin structural domains.\n\nThese results underscore the importance of comprehensive structural variant identification and indicate that non-coding structural variation may be an underappreciated mutational process in cancer genomes.

genomics

Comparative Genomics Shows That Viral Integrations Are Abundant And Express piRNAs In The Arboviral Vectors Aedes aegypti And Aedes albopictus

BackgroundArthropod-borne viruses (arboviruses) transmitted by mosquito vectors cause many important emerging or resurging infectious diseases in humans including dengue, chikungunya and Zika. Understanding the co-evolutionary processes among viruses and vectors is essential for the development of novel transmission-blocking strategies. Arboviruses form episomal viral DNA fragments upon infection of mosquito cells and adults. Additionally, sequences from insect-specific viruses and arboviruses have been found integrated into mosquito genomes.\n\nResultsWe used a bioinformatic approach to analyze the presence, abundance, distribution, and transcriptional activity of integrations from 425 non-retroviral viruses, including 133 arboviruses, across the presently available 22 mosquito genome sequences. Large differences in abundance and types of viral integrations were observed in mosquito species from the same region. Viral integrations are unexpectedly abundant in the arboviral vector species Aedes aegypti and Ae. albopictus, but are [~]10-fold less abundant in all other mosquitoes analysed. Additionally, viral integrations are enriched in piRNA clusters of both the Ae. aegypti and Ae. albopictus genomes and, accordingly, they express piRNAs, but not siRNAs.\n\nConclusionsDifferences in number of viral integrations in the genomes of mosquito species from the same geographic area support the conclusion that integrations of viral sequences is not dependent on viral exposure, but that lineage-specific interactions exits. Viral integrations are abundant in Ae. aegypti and Ae. albopictus, and represent a thus far unappreciated component of their genomes. Additionally, the genome locations of viral integrations and their production of piRNAs indicate a functional link between viral integrations and the piRNA pathway. These results greatly expand the breadth and complexity of small RNA-mediated regulation and suggest a role for viral integrations in antiviral defense in these two mosquito species.

genomics

Framework For Quality Assessment Of Whole Genome, Cancer Sequences

Working with cancer whole genomes sequenced over a period of many years in different sequencing centres requires a validated framework to compare the quality of these sequences. The Pan-Cancer Analysis of Whole Genomes (PCAWG) of the International Cancer Genome Consortium (ICGC), a project a cohort of over 2800 donors provided us with the challenge of assessing the quality of the genome sequences. A non-redundant set of five quality control (QC) measurements were assembled and used to establish a star rating system. These QC measures reflect known differences in sequencing protocol and provide a guide to downstream analyses of these whole genome sequences. The resulting QC measures also allowed for exclusion samples of poor quality, providing researchers within PCAWG, and when the data is released for other researchers, a good idea of the sequencing quality. For a researcher wishing to apply the QC measures for their data we provide a Docker Container of the software used to calculate them. We believe that this is an effective framework of quality measures for whole genome, cancer sequences, which will be a useful addition to analytical pipelines, as it has to the PCAWG project.

genomics

A Model-Free Approach For Detecting Genomic Regions Of Deep Divergence Using The Distribution Of Haplotype Distances

Recent advances in comparative genomics have revealed that divergence between populations is not necessarily uniform across all parts of the genome. There are examples of regions with divergent haplotypes that are substantially more different from each other that the genomic average.\n\nTypically, these regions are of interest, as their persistence over long periods of time may reflect balancing selection. However, they are hard to detect unless the divergent sub-populations are known prior to analysis.\n\nHere, we introduce HaploDistScan, an R-package implementing model-free detection of deep-divergence genomic regions based on the distribution of pair-wise haplotype distances, and show that it can detect such regions without use of a priori information about population sub-division. We apply the method to real-world data sets, from ruff and Darwins finches, and show that we are able to recover known instances of balancing selection - originally identified in studies reliant on detailed phenotyping - using only genotype data. Furthermore, in addition to replicating previously known divergent haplotypes as a proof-of-concept, we identify novel regions of interest in the Darwins finch genome and propose a plausible, data-driven evolutionary history for each novel locus individually.\n\nIn conclusion, HaploDistScan requires neither phenotypic nor demographic input data, thus filling a gap in the existing set of methods for genome scanning, and provides a useful tool for identification of regions under balancing selection or similar evolutionary processes.

genomics

The Plastid Genome In Cladophorales Green Algae Is Encoded By Hairpin Plasmids

Virtually all plastid (chloroplast) genomes are circular double-stranded DNA molecules, typically between 100-200 kb in size and encoding circa 80-250 genes. Exceptions to this universal plastid genome architecture are very few and include the dinoflagellates where genes are located on DNA minicircles. Here we report on the highly deviant chloroplast genome of Cladophorales green algae, which is entirely fragmented into hairpin plasmids. Short and long read high-throughput sequencing of DNA and RNA demonstrated that the chloroplast genes of Boodlea composita are encoded on 1-7 kb DNA contigs with an exceptionally high GC-content, each containing a long inverted repeat with one or two protein-coding genes and conserved non-coding regions putatively involved in replication and/or expression. We propose that these contigs correspond to linear single-stranded DNA molecules that fold onto themselves to form hairpin plasmids. The Boodlea chloroplast genes are highly divergent from their corresponding orthologs. The origin of this highly deviant chloroplast genome likely occurred before the emergence of the Cladophorales, and coincided with an elevated transfer of chloroplast genes to the nucleus. A chloroplast genome that is composed only of linear DNA molecules is unprecedented among eukaryotes and highlights unexpected variation in the plastid genome architecture.

genomics

High contiguity Arabidopsis thaliana genome assembly with a single nanopore flow cell

While many evolutionary questions can be answered by short read re-sequencing, presence/absence polymorphisms of genes and/or transposons have been largely ignored in large-scale intraspecific evolutionary studies. To enable the rigorous analysis of such variants, multiple high quality and contiguous genome assemblies are essential. Similarly, while genome assemblies based on short reads have made genomics accessible for non-reference species, these assemblies have limitations due to low contiguity. Long-read sequencers and long-read technologies have ushered in a new era of genome sequencing where the lengths of reads exceed those of most repeats. However, because these technologies are not only costly, but also time and compute intensive, it has been unclear how scalable they are. Here we demonstrate a fast and cost effective reference assembly for an Arabidopsis thaliana accession using the USB-sized Oxford Nanopore MinION sequencer and typical consumer computing hardware (4 Cores, 16Gb RAM). We assemble the accession KBS-Mac-74 into 62 contigs with an N50 length of 12.3 Mb covering 100% (119 Mb) of the non-repetitive genome. We demonstrate that the polished KBS-Mac-74 assembly is highly contiguous with BioNano optical genome maps, and of high per-base quality against a likewise polished Pacific Biosciences long-read assembly. The approach we implemented took a total of four days at a cost of less than 1,000 USD for sequencing consumables including instrument depreciation.

genomics

De novo draft assembly of the Botrylloides leachii genome provides further insight into tunicate evolution.

Tunicates are marine invertebrates that compose the closest phylogenetic group to the vertebrates. This chordate subphylum contains a particularly diverse range of reproductive methods, regenerative abilities and life-history strategies. Consequently, tunicates provide an extraordinary perspective into the emergence and diversity of chordate traits. To gain further insights into the evolution of the tunicate phylum, we have sequenced the genome of the colonial Stolidobranchian Botrylloides leachii.\n\nWe have produced a high-quality (90 % BUSCO genes) 159 Mb assembly, containing 82 % of the predicted total 194 Mb genomic content. The B. leachii genome is much smaller than that of Botryllus schlosseri (725 Mb), but comparable to those of Ciona robusta and Molgula oculata (both 160 Mb). We performed an orthologous clustering between five tunicate genomes that highlights sets of genes specific to some species, including a large group unique to colonial ascidians with gene ontology terms including cell communication and immune response.\n\nBy analysing the structure and composition of the conserved gene clusters, we identified many examples of multiple cluster breaks and gene dispersion, suggesting that several lineage-specific genome rearrangements occurred during tunicate evolution. In addition, we investigate lineage-specific gene gain and loss within the Wnt, Notch and retinoic acid pathways. Such examples of genetic change within these highly evolutionary conserved pathways commonly associated with regeneration and development may underlie some of the diverse regenerative abilities observed in the tunicate subphylum. These results supports the widely held view that tunicate genomes are evolving particularly rapidly.

genomics

Analysis of the Aedes albopictus C6/36 genome provides insight into cell line adaptations to in vitro viral propagation

BackgroundThe 50-year old Aedes albopictus C6/36 cell line is a resource for the detection, amplification, and analysis of mosquito-borne viruses including Zika, dengue, and chikungunya. The cell line is derived from an unknown number of larvae from an unspecified strain of Aedes albopictus mosquitoes. Toward improved utility of the cell line for research in virus transmission, we present an annotated assembly of the C6/36 genome.\n\nResultsThe C6/36 genome assembly has the largest contig N50 (3.3 Mbp) of any mosquito assembly, presents the sequences of both haplotypes for most of the diploid genome, reveals independent null mutations in both alleles of the Dicer locus, and indicates a male-specific genome. Gene annotation was computed with publicly available mosquito transcript sequences. Gene expression data from cell line RNA sequence identified enrichment of growth-related pathways and conspicuous deficiency in aquaporins and inward rectifier K+ channels. As a test of utility, RNA sequence data from Zika-infected cells was mapped to the C6/36 genome and transcriptome assemblies. Host subtraction reduced the data set by 89%, enabling faster characterization of non-host reads.\n\nConclusionsThe C6/36 genome sequence and annotation should enable additional uses of the cell line to study arbovirus vector interactions and interventions aimed at restricting the spread of human disease.

genomics

The first near-complete assembly of the hexaploid bread wheat genome, Triticum aestivum

Common bread wheat, Triticum aestivum, has one of the most complex genomes known to science, with 6 copies of each chromosome, enormous numbers of near-identical sequences scattered throughout, and an overall size of more than 15 billion bases. Multiple past attempts to assemble the genome have failed. Here we report the first successful assembly of T. aestivum, using deep sequencing coverage from a combination of short Illumina reads and very long Pacific Biosciences reads. The final assembly contains 15,344,693,583 bases and has a weighted average (N50) contig size of of 232,659 bases. This represents by far the most complete and contiguous assembly of the wheat genome to date, providing a strong foundation for future genetic studies of this important food crop. We also report how we used the recently published genome of Aegilops tauschii, the diploid ancestor of the wheat D genome, to identify 4,179,762,575 bp of T. aestivum that correspond to its D genome components.

genomics

Pan-cancer analysis of whole genomes

We report the integrative analysis of more than 2,600 whole cancer genomes and their matching normal tissues across 39 distinct tumour types. By studying whole genomes we have been able to catalogue non-coding cancer driver events, study patterns of structural variation, infer tumour evolution, probe the interactions among variants in the germline genome, the tumour genome and the transcriptome, and derive an understanding of how coding and non-coding variations together contribute to driving individual patient's tumours. This work represents the most comprehensive look at cancer whole genomes to date. NOTE TO READERS: This is an incomplete draft of the marker paper for the Pan-Cancer Analysis of Whole Genomes Project, and is intended to provide the background information for a series of in-depth papers that will be posted to BioRixv during the summer of 2017.

cancer biology