Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

SplitThreader: Exploration and analysis of rearrangements in cancer genomes

Genomic rearrangements and associated copy number changes are important drivers in cancer as they can alter the expression of oncogenes and tumor suppressors, create gene fusions, and misregulate gene expression. Here we present SplitThreader (http://splitthreader.com), an open-source interactive web application for analysis and visualization of genomic rearrangements and copy number variation in cancer genomes. SplitThreader constructs a sequence graph of genomic rearrangements in the sample and uses a priority queue breadth-first search algorithm on the graph to search for novel interactions. This is applied to detect gene fusions and other novel sequences, as well as to evaluate distances in the rearranged genome between any genomic regions of interest, especially the repositioning of regulatory elements and their target genes. SplitThreader also analyzes each variant to categorize it by its relation to other variants and by its copy number concordance. This identifies balanced translocations, identifies simple and complex variants, and suggests likely false positives when copy number is not concordant across a candidate breakpoint. It also provides explanations when multiple variants affect the copy number state and obscure the contribution of a single variant, such as a deletion within a region that is overall amplified. Together, these categories triage the variants into groups and provide a starting point for further systematic analysis and manual curation. To demonstrate its utility, we apply SplitThreader to three cancer cell lines, MCF-7 and A549 with Illumina paired-end sequencing, and SK-BR-3, with long-read PacBio sequencing. Using SplitThreader, we examine the genomic rearrangements responsible for previously observed gene fusions in SK-BR-3 and MCF-7, and discover many of the fusions involved a complex series of multiple genomic rearrangements. We also find notable differences in the types of variants between the three cell lines, in particular a much higher proportion of reciprocal variants in SK-BR-3 and a distinct clustering of interchromosomal variants in SK-BR-3 and MCF-7 that is absent in A549.

bioinformatics

Chromosome assembly of large and complex genomes using multiple references

Despite the rapid development of sequencing technologies, assembly of mammalian-scale genomes into complete chromosomes remains one of the most challenging problems in bioinformatics. To help address this difficulty, we developed Ragout, a reference-assisted assembly tool that now works for large and complex genomes. Taking one or more target assemblies (generated from an NGS assembler) and one or multiple related reference genomes, Ragout infers the evolutionary relationships between the genomes and builds the final assemblies using a genome rearrangement approach. Using Ragout, we transformed NGS assemblies of 15 different Mus musculus and one Mus spretus genomes into sets of complete chromosomes, leaving less than 5% of sequence unlocalized per set. Various benchmarks, including PCR testing and realigning of long PacBio reads, suggest only a small number of structural errors in the final assemblies, comparable with direct assembly approaches. Additionally, we applied Ragout to Mus caroli and Mus pahari genomes, which exhibit karyotype-scale variations compared to other genomes from the Muridae family. Chromosome color maps confirmed most large-scale rearrangements that Ragout detected.

bioinformatics

Inference of candidate germline mutator loci in humans from genome-wide haplotype data

The rate of germline mutation varies widely between species but little is known about the extent of variation in the germline mutation rate between individuals of the same species. Here we demonstrate that an allele that increases the rate of germline mutation can result in a distinctive signature in the genomic region linked to the affected locus, characterized by a number of haplotypes with a locally high proportion of derived alleles, against a background of haplotypes carrying a typical proportion of derived alleles. We searched for this signature in human haplotype data from phase 3 of the 1000 Genomes Project and report a number of candidate mutator loci, several of which are located close to or within genes involved in DNA repair or the DNA damage response. To investigate whether mutator alleles remained active at any of these loci, we used de novo mutation counts from human parent-offspring trios in the 1000 Genomes and Genome of the Netherlands cohorts, looking for an elevated number of de novo mutations in the offspring of parents carrying a candidate mutator haplotype at each of these loci. We found some support for two of the candidate loci, including one locus just upstream of the BRSK2 gene, which is expressed in the testis and has been reported to be involved in the response to DNA damage.\n\nAuthor SummaryEach time a genome is replicated there is the possibility of error resulting in the incorporation of an incorrect base or bases in the genome sequence. When these errors occur in cells that lead to the production of gametes they can be incorporated into the germline. Such germline mutations are the basis of evolutionary change; however, to date there has been little attempt to quantify the extent of genetic variation in human populations in the rate at which they occur. This is particularly important because new spontaneous mutations are thought to make an important contribution to many human diseases. Here we present a new way to identify genetic loci that may be associated with an elevated rate of germline mutation and report the application of this method to data from a large number of human genomes, generated by the 1000 Genomes Project. Several of the candidate loci we report are in or near genes involved in DNA repair and some were supported by direct measurement of the mutation rate obtained from parent-offspring trios.

genetics

BGDMdocker: an workflow base on Docker for analysis and visualization pan-genome and biosynthetic gene clusters of Bacterial

MotivationAt present Docker technology has received increasing level of attention throughout the bioinformatics community. However, its implementation details have not yet been mastered by most biologists and applied widely in biological researches. In order to popularizing this technology in the bioinformatics and sufficiently use plenty of public resources of bioinformatics tools (Dockerfile and image of scommunity, officially and privately) in Docker Hub Registry and other Docker sources based on Docker, we introduced full and accurate instance of a bioinformatics workflow based on Docker to analyse and visualize pan-genome and biosynthetic gene clusters of a bacteria in this article, provided the solutions for mining bioinformatics big data from various public biology databases. You could be guided step-by-step through the workflow process from docker file to build up your own images and run an container fast creating an workflow.\n\nResultsWe presented a BGDMdocker (bacterial genome data mining docker-based) workflow based on docker. The workflow consists of three integrated toolkits, Prokka v1.11, panX, and antiSMASH3.0. The dependencies were all written in Dockerfile, to build docker image and run container for analysing pan-genome of total 44 Bacillus amyloliquefaciens strains, which were retrieved from public? database. The pan-genome totally includes 172,432 gene, 2,306 Core gene cluster. The visualized pan-genomic data such as alignment, phylogenetic trees, maps mutations within that cluster to the branches of the tree, infers loss and gain of genes on the core-genome phylogeny for each gene cluster were presented. Besides, 997 known (MIBiG database) and 553 unknown (antiSMASH-predicted clusters and Pfam database) genes of biosynthesis gene clusters types and orthologous groups were mined in all strains. This workflow could also be used for other species pan-genome analysis and visualization. The display of visual data can completely duplicated as well as done in this paper. All result data and relevant tools and files can be downloaded from our website with no need to register. The pan-genome and biosynthetic gene clusters analysis and visualization can be fully reusable immediately in different computing platforms (Linux, Windows, Mac and deployed in the cloud), achieved cross platform deployment flexibility, rapid development integrated software package.\n\nAvailability and implementationBGDMdocker is available at http://42.96.173.25/bapgd/ and the source code under GPL license is available at https://github.com/cgwyx/debian_prokka_panx_antismash_biodocker.\n\nContactchenggongwyx@foxmail.com\n\nSupplementary informationSupplementary data are available at biorxiv online.

bioinformatics

OrthoFiller: utilising data from multiple species to improve the completeness of genome annotations.

BackroundComplete and accurate annotation of sequenced genomes is of paramount importance to their utility and analysis. Differences in gene prediction pipelines mean that genome sequences for a species can differ considerably in the quality and quantity of their predicted genes. Furthermore, genes that are present in genome sequences sometimes fail to be detected by computational gene prediction methods. Erroneously unannotated genes can lead to oversights and inaccurate assertions in biological investigations, especially for smaller-scale genome projects which rely heavily on computational prediction.\n\nResultsHere we present OrthoFiller, a tool designed to address the problem of finding and adding such missing genes to genome annotations. OrthoFiller leverages information from multiple related species to identify those genes whose existence can be verified through comparison with known gene families, but which have not been predicted. By simulating missing gene annotations in real sequence datasets from both plants and fungi we demonstrate the accuracy and utility of OrthoFiller for finding missing genes and improving genome annotation. Furthermore, we show that applying OrthoFiller to existing \"complete\" genome annotations can identify and correct substantial numbers of erroneously missing genes in these two sets of species.\n\nConclusionsWe show that significant improvements in the completeness of genome annotations can be made by leveraging information from multiple species.

bioinformatics

The mitochondrial genomes of the acoelomorph wormsParatomella rubra and Isodiametra pulchra

Acoels are small, ubiquitous, but understudied, marine worms with a very simple body plan. Their internal phylogeny is still in parts unresolved, and the position of their proposed phylum Xenacoelomorpha (Xenoturbella+Acoela) is still debated.\n\nHere we describe mitochondrial genome sequences from two acoel species: Paratomella rubra and Isodiametra pulchra. The 14,954 nucleotide-long P. rubra sequence is typical for metazoans in size and gene content. The larger I. pulchra mitochondrial genome contains both ribosomal genes, 21 tRNAs, but only 11 protein-coding genes. We find evidence suggesting a duplicated sequence in the I. pulchra mitochondrial genome.\n\nMitochondrial sequences for both P. rubra and I. pulchra have a unique genome organisation in comparison to other published metazoan mitochondrial genomes. We found a large degree of protein-coding gene and tRNA overlap in P. rubra, with little non-coding sequence making the genome compact. Conversely, the I. pulchra mitochondrial genome has many long non-coding sequences between genes, likely driving the genome size expansion. Phylogenetic trees inferred from concatenated alignments of mitochondrial genes grouped the fast-evolving Acoela and Tunicata, almost certainly due to the systematic error of long branch attraction: a reconstruction artefact that is probably compounded by the fast substitution rate of mitochondrial genes in this taxon.

evolutionary biology

VICTOR: Genome-based Phylogeny and Classification of Prokaryotic Viruses

Bacterial and archaeal viruses (\"phages\") play an enormous role in global life cycles and have recently regained importance as therapeutic agents to fight serious infections by multi-resistant bacterial strains. Nevertheless, taxonomic classification of phages is up to now only insufficiently informed by genome sequencing. Despite thousands of publicly available phage genomes, it still needs to be investigated how this wealth of information can be used for the fast, universal and accurate classification of phages. The Genome BLAST Distance Phylogeny (GBDP) approach is a truly whole-genome method currently used for in silico DNA: DNA hybridization and phylogenetic inference from prokaryotic genomes. Based on the principles of phylogenetic systematics, we here established GBDP for phage phylogeny and classification, using the common subset of genome-sequenced and officially classified phages. Trees inferred with the best GBDP variants showed only few deviations from the official phage classification, which were uniformly due to incorrectly annotated GenBank entries. Except for low resolution at the family level, the majority of taxa was well supported as monophyletic. Clustering genome sequences with distance thresholds optimized for the agreement with the classification turned out to be phylogenetically reasonable. Accordingly modifying genera and species is taxonomically optional but would yield more uniform sequence divergence as well as stronger branch support. Analysing an expanded data set containing > 4000 phage genomes from public databases allowed for extrapolating regarding the number, composition and host specificity of future phage taxa. The selected methods are implemented in an easy-to-use web service \"VICTOR\" freely available at http://ggdc.dsmz.de/victor.php.

evolutionary biology

Dense And Accurate Whole-Chromosome Haplotyping Of Individual Genomes

The diploid nature of the genome is neglected in many analyses done today, where a genome is perceived as a set of unphased variants with respect to a reference genome. Many important biological phenomena such as compound heterozygosity and epistatic effects between enhancers and target genes, however, can only be studied when haplotype-resolved genomes are available. This lack of haplotype-level analyses can be explained by a dearth of methods to produce dense and accurate chromosome-length haplotypes at reasonable costs. Here we introduce an integrative phasing strategy that combines global, but sparse haplotypes obtained from strand-specific single cell sequencing (Strand-seq) with dense, yet local, haplotype information available through long-read or linked-read sequencing. Our experiments provide comprehensive guidance on favorable combinations of Strand-seq libraries and sequencing coverages to obtain complete and genome-wide haplotypes of a single individual genome (NA12878) at manageable costs. We were able to reliably assign > 95% of alleles to their parental haplotypes using as few as 10 Strand-seq libraries in combination with 10-fold coverage PacBio data or, alternatively, 10X Genomics linked-read sequencing data. We conclude that the combination of Strand-seq with different sequencing technologies represents an attractive solution to chart the unique genetic variation of diploid genomes.

bioinformatics

Pooled CRISPR interference screens enable high-throughput functional genomics study and elucidate new rules for guide RNA library design in Escherichia coli

Clustered regularly interspaced short palindromic repeat (CRISPR)/Cas9 technology provides potential advantages in high-throughput functional genomics analysis in prokaryotes over previously established platforms based on recombineering or transposon mutagenesis. In this work, as a proof-of-concept to adopt CRISPR/Cas9 method as a pooled functional genomics analysis platform in prokaryotes, we developed a CRISPR interference (CRISPRi) library consisting of 3,148 single guide RNAs (sgRNAs) targeting the open reading frame (ORF) of 67 genes with known knockout phenotypes and performed pooled screens under two stressed conditions (minimal and acidic medium) in Escherichia coli. Our approach confirmed most of previously described gene-phenotype associations while maintaining < 5% false positive rate, suggesting that CRISPRi screen is both sensitive and specific. Our data also supported the ability of this method to narrow down the candidate gene pool when studying operons, a unique structure in prokaryotic genome. Meanwhile, assessment of multiple loci across treatments enables us to extract several guidelines for sgRNA design for such pooled functional genomics screen. For instance, sgRNAs locating at the first 5% upstream region within ORF exhibit enhanced activity and 10 sgRNAs per gene is suggested to be enough for robust identification of gene-phenotype associations. We also optimized the hit-gene calling algorithm to identify target genes more robustly with even fewer sgRNAs. This work showed that CRISPRi could be adopted as a powerful functional genomics analysis tool in prokaryotes and provided the first guideline for the construction of sgRNA libraries in such applications.\n\nImportanceTo fully exploit the valuable resource of explosive sequenced microbial genomes, high-throughput experimental platform is needed to associate genes and phenotypes at the genome level, giving microbiologists the insight about the genetic structure and physiology of a microorganism. In this work, we adopted CRISPR interference method as a pooled high-throughput functional genomics platform in prokaryotes with Escherichia coli as the model organism. Our data suggested that this method was highly sensitive and specific to map genes with previously known phenotypes, potent to act as a new strategy for high-throughput microbial genetics study with advantages over previously established methods. We also provided the first guideline for the sgRNA library design by comprehensive analysis of the screen data. The concept, gRNA library design rules and open-source scripts of this work should benefit prokaryotic genetics community to apply high-throughput mapping of defined gene set with phenotypes in a broad spectrum of microorganisms.

microbiology

Ultrafast comparison of personal genomes

We present an ultra-fast method for comparing personal genomes. We transform the standard genome representation (lists of variants relative to a reference) into genome fingerprints that can be readily compared across sequencing technologies and reference versions. Because of their reduced size, computation on the genome fingerprints is fast and requires little memory. This enables scaling up a variety of important genome analyses, including quantifying relatedness, recognizing duplicative sequenced genomes in a set, population reconstruction, and many others. The original genome representation cannot be reconstructed from its fingerprint; the method thus has significant implications for privacy-preserving genome analytics.

bioinformatics

Enhanced Guide-RNA Design And Targeting Analysis For Precise CRISPR Genome Editing Of Single And Consortia Of Industrially Relevant And Non-Model Organisms

MotivationGenetic diversity of non-model organisms offers a repertoire of unique phenotypic features for exploration and cultivation for synthetic biology and metabolic engineering applications. To realize this enormous potential, it is critical to have an efficient genome editing tool for rapid strain engineering of these organisms to perform novel programmed functions.\n\nResultsTo accommodate the use of CRISPR/Cas systems for genome editing across organisms, we have developed a novel method, named CASPER (CRISPR Associated Software for Pathway Engineering and Research), for identifying on- and off-targets with enhanced predictability coupled with an analysis of non-unique (repeated) targets to assist in editing any organism with various endonucleases. Utilizing CASPER, we demonstrated a modest 2.4% and significant 30.2% improvement (F-test, p<0.05) over the conventional methods for predicting on- and off-target activities, respectively. Further we used CASPER to develop novel applications in genome editing: multitargeting analysis (i.e. simultaneous multiple-site modification on a target genome with a sole guide-RNA (gRNA) requirement) and multispecies population analysis (i.e. gRNA design for genome editing across a consortium of organisms). Our analysis on a selection of industrially relevant organisms revealed a number of non-unique target sites associated with genes and transposable elements that can be used as potential sites for multitargeting. The analysis also identified shared and unshared targets that enable genome editing of single or multiple genomes in a consortium of interest. We envision CASPER as a useful platform to enhance the precise CRISPR genome editing for metabolic engineering and synthetic biology applications.

bioinformatics

Precision genome editing using synthesis-dependent repair of Cas9-induced DNA breaks

The RNA-guided DNA endonuclease Cas9 has emerged as a powerful new tool for genome engineering. Cas9 creates targeted double-strand breaks (DSBs) in the genome. Knock-in of specific mutations (precision genome editing) requires homology-directed repair (HDR) of the DSB by synthetic donor DNAs containing the desired edits, but HDR has been reported to be variably efficient. Here, we report that linear DNAs (single and double-stranded) engage in a high-efficiency HDR mechanism that requires only [~]35 nucleotides of homology with the targeted locus to introduce edits ranging from 1 to 1000 nucleotides. We demonstrate the utility of linear donors by introducing fluorescent protein tags in human cells and mouse embryos using PCR fragments. We find that repair is local, polarity-sensitive, and prone to template switching, characteristics that are consistent with gene conversion by synthesis-dependent strand-annealing (SDSA). Our findings enable rational design of synthetic donor DNAs for efficient genome editing.\n\nSignificanceGenome editing, the introduction of precise changes in the genome, is revolutionizing our ability to decode the genome. Here we describe a simple method for genome editing that takes advantage of an efficient mechanism for DNA repair called synthesis-dependent strand annealing. We demonstrate that synthetic linear DNAs (ssODNs and PCR fragments) with [~]35bp homology arms function as efficient donors for SDSA repair of Cas9-induced double-strand breaks. Edits from 1 to 1000 base pairs can be introduced in the genome without cloning or selection.

synthetic biology

Mitotic systemic genomic instability in yeast

Conventional models of genome evolution generally include the assumption that mutations accumulate gradually and independently over time. We characterized the occurrence of sudden spikes in the accumulation of genome-wide loss-of-heterozygosity (LOH) in Saccharomyces cerevisiae, suggesting the existence of a mitotic systemic genomic instability process (mitSGI). We characterized the emergence of a rough colony morphology phenotype resulting from an LOH event spanning a specific locus (ACE2/ace2-A7). Surprisingly, half of the clones analyzed also carried unselected secondary LOH tracts elsewhere in their genomes. The number of secondary LOH tracts detected was 20-fold higher than expected assuming independence between mutational events. Secondary LOH tracts were not detected in control clones without a primary selected LOH event. We then measured the rates of single and double LOH at different chromosome pairs and found that coincident LOH accumulated at rates 30-100 fold higher than expected if the two underlying single LOH events occurred independently. These results were consistent between two different strain backgrounds, and in mutant strains incapable of entering meiosis. Our results indicate that a subset of mitotic cells within a population experience systemic genomic instability episodes, resulting in multiple chromosomal rearrangements over one or few generations. They are reminiscent of early reports from the classic yeast genetics literature, as well as recent studies in humans, both in the cancer and genomic disorder contexts, all of which challenge the idea of gradual accumulation of structural genomic variation. Our experimental approach provides a model to further dissect the fundamental mechanisms responsible for mitSGI.\n\nSIGNIFICANCE STATEMENTPoint mutations and alterations in chromosome structure are generally thought to accumulate gradually and independently over many generations. Here, we combined complementary genetic approaches in budding yeast to track the appearance of chromosomal changes resulting in loss-of-heterozygosity (LOH). Contrary to expectations, our results provided evidence for the occurrence of non-independent accumulation of multiple LOH events over one or a few cell generations. These results are analogous to recent reports of bursts of chromosomal instability in humans. Our experimental approach provides a framework to further dissect the fundamental mechanisms underlying systemic chromosomal instability processes, including in the human cancer and genomic disorder contexts.

genetics

Genome downsizing, physiological novelty, and the global dominance of flowering plants

During the Cretaceous (145-66 Ma), early angiosperms rapidly diversified, eventually outcompeting the ferns and gymnosperms previously dominating most ecosystems. Heightened competitive abilities of angiosperms are often attributed to higher rates of transpiration facilitating faster growth. This hypothesis does not explain how angiosperms were able to develop leaves with smaller, but densely packed stomata and highly branched venation networks needed to support increased gas exchange rates. Although genome duplication and reorganization have likely facilitated angiosperm diversification, here we show that genome downsizing facilitated reductions in cell size necessary to construct leaves with a high density stomata and veins. Rapid genome downsizing during the early Cretaceous allowed angiosperms to push the frontiers of anatomical trait space. In contrast, during the same time period ferns and gymnosperms exhibited no such changes in genome size, stomatal size, or vein density. Further reinforcing the effect of genome downsizing on increased gas exchange rates, we found that species employing water-loss limiting crassulacean acid metabolism (CAM) photosynthesis, have significantly larger genomes than C3 and C4 species. By directly affecting cell size and gas exchange capacity, genome downsizing brought actual primary productivity closer to its maximum potential. These results suggest species with small genomes, exhibiting a larger range of final cell size, can more finely tune their leaf physiology to environmental conditions and inhabit a broader range of habitats.

ecology

iSeg: an efficient algorithm for segmentation of genomic and epigenomic data

BackgroundIdentification of functional elements of a genome often requires dividing a sequence of measurements along a genome into segments where adjacent segments have different properties, such as different mean values. This problem is often called the segmentation problem in the field of genomics, and the change-point problem in other scientific disciplines. Despite dozens of algorithms developed to address this problem in genomics research, methods with improved accuracy and speed are still needed to effectively tackle both existing and emerging genomic and epigenomic segmentation problems.\n\nResultsWe designed an efficient algorithm, called iSeg, for segmentation of genomic and epigenomic profiles. iSeg first utilizes dynamic programming to identify candidate segments and test for significance. It then uses a novel data structure based on two coupled balanced binary trees to detect overlapping significant segments and update them simultaneously during searching and refinement stages. Refinement and merging of significant segments are performed at the end to generate the final set of segments. By using an objective function based on the p-values of the segments, the algorithm can serve as a general computational framework to be combined with different assumptions on the distributions of the data. As a general segmentation method, it can segment different types of genomic and epigenomic data, such as DNA copy number variation, nucleosome occupancy, nuclease sensitivity, and differential nuclease sensitivity data. Using simple t-tests to compute p-values across multiple datasets of different types, we evaluate iSeg using both simulated and experimental datasets and show that it performs satisfactorily when compared with some other popular methods, which often employ more sophisticated statistical models. Implemented in C++, iSeg is also very computationally efficient, well suited for large numbers of input profiles and data with very long sequences.\n\nConclusionsWe have developed an effective and efficient general-purpose segmentation tool for sequential data and illustrated its use in segmentation of genomic and epigenomic profiles.

bioinformatics

MatchMiner: An open source computational platform for real-time matching of cancer patients to precision medicine clinical trials using genomic and clinical criteria

BackgroundMolecular profiling of cancers is now routine at many cancer centers, and the number of precision cancer medicine clinical trials, which are informed by profiling, is steadily rising. Additionally, these trials are becoming increasingly complex, often having multiple arms and many genomic eligibility criteria. Currently, it is a challenging for physicians to match patients to relevant clinical trials using the patients genomic profile, which can lead to missed opportunities. Automated matching against uniformly structured and encoded genomic eligibility criteria is essential to keep pace with the complex landscape of precision medicine clinical trials.\n\nResultsTo meet these needs, we built and deployed an automated clinical trial matching platform called MatchMiner at the Dana-Farber Cancer Institute (DFCI). The platform has been integrated with Profile, DFCIs enterprise genomic profiling project, which contains tumor profile data for >20,000 patients, and has been made available to physicians across the Institute. As no current standard exists for encoding clinical trial eligibility criteria, a new language called Clinical Trial Markup Language (CTML) was developed, and over 178 genomically-driven clinical trials were encoded using this language. The platform is open source and freely available for adoption by other institutions.\n\nConclusionMatchMiner is the first open platform developed to enable computational matching of patient-specific genomic profiles to precision cancer medicine clinical trials. Creating MatchMiner required developing clinical trial eligibility standards to support genome-driven matching and developing intuitive interfaces to support practical use-cases. Given the complexity of tumor profiling and the rapidly changing multi-site nature of genome-driven clinical trials, open source software is the most efficient, scalable, and economical option for matching cancer patients to clinical trials.

bioinformatics

Delta integrates 3D physical structure with topology and genomic data of chromosomes

MotivationThe regulation of gene transcription and DNA replication are tightly associated with the 3D chromosomal structures and genomic features, e.g. epigenetic marks, transcription factor bindings and non-coding RNAs. The interaction between the features and the chromosomal structures forming a multilayer 3D regulatory network. Therefore, it is necessary to integrate the physical 3D architecture of genome and features to comprehensive depict their connection to gene regulation.\n\nResultsHere, we present an integrative visualization and analysis platform, Delta, to facilitate visually annotating and exploring the 3D physical architecture of genomes. Delta takes Hi-C or ChIA-PET contact matrix as input and predicts the topology associated domains and chromatin loops in the genome, and generates a physical 3D model which represents the plausible consensus 3D structure of the genome. Delta features a highly interactive visualization tool, which enhanced the integration of genome topology/physical structure and extensive genome annotation, by juxtaposition of the 3D model with diverse genomic assay outputs. Finally, we showcased that Delta could be helpful to reveal potentially interesting findings by a case study on the {beta}-globin gene region.\n\nAvailability and implementationhttp://delta.big.ac.cn/.\n\nContacttangbx@big.ac.cn.\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Stable isotope informed genome-resolved metagenomics reveals that Saccharibacteria utilize microbially processed plant derived carbon

BackgroundThe transformation of plant photosynthate into soil organic carbon and its recycling to CO2 by soil microorganisms is one of the central components of the terrestrial carbon cycle. There are currently large knowledge gaps related to which soil-associated microorganisms take up plant carbon in the rhizosphere and the fate of that carbon.\n\nResultsWe conducted an experiment in which common wild oats (Avena fatua) were grown in a 13CO2 atmosphere and the rhizosphere and non-rhizosphere soil was sampled for genomic analyses. Density gradient centrifugation of DNA extracted from soil samples enabled distinction of microbes that did and did not incorporate the 13C into their DNA. A 1.45 Mbp genome of a Saccharibacteria (TM7) was identified and, despite the microbial complexity of rhizosphere soil, curated to completion. The genome lacks many biosynthetic pathways, including genes required to synthesize DNA de novo. Rather, it requires externally-derived nucleotides for DNA and RNA synthesis. Given this, we conclude that rhizosphere-associated Saccharibacteria recycle DNA from bacteria that live off plant exudates and/or phage that acquired 13C because they preyed upon these bacteria and/or directly from the labelled plant DNA. Isotopic labeling indicates that the population was replicating during the six-week period of plant growth. Interestingly, the genome is ~30% larger than other complete Saccharibacteria genomes from non-soil environments, largely due to more genes for complex carbon utilization and amino acid metabolism. Given the ability to degrade cellulose, hemicellulose, pectin, starch and 1,3-{beta}-glucan, we predict that this Saccharibacteria generates energy by fermentation of soil necromass and plant root exudates to acetate or lactate. The genome encodes a linear electron transport chain featuring a terminal oxidase, suggesting that this Saccharibacteria may respire aerobically. The genome encodes a hydrolase that could breakdown salicylic acid, a plant defense signaling molecule, and genes to make a variety of isoprenoids, including the plant hormone zeatin.\n\nConclusionsRhizosphere Saccharibacteria likely depend on other bacteria for basic cellular building blocks. We propose that isotopically labeled CO2 is incorporated into plant-derived carbon and then into the DNA of rhizosphere organisms capable of nucleotide synthesis, and the nucleotides are recycled into Saccharibacterial genomes.

microbiology