Search bioRxivSearch

Biology subjects

Salzberg, S.

Publications and source records attributed to Salzberg, S..

6 recordsLinked to original sources

Life Inside A Dinosaur Bone: A Thriving Microbiome

Fossils were long thought to lack original organic material, but the discovery of organic molecules in fossils and sub-fossils, thousands to millions of years old, has demonstrated the potential of fossil organics to provide radical new insights into the fossil record. How long different organics can persist remains unclear, however. Non-avian dinosaur bone has been hypothesised to preserve endogenous organics including collagen, osteocytes, and blood vessels, but proteins and labile lipids are unstable during diagenesis or over long periods of time. Furthermore, bone is porous and an open system, allowing microbial and organic flux. Some of these organics within fossil bone have therefore been identified as either contamination or microbial biofilm, rather than original organics. Here, we use biological and chemical analyses of Late Cretaceous dinosaur bones and sediment matrix to show that dinosaur bone hosts a diverse microbiome. Fossils and matrix were freshly-excavated, aseptically-acquired, and then analysed using microscopy, spectroscopy, chromatography, spectrometry, DNA extraction, and 16S rRNA amplicon sequencing. The fossil organics differ from modern bone collagen chemically and structurally. A key finding is that 16S rRNA amplicon sequencing reveals that the subterranean fossil bones host a unique, living microbiome distinct from that of the surrounding sediment. Even in the subsurface, dinosaur bone is biologically active and behaves as an open system, attracting microbes that might alter original organics or complicate the identification of original organics. These results suggest caution regarding claims of dinosaur bone soft tissue preservation and illustrate a potential role for microbial communities in post-burial taphonomy.

paleontology

Thousands of large-scale RNA sequencing experiments yield a comprehensive new human gene list and reveal extensive transcriptional noise

We assembled the sequences from 9,795 RNA sequencing experiments, collected from 31 human tissues and hundreds of subjects as part of the GTEx project, to create a new, comprehensive catalog of human genes and transcripts. The new human gene database contains 43,162 genes, of which 21,306 are protein-coding and 21,856 are noncoding, and a total of 323,824 transcripts, for an average of 7.5 transcripts per gene. Our expanded gene list includes 4,998 novel genes (1,178 coding and 3,819 noncoding) and 97,511 novel splice variants of protein-coding genes as compared to the most recent human gene catalogs. We detected over 30 million additional transcripts at more than 650,000 sites, nearly all of which are likely to be nonfunctional, revealing a heretofore unappreciated amount of transcriptional noise in human cells.

genomics

HISAT-genotype: Next Generation Genomic Analysis Platform on a Personal Computer

Rapid advances in next-generation sequencing technologies have dramatically changed our ability to perform genome-scale analyses of human genomes. The human reference genome used for most genomic analyses represents only a small number of individuals, limiting its usefulness for genotyping. We designed a novel method, HISAT-genotype, for representing and searching an expanded model of the human reference genome, in which a comprehensive catalogue of known genomic variants and haplotypes is incorporated into the data structure used for searching and alignment. This strategy for representing a population of genomes, along with a very fast and memory-efficient search algorithm, enables more detailed and accurate variant analyses than previous methods. We demonstrate HISAT-genotypes accuracy for HLA typing, a critical task in human organ transplantation, and for the DNA fingerprinting tests widely used in forensics. In both applications, HISAT-genotype not only improves upon earlier computational methods, but matches or exceeds the accuracy of laboratory-based assays.\n\nOne Sentence SummaryHISAT-genotype is a software platform that has the ability to genotype all the genes in an individuals genome within a few hours on a desktop computer.

bioinformatics

Removing Contaminants from Metagenomic Databases

Metagenomic sequencing of patient samples is a very promising method for the diagnosis of human infections. Sequencing has the ability to capture all the DNA or RNA from pathogenic organisms in a human sample. However, complete and accurate characterization of the sequence, including identification of any pathogens, depends on the availability and quality of genomes for comparison. Thousands of genomes are now available, and as these numbers grow, the power of metagenomic sequencing for diagnosis should increase. However, recent studies have exposed the presence of contamination in published genomes, which when used for diagnosis increases the risk of falsely identifying the wrong pathogen.\n\nTo address this problem, we have developed a bioinformatics system for eliminating contamination as well as low-complexity genomic sequences in the draft genomes of eukaryotic pathogens. We applied this software to identify and remove human, bacterial, archaeal, and viral sequences present in a comprehensive database of all sequenced eukaryotic pathogen genomes. We also removed low-complexity genomic sequences, another source of false positives. Using this pipeline, we have produced a database of \"clean\" eukaryotic pathogen genomes for use with bioinformatics classification and analysis tools. We demonstrate that when attempting to find eukaryotic pathogens in metagenomic samples, the new database provides better sensitivity than one using the original genomes while offering a dramatic reduction in false positives.

bioinformatics

First Draft Genome Sequence of the Pathogenic Fungus Lomentospora prolificans (formerly Scedosporium prolificans)

Here we describe the sequencing and assembly of the pathogenic fungus Lomentospora prolificans using a combination of short, highly accurate Illumina reads and additional coverage in very long Oxford Nanopore reads. The resulting assembly is highly contiguous, containing a total of 37,630,066 bp with over 98% of the sequence in just 26 scaffolds. Annotation identified 8,656 protein-coding genes. Pulsed-field gel analysis suggests that this organism contains at least 7 and possibly 11 chromosomes, the two longest of which have sizes corresponding closely to the sizes of the longest scaffolds, at 6.6 and 5.7 Mb.

microbiology

16GT: a fast and sensitive variant caller using a 16-genotype probabilistic model

Summary16GT is a variant caller for Illumina WGS and WES germline data. It uses a new 16-genotype probabilistic model to unify SNP and indel calling in a single variant calling algorithm. In benchmark comparisons with five other widely used variant callers on a modern 36-core server, 16GT ran faster and demonstrated improved sensitivity in calling SNPs, and it provided comparable sensitivity and accuracy in calling indels as compared to the GATK HaplotypeCaller.\n\nAvailability and implementationhttps://github.com/aquaskyline/16GT\n\nContactrluo5@jhu.edu\n\nSupplementary informationSupplementary tables and notes are available at Bioinformatics online.

bioinformatics