Search bioRxivSearch

Biology subjects

Dunning Hotopp, J. C.

Publications and source records attributed to Dunning Hotopp, J. C..

3 recordsLinked to original sources

Cost Effective, Experimentally Robust Differential Expression Analysis for Human/Mammalian, Pathogen, and Dual-Species Transcriptomics

As sequencing read length has increased, researchers have quickly adopted longer reads for their experiments. Here, we examine host-pathogen interaction studies to assess if using longer reads is warranted. Six diverse datasets encountered in studies of host-pathogen interactions were used to assess what genomic attributes might affect the outcome of differential gene expression analysis including: gene density, operons, gene length, number of introns/exons, and intron length. Principal components analysis, hierarchical clustering with bootstrap support, and regression analyses of pairwise comparisons were undertaken on the same reads, looking at all combinations of paired and unpaired reads trimmed to 36,54,72, and 101-bp. For E coli, 36-bp single end reads performed as well as any other read length and as well as paired end reads. For all other comparisons, 54-bp and 72-bp reads were typically equivalent and different from 36-bp and 101-bp reads. Read pairing improved the outcome in several, but not all, comparisons in no discernable pattern, such that using paired reads is recommended in most scenarios. No specific genome attribute appeared to influence the data. However, experiments with an a priori expected greater biological complexity had more variable results with all read lengths relative to those with decreased complexity. When combined with cost, 54-bp paired end reads provided the most robust, internally reproducible results across all comparisons. However, using 36-bp single end reads may be desirable for bacterial samples, although possibly only if the transcriptional response is expected a priori to be robust.\n\nDATA SUMMARYO_LIThe human only CSHL Encode data set (1) was downloaded from ftp://hgdownload.cse.ucsc.edu/goldenPath/hgl9/encodeDCC/wgEncodeCshlLongRnaSeq/.\nC_LIO_LIThe data from mice vaginas infected with Candida albicans (2) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP057050).\nC_LIO_LIThe data from Aspergillus fumigatus cells in contact with human cells was downloaded from the SRA (url - https://www.ncbi.nlm.nih.gov/bioproject/399754).\nC_LIO_LIThe data from a strand-specific library from a study comparing C. albicans cells in contact with human cells with those in media (3) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP011085).\nC_LIO_LIThe data from C. albicans in culture media (3) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP011085).\nC_LIO_LIThe data from Escherichia coli grown in different media (4) was downloaded from the SRA (url - https://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP056578).\nC_LI\n\nI/We confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. {boxtimes}\n\nIMPACT STATEMENTAs sequencing technologies improve, sequencing costs decrease and read lengths increase. We examine host-pathogen interaction studies to assess if using these longer reads is warranted given their increased cost relative to using the same number of shorter reads. To this end we compared the use of various read lengths and read pairing for six diverse host-pathogen datasets with varying genomic attributes including: gene density, operons, gene length, number of introns/exons, and intron length. We find that in the bacterial sample, 36-bp single end reads performed as well as any other read length and as well as paired end reads. When combined with cost, 54-bp paired end reads provided the most robust, internally reproducible results for all other comparisons. Read pairing improved the outcome in several, but not all, comparisons in no discernable pattern, such that using paired reads is recommended in most scenarios. No specific genome attribute appeared to influence the data.

genomics

Using Core Genome Alignments to Assign Bacterial Species

With the exponential increase in the number of bacterial taxa with genome sequence data, a new standardized method is needed to assign bacterial species designations using genomic data that is consistent with the classically-obtained taxonomy. This is particularly acute for unculturable obligate intracellular bacteria like those in the Rickettsiales, where classical methods like DNA-DNA hybridization cannot be used to define species. Within the Rickettsiales, species designations have been applied inconsistently, often obfuscating the relationship between organisms and the context for experimental results. In this study, we generated core genome alignments for a wide range of genera with classically defined species, including Arcobacter, Caulobacter, Erwinia, Neisseria, Polaribacter, Ralstonia, Thermus, as well as genera within the Rickettsiales including Rickettsia, Orientia, Ehrlichia, Neoehrlichia, Anaplasma, eorickettsia, and Wolbachia. A core genome alignment sequence identity (CGASI) threshold of 96.8% was found to maximize the prediction of classically-defined species. Using the CGASI cutoff, the Wolbachia genus can be delineated into species that differ from the currently used supergroup designations, while the Rickettsia genus is delineated into nine species, as opposed to the current 27 species. Additionally, we find that core genome alignments cannot be constructed between genomes belonging to different genera, establishing a bacterial genus cutoff that suggests the need to create new genera from the Anaplasma and Neorickettsia. By using core genome alignments to assign taxonomic designations, we aim to provide a high-resolution, robust method for bacterial nomenclature that is aligned with classically-obtained results.

microbiology

Grafting or Pruning in the Animal Tree: Lateral Gene Transfer and Gene Loss?

Lateral gene transfer (LGT) into multicellular eukaryotes with differentiated tissues, particularly gonads, continues to be met with skepticism by many prominent evolutionary and genomic biologists. A detailed examination of 26 animal genomes identifed putative LGTs in invertebrate and vertebrate genomes, concluding that there are fewer predicted LGTs in vertebrates/chordates than invertebrates, but there is still evidence of LGT into chordates, including humans. More recently, a reanalysis a subset of these putative LGTs into vertebrates concluded that there is not horizontal gene transfer in the human genome. One of the genes in dispute is an N-acyl-aromatic-L-amino acid amidohydrolase (ENSG00000132744), which encodes ACY3, which was initially identified as a putative bacteria-chordate LGT but was later debunked has a significant BLAST match to a more recently deposited genome of Saccoglossus kowalevskii, a flatworm, Metazoan, and hemichordate. Using BLAST searches, HMM searches, and phylogenetics to better understand the evidence for lateral gene transfer, gene loss, and rate variation in ACY3/ASPA homologues, the most parsimonious explanation for the distribution of ACY3/ASPA genes in eukaryotes likely involves both gene loss and lateral gene transfer, albeit lateral gene transfer that occurred hundreds of millions of years ago prior to the divergence of gnathostomes and even longer and prior to the divergence of bilateria. Given the many known, well-characterized, and adaptive lateral gene transfers from bacteria to insects and nematodes, lateral gene transfers at these time scales in the ancestors of humans is expected.

genomics