Search bioRxivSearch

Biology subjects

Timp, W.

Publications and source records attributed to Timp, W..

4 recordsLinked to original sources

Assessing 16S marker gene survey data analysis methods using mixtures of human stool sample DNA extracts.

BackgroundAnalysis of 16S rRNA marker-gene surveys, used to characterize prokaryotic microbial communities, may be performed by numerous bioinformatic pipelines and downstream analysis methods. However, there is limited guidance on how to decide between methods, appropriate data sets and statistics for assessing these methods are needed. We developed a mixture dataset with real data complexity and an expected value for assessing 16S rRNA bioinformatic pipelines and downstream analysis methods. We generate an assessment dataset using a two-sample titration mixture design. The sequencing data were processed using multiple bioinformatic pipelines, i) DADA2 a sequence inference method, ii) Mothur a de novo clustering method, and iii) QIIME with open-reference clustering. The mixture dataset was used to qualitatively and quantitatively assess count tables generated using the pipelines.\n\nResultsThe qualitative assessment was used to evalute features only present in unmixed samples and titrations. The abundance of Mothur and QIIME features specific to unmixed samples and titrations were explained by sampling alone. However, for DADA2 over a third of the unmixed sample and titration specific feature abundance could not be explained by sampling alone. The quantitative assessment evaluated pipeline performance by comparing observed to expected relative and differential abundance values. Overall the observed relative abundance and differential abundance values were consistent with the expected values. Though outlier features were observed across all pipelines.\n\nConclusionsUsing a novel mixture dataset and assessment methods we quantitatively and qualitatively evaluated count tables generated using three bioinformatic pipelines. The dataset and methods developed for this study will serve as a valuable community resource for assessing 16S rRNA marker-gene survey bioinformatic methods.

bioinformatics

Epigenetic changes induced by Bacteroides fragilis toxin

Enterotoxigenic Bacteroides fragilis (ETBF) is a gram negative, obligate anaerobe member of the gut microbial community in up to 40% of healthy individuals. This bacterium is found more frequently in people with colorectal cancer (CRC) and causes tumor formation in the distal colon of mice heterozygous for the adenomatous polyposis coli gene (Apc+/-); tumor formation is dependent on ETBF-secreted Bacteroides fragilis toxin (BFT). Though some of the immediate downstream effects of BFT on colon epithelial cells (CECs) are known, we still do not understand how this potent exotoxin causes changes in CECs that lead to tumor formation and growth. Because of the extensive data connecting alterations in the epigenome with tumor formation, initial experiments attempting to connect BFT-induced tumor formation with methylation in CECs have been performed, but the effect of BFT on other epigenetic processes, such as chromatin structure, remains unexplored. Here, the changes in chromatin accessibility (ATAC-seq) and gene expression (RNA-seq) induced by treatment of HT29/C1 cells with BFT for 24 and 48 hours is examined. Our data show that several genes are differentially expressed after BFT treatment and these changes correlate with changes in chromatin accessibility. Also, sites of increased chromatin accessibility are associated with a lower frequency of common single nucleotide variants (SNVs) in CRC and with a higher frequency of common differentially methylated regions (DMRs) in CRC. These data provide insight into the mechanisms by which BFT induces tumor formation. Further understanding of how BFT impacts nuclear structure and function in vivo is needed.\n\nImportanceColorectal cancer (CRC) is a major public health concern; there were approximately 135,430 new cases in 2017, and CRC is the second leading cause of cancer-related deaths for both men and women in the US (1). Many factors have been linked to CRC development, the most recent of which is the gut microbiome. Pre-clinical models support that enterotoxigenic Bacteroides fragilis (ETBF), among other bacteria, induce colon carcinogenesis. However, it remains unclear if the virulence determinants of any pro-carcinogenic colon bacterium induce DNA mutations or changes that initiate clonal CEC expansion. Using a reductionist model, we demonstrate that BFT rapidly alters chromatin structure and function consistent with capacity to contribute to CRC pathogenesis.

genomics

First Draft Genome Sequence of the Pathogenic Fungus Lomentospora prolificans (formerly Scedosporium prolificans)

Here we describe the sequencing and assembly of the pathogenic fungus Lomentospora prolificans using a combination of short, highly accurate Illumina reads and additional coverage in very long Oxford Nanopore reads. The resulting assembly is highly contiguous, containing a total of 37,630,066 bp with over 98% of the sequence in just 26 scaffolds. Annotation identified 8,656 protein-coding genes. Pulsed-field gel analysis suggests that this organism contains at least 7 and possibly 11 chromosomes, the two longest of which have sizes corresponding closely to the sizes of the longest scaffolds, at 6.6 and 5.7 Mb.

microbiology

Single molecule, full-length transcript sequencing provides insight into the extreme metabolism of ruby-throated hummingbird Archilochus colubris

Hummingbirds can support their high metabolic rates exclusively by oxidizing ingested sugars, which is unsurprising given their sugar-rich nectar diet and use of energetically expensive hovering flight. However, they cannot rely on dietary sugars as a fuel during fasting periods, such as during the night, at first light, or when undertaking long-distance migratory flights, and must instead rely exclusively on onboard lipids. This metabolic flexibility is remarkable both in that the birds can switch between exclusive use of each fuel type within minutes and in that de novo lipogenesis from dietary sugar precursors is the principle way in which fat stores are built, sometimes at exceptionally high rates, such as during the few days prior to a migratory flight. The hummingbird hepatopancreas is the principle location of de novo lipogenesis and likely plays a key role in fuel selection, fuel switching, and glucose homeostasis. Yet understanding how this tissue, and the whole organism, achieves and moderates high rates of energy turnover is hampered by a fundamental lack of information regarding how genes coding for relevant enzymes differ in their sequence, expression, and regulation in these unique animals. To address this knowledge gap, we generated a de novo transcriptome of the hummingbird liver using PacBio full-length cDNA sequencing (Iso-Seq), yielding a total of 8.6Gb of sequencing data, or 2.6M reads from 4 different size fractions. We analyzed data using the SMRTAnalysis v3.1 Iso-Seq pipeline, including classification of reads and clustering of isoforms (ICE) followed by error-correction (Arrow). With COGENT, we clustered different isoforms into gene families to generate de novo gene contigs. We performed orthology analysis to identify closely related sequences between our transcriptome and other avian and human gene sets. We also aligned our transcriptome against the Calypte anna genome where possible. Finally, we closely examined homology of critical lipid metabolic genes between our transcriptome data and avian and human genomes. We confirmed high levels of sequence divergence within hummingbird lipogenic enzymes, suggesting a high probability of adaptive divergent function in the hepatic lipogenic pathways. Our results have leveraged cutting-edge technology and a novel bioinformatics pipeline to provide a compelling first direct look at the transcriptome of this incredible organism.

genomics