Search bioRxivSearch

Biology subjects

Wang, L. C.

Publications and source records attributed to Wang, L. C..

2 recordsLinked to original sources

Dual RNA-seq provides insight into the biology of the neglected intracellular human pathogen Orientia tsutsugamushi

Emerging and neglected diseases pose challenges as their biology is frequently poorly understood, and genetic tools often do not exist to manipulate the responsible pathogen. Organism agnostic sequencing technologies offer a promising approach to understand the molecular processes underlying these diseases. Here we apply dual RNA-seq to Orientia tsutsugamushi (Ot), an obligate intracellular bacterium and the causative agent of the vector-borne human disease scrub typhus. Half the Ot genome is composed of repetitive DNA, and there is minimal collinearity in gene order between strains. Integrating RNA-seq, comparative genomics, proteomics, and machine learning, we investigated the transcriptional architecture of Ot, including operon structure and non-coding RNAs, and found evidence for wide-spread post-transcriptional antisense regulation. We compared the host response to two clinical isolates and identified distinct immune response networks that are up-regulated in response to each strain, leading to predictions of relative virulence which were confirmed in a mouse infection model. Thus, dual RNA-seq can provide insight into the biology and host-pathogen interactions of a poorly characterized and genetically intractable organism such as Ot.

microbiology

Double triage to identify poorly annotated genes in Maize: The missing link in community curation

The sophistication of gene prediction algorithms and the abundance of RNA-based evidence for the maize genome may suggest that manual curation of gene models is no longer necessary. However, quality metrics generated by the MAKER-P gene annotation pipeline identified 17,225 of 130,330 (13%) protein-coding transcripts in the B73 Reference Genome V4 gene set with models of low concordance to available biological evidence. Working with eight graduate students, we used the Apollo annotation editor to curate 86 transcript models flagged by quality metrics and a complimentary method using the Gramene gene tree visualizer. All of the triaged models had significant errors - including missing or extra exons, non-canonical splice sites, and incorrect UTRs. A correct transcript model existed for about 60% of genes (or transcripts) flagged by quality metrics; we attribute this to the convention of elevating the transcript with the longest coding sequence (CDS) to the canonical, or first, position. The remaining 40% of flagged genes resulted in novel annotations and represent a manual curation space of about 10% of the maize genome (~4,000 protein-coding genes). MAKER-P metrics have a specificity of 100%, and a sensitivity of 85%; the gene tree visualizer has a specificity of 100%. Together with the Apollo graphical editor, our double triage provides an infrastructure to support the community curation of eukaryotic genomes by scientists, students, and potentially even citizen scientists.

bioinformatics