Search bioRxiv⌕ Search

Biology subjects

Li, P.-E.

Publications and source records attributed to Li, P.-E..

3 recordsLinked to original sources

Global untreated wastewater hosts a vast reservoir of previously uncharacterized microbial lineages

Untreated wastewater are reservoirs of microbial communities originating from different sources within an urbanized human-populated area and is now regularly used for tracking and assessing levels of certain human pathogens like SARS-CoV-2, Salmonella, and Poliomyelitis. Wastewater microbial communities can exhibit high diversity, comprised of mostly non-pathogens alongside a smaller population of pathogens. Much of the research on wastewater has focused on either pathogens or microbiomes in the treatment plants, but not on the microbiome of untreated wastewater, which is perhaps more reflective of community health. Moreover, as wastewater usage expands to more known and unknown pathogen detection, a deeper understanding of both wastewaters overall genetic diversity and diversity over time is pivotal. Towards that goal, we characterized the observed microbial diversity of global untreated wastewater by analyzing 1,344 publicly available metagenomes, representing 6 continents, and created a global wastewater database hosted at https://doi.org/10.5281/zenodo.17794790. We elucidated the full spectrum of prokaryotic diversity by recovering and characterizing high quality bins, contigs, and reads from all the data. We provide the first comprehensive, publicly available database describing global wastewater by uncovering previously uncharacterized microbes, identifying geographically specific microbial populations, and creating a high-quality dataset to advance future microbiome, epidemiological, and biosurveillance research.

genomics↗

Towards increased accuracy and reproducibility in SARS-CoV-2 next generation sequence analysis for public health surveillance

During the COVID-19 pandemic, SARS-CoV-2 surveillance efforts integrated genome sequencing of clinical samples to identify emergent viral variants and to support rapid experimental examination of genome-informed vaccine and therapeutic designs. Given the broad range of methods applied to generate new viral genomes, it is critical that consensus and variant calling tools yield consistent results across disparate pipelines. Here we examine the impact of sequencing technologies (Illumina and Oxford Nanopore) and 7 different downstream bioinformatic protocols on SARS-CoV-2 variant calling as part of the NIH Accelerating COVID-19 Therapeutic Interventions and Vaccines (ACTIV) Tracking Resistance and Coronavirus Evolution (TRACE) initiative, a public-private partnership established to address the COVID-19 outbreak. Our results indicate that bioinformatic workflows can yield consensus genomes with different single nucleotide polymorphisms, insertions, and/or deletions even when using the same raw sequence input datasets. We introduce the use of a specific suite of parameters and protocols that greatly improves the agreement among pipelines developed by diverse organizations. Such consistency among bioinformatic pipelines is fundamental to SARS-CoV-2 and future pathogen surveillance efforts. The application of analysis standards is necessary to more accurately document phylogenomic trends and support data-driven public health responses.

bioinformatics↗

PanGIA: A Metagenomics Analytical Framework for RoutineBiosurveillance and Clinical Pathogen Detection

Metagenomics is emerging as an important tool in biosurveillance, public health, and clinical applications. However, ease-of-use for execution and data analysis remains a barrier-of-entry to the adoption of metagenomics in applied health and forensics settings. In addition, these venues often have more stringent requirements for reporting, accuracy, and precision than the traditional ecological research role of the technology. Here, we present PanGIA (Pan-Genomics for Infectious Agents), a novel bioinformatics analysis platform for hosting, processing, analyzing, and reporting shotgun metagenomics data of complex samples suspected of containing one or more pathogens. PanGIA was developed to address gaps that often preclude clinicians, medical technicians, forensics personnel, or other non-expert end-users from the routine application of metagenomics for pathogen identification. Though primarily designed to detect pathogenic microorganisms within clinical and environmental metagenomics data, PanGIA also serves as an analytical framework for microbial community profiling and comparative metagenomics. To provide statistical confidence in PanGIAs taxonomic assignments, the system provides two independent estimations of probability for species and strain level detection. First, PanGIA integrates coverage data with uniqueness information mapped across each reference genome for a stand-alone determination of confidence for each query sequence at each taxonomy level. Second, if a negative-control sample is provided, PanGIA compares this sample with a corresponding experimental unknown sample and determines a measure of confidence associated with detection above background. An integrated graphical user interface allows interactive interrogation and enables users to summarize multiple sample results by confidence score, normalized read abundance, reference genome linear coverage, depth-of-coverage, RPKM, and other metrics to detect specific organisms-of-interest. Comparison testing of the PanGIA algorithm against a number of recent k-mer, read-mapping, and marker-gene based taxonomy classifiers across various real-world datasets with spiked targets shows superior mean positive predictive value, sensitivity, and specificity. PanGIA can process a five million paired-end read dataset in under 1 hour on commodity computational hardware. The source code and documentation are publicly available at https://github.com/LANL-Bioinformatics/PanGIA or https://github.com/mriglobal/PanGIA. The database for PanGIA can be downloaded from ftp://bioinformatics.mriglobal.org/. The full GUI-based PanGIA analysis environment is available in a Docker container and can be installed from https://hub.docker.com/r/poeli/pangia/.

bioinformatics↗