Search bioRxivSearch

Biology subjects

Ning, K.

Publications and source records attributed to Ning, K..

7 recordsLinked to original sources

Microbial contamination screening and interpretation for biological laboratory environments

Advances in microbiome researches have led us to the realization that the composition of microbial communities of indoor environment is profoundly affected by the function of buildings, and in turn may bring detrimental effects to the indoor environment and the occupants. Thus investigation is warranted for a deeper understanding of the potential impact of the indoor microbial communities. Among these environments, the biological laboratories stand out because they are relatively clean and yet are highly susceptible to microbial contaminants. In this study, we assessed the microbial compositions of samples from the surfaces of various sites across different types of biological laboratories. We have qualitatively and quantitatively assessed these possible microbial contaminants, and found distinct differences in their microbial community composition. We also found that the type of laboratories has a larger influence than the sampling site in shaping the microbial community, in terms of both structure and richness. On the other hand, the public areas of the different types of laboratories share very similar sets of microbes. Tracing the main sources of these microbes, we identified both environmental and human factors that are important factors in shaping the diversity and dynamics of these possible microbial contaminations in biological laboratories. These possible microbial contaminants that we have identified will be helpful for people who aim to eliminate them from samples.\n\nImportanceMicrobial communities from biological laboratories might hamper the conduction of molecular biology experiments, yet these possible contaminations are not yet carefully investigated. In this work, a metagenomic approach has been applied to identify the possible microbial contaminants and their sources, from the surfaces of various sites across different types of biological laboratories. We have found distinct differences in their microbial community compositions. We have also identified the main sources of these microbes, as well as important factors in shaping the diversity and dynamics of these possible microbial contaminations. The identification and interpretation of these possible microbial contaminants in biological laboratories would be helpful for alleviate their potential detrimental effects.

microbiology

Using QC-Blind for quality control and contamination screening of bacteria DNA sequencing data without reference genome

Quality control in next generation sequencing has become increasingly important as the technique becomes widely used. Tools have been developed for filtering possible contaminants in the sequencing data of species with known reference genome. Unfortunately, reference genomes for all the species involved, including the contaminants, are required for these tools to work. This precludes many real-life samples that have no information about the complete genome of the target species, and are contaminated with unknown microbial species.\n\nIn this work we propose QC-Blind, a novel quality control pipeline for removing contaminants without any use of reference genomes. The pipeline requires only very little information from the marker genes of the target species. The entire pipeline consists of unsupervised read assembly, contig binning, read clustering and marker gene assignment.\n\nWhen evaluated on in silico, ab initio and in vivo datasets, QC-Blind proved effective in removing unknown contaminants with high specificity and accuracy, while preserving most of the genomic information of the target bacterial species. Therefore, QC-Blind could serve well in situations where limited information is available for both target and contamination species.\n\nIMPORTANCEAt present, many sequencing projects are still performed on potentially contaminated samples, which bring into question their accuracies. However, current reference-based quality control method are limited as they need either the genome of target species or contaminations. In this work we propose QC-Blind, a novel quality control pipeline for removing contaminants without any use of reference genomes. When evaluated on in silico, ab initio and in vivo datasets, QC-Blind proved effective in removing unknown contaminants with high specificity and accuracy, while preserving most of the genomic information of the target bacterial species. Therefore, QC-Blind is suitable for real-life samples where limited information is available for both target and contamination species.

bioinformatics

Gut microbiome of an unindustrialized population have characteristic enrichment of SNPs in species and functions with the succession of seasons

Most studies investigating human gut microbiome dynamics are conducted in modern populations. However, unindustrialized populations are arguably better subjects in answering human-gut microbiome coevolution questions due to their lower exposure to antibiotics and higher dependence on natural resources. Hadza hunter-gatherers in Tanzania have been found to exhibit high biodiversity and seasonal patterns in their gut microbiome composition at family level, where some taxa disappear in one season and reappear at later time. However, such seasonal changes have previously been profiled only according to species abundances, with genome-level variant dynamics unexplored. As a result it is still elusive how microbial communities change at the genome-level under environmental pressures caused by seasonal changes. Here, a strain-level SNP analysis of Hadza gut metagenome is performed for 40 Hadza fecal samples collected in three seasons. First, we benchmarked three SNP calling tools based on simulated sequencing reads, and selected VarScan2 that has highest accuracy and sensitivity after a filtering step. Second, we applied VarScan2 on Hadza gut microbiome, with results showing that: with more SNP presented in wet season in general, eight prevalent species have significant SNP enrichments in wet season of which only three species have relatively high abundances. This indicates that SNP characteristics are independent of species abundances, and provides us a unique lens towards microbial community dynamics. Finally, we identify 83 genes with the most characteristic SNP distributions between wet season and dry season. Many of these genes are from Ruminococcus obeum, and mainly from metabolic pathways like carbon metabolism, pyruvate metabolism and glycolysis, as shown by KEGG annotation. This implies that the seasonal changes might indirectly impact the mutational patterns for specific species and functions for gut microbiome of an unindustrialized population, indicating the role of these variants in their adaptation to the changing environment and diets.\n\nImportanceBy analyzing the changes of SNP enrichments in different seasons, we have found that SNP characteristics are independent of species abundances, and could provide us a unique lens towards microbial community dynamics at the genomic level. Many of the genes in microbiome also presented characteristic SNP distributions between wet season and dry season, indicating the role of variants in specific species in their adaptation to the changing environment for an unindustrialized population.

microbiology

Bi-clustering interpretation and prediction of correlation between gene expression and protein abundance

Most organisms transcript and protein level only moderately correlate for various reasons, such as regulation of transcription and protein degradation. Better prediction and understanding the correlation between gene expression and protein abundance has been possible by harnessing the matching RNA/protein datasets produced by modern high-throughput RNA-Seq and mass spectrometry methods. In this work, we have utilized some well-studied matching RNA/protein datasets, and explored for the first time a bi-clustering method to cluster genes that have consistent correlation patterns between gene expression and protein abundance. The clustering results have been interpreted from the perspective of both transcriptomic and proteomic features, which show that mRNA half-life, protein half-life and protein structure in concert significantly affect the correlation of gene expression and protein abundance. With these and other carefully selected features, a prediction model based on individual clusters, called Cluster-based Linear prediction Model (CLM), was built and tested on mouse liver mitochondrial, mouse brainstem mitochondrial, Saccharomyces cerevisiae and Danio rerio datasets. CLM could find genes for which protein abundance can be predicted from mRNA data. In summary, based on bi-clustering, feature selection and CLM model, we have established a new and valuable cluster-based protein abundance prediction method.

bioinformatics

Holmes-ITS2: Consolidated ITS2 resources and search engines for plant DNA-based marker analyses

Plants are valuable resources for a variety of products in modern societies. Plant species identification is an integral part of research and practical application on plants. In parallel with high-throughput sequencing technology, the high-throughput screening of species is in high demand. Highly accurate and efficient DNA-based marker identification is essential for the effective analysis of plant species or biological constituents of a mixture of plants as well. Therefore, it is of general interests and significance to generate a comprehensive and accurate DNA-based marker sequence resource, as well as to build efficient sequence search engines, for the accurate and fast identification of plant species.\n\nIn this work, we have firstly established a high-quality ITS2 sequence database of plant species containing more than 150,000 entries, through the systematical collection and manually collation of the published ITS2 sequencing data of plant species, data quality control, as well as representative sequence refinement based on clustering method. Secondly, an accurate and efficient plant species identification system based on ITS2 sequence was constructed, which is the proper combination of sequence search algorithms including BLAST and Kraken. Through the deployment of high-performance and frequently updated web service, its expected to serve for a wide range of researchers involving the taxonomy classification of plant species, as well as for deciphering of plant mixed systems including herbal materials in TCM preparations.\n\nThe Holmes-ITS2 web service is freely accessible at: http://its2.tcm.microbioinformatics.org/. The input of this web service could be multiple sequences in a single fasta format, to search for matching ITS2 biomarker sequences already annotated in the database. This sequence-based search is based on two engines: BLAST, and k-mer based Kraken. Alternatively, users can directly search for species name for the corresponding ITS2 biomarker sequences. The web service has been put to the test by more than 50 experts from China, Denmark and US, and the average running time for the search ranges from 3-30 seconds for up to 100 sequences as a batch query.

bioinformatics

Data-mining of Antibiotic Resistance Genes Provides Insight into the Community Structure of Ocean Microbiome

BackgroundAntibiotics have been spread widely in environments, asserting profound effects on environmental microbes as well as antibiotic resistance genes (ARGs) within these microbes. Therefore, investigating the associations between ARGs and bacterial communities become an important issue for environment protection. Ocean microbiomes are potentially large ARG reservoirs, but the marine ARG distribution and its associations with bacterial communities remain unclear.\n\nMethodswe have utilized the big-data mining techniques on ocean microbiome data to analysis the marine ARGs and bacterial distribution on a global scale, and applied comprehensive statistical analysis to unveil the associations between ARG contents, ocean microbial community structures, and environmental factors by reanalyzing 132 metagenomic samples from the Tara Oceans project.\n\nResultsWe identified in total 1,926 unique ARGs and found that: firstly, ARGs are more abundant and diverse in the mesopelagic zone than other water layers. Additionally, ARG-enriched genera are closely connected in co-occurrence network. We also found that ARG-enriched genera are often more abundant than their ARG-less neighbors. Furthermore, we found that samples from the Mediterranean that is surrounded by human activities often contain more ARGs.\n\nConclusionOur research for investigating the marine ARG distribution and revealing the association between ARG and bacterial communities provide a deeper insight into the marine bacterial communities. We found that ARG-enriched genera were often more abundant than their ARG-less neighbors in the same environment, indicating that genera enriched with ARGs might possess an advantage over others in the competition for survival in the oceanic microbial communities.

microbiology

Agricultural Pollution Risks Influence Microbial Ecology in Honghu Lake

BackgroundAgricultural activities, such as stock-farming, planting industry, and fish aquaculture, can influence the physicochemistry and biology of freshwater lakes. However, the extent to which these agricultural activities, especially those that result in eutrophication and antibiotic pollution, effect water and sediment-associated microbial ecology, remains unclear.\n\nMethodsWe performed a geospatial analysis of water and sediment associated microbial community structure, as well as physicochemical parameters and antibiotic pollution, across 18 sites in Honghu lake, which range from impacted to less-impacted by agricultural pollution. Furthermore, the co-occurrence network of water and sediment were built and compared accorded to the agricultural activities.\n\nResultsPhysicochemical properties including TN, TP, NO3--N, and NO2--N were correlated with microbial compositional differences in water samples. Likewise, in sediment samples, Sed-OM and Sed-TN correlated with microbial diversity. Oxytetracycline and tetracycline concentration described the majority of the variance in taxonomic and predicted functional diversity between impacted and less-impacted sites in water and sediment samples, respectively. Finally, the structure of microbial co-associations was influenced by the eutrophication and antibiotic pollution.\n\nConclusionThese analyses of the composition and structure of water and sediment microbial communities in anthropologically-impacted lakes are imperative for effective environmental pollution monitoring. Likewise, the exploration of the associations between environmental variables (e.g. physicochemical properties, and antibiotics) and community structure is important in the assessment of lake water quality and its ability to sustain agriculture. These results show agricultural practices can negatively influence not only the physicochemical properties, but also the biodiversity of microbial communities associated with the Honghu lake ecosystem. And these results provide compelling evidence that the microbial community can be used as a sentinel of eutrophication and antibiotics pollution risk associated with agricultural activity; and that proper monitoring of this environment is vital to maintain a sustainable environment in Honghu lake.

microbiology