Search bioRxivSearch

Biology subjects

Han, M.

Publications and source records attributed to Han, M..

8 recordsLinked to original sources

Using QC-Blind for quality control and contamination screening of bacteria DNA sequencing data without reference genome

Quality control in next generation sequencing has become increasingly important as the technique becomes widely used. Tools have been developed for filtering possible contaminants in the sequencing data of species with known reference genome. Unfortunately, reference genomes for all the species involved, including the contaminants, are required for these tools to work. This precludes many real-life samples that have no information about the complete genome of the target species, and are contaminated with unknown microbial species.\n\nIn this work we propose QC-Blind, a novel quality control pipeline for removing contaminants without any use of reference genomes. The pipeline requires only very little information from the marker genes of the target species. The entire pipeline consists of unsupervised read assembly, contig binning, read clustering and marker gene assignment.\n\nWhen evaluated on in silico, ab initio and in vivo datasets, QC-Blind proved effective in removing unknown contaminants with high specificity and accuracy, while preserving most of the genomic information of the target bacterial species. Therefore, QC-Blind could serve well in situations where limited information is available for both target and contamination species.\n\nIMPORTANCEAt present, many sequencing projects are still performed on potentially contaminated samples, which bring into question their accuracies. However, current reference-based quality control method are limited as they need either the genome of target species or contaminations. In this work we propose QC-Blind, a novel quality control pipeline for removing contaminants without any use of reference genomes. When evaluated on in silico, ab initio and in vivo datasets, QC-Blind proved effective in removing unknown contaminants with high specificity and accuracy, while preserving most of the genomic information of the target bacterial species. Therefore, QC-Blind is suitable for real-life samples where limited information is available for both target and contamination species.

bioinformatics

Genome-wide identification and expression specificity analysis of the DNA methyltransferase gene family under adversity stresses in cotton

DNA methylation is an important epigenetic mode of genomic DNA modification that is an important part of maintaining epigenetic content and regulating gene expression. DNA methyltransferases (MTases) are the key enzymes in the process of DNA methylation. Thus far, there has been no systematic analysis the DNA MTases found in cotton. In this study, the whole genome of cotton C5-Mtase coding genes was identified and analyzed using a bioinformatics method based on information from the cotton genome. In this study, 51 DNA MTase genes were identified, of which 8 belonged to G. raimondii (group D), 9 belonged to G. arboretum L. (group A), 16 belonged to G. hirsutum L. (group AD1) and 18 belonged to G. barbadebse L. (group AD2). Systematic evolutionary analysis divided the 51 genes into four subfamilies, including 7 MET homologous proteins, 25 CMT homologous proteins, 14 DRM homologous proteins and 5 DNMT2 homologous proteins. Further studies showed that the DNA MTases in cotton were more phylogenetically conserved. The comparison of their protein domains showed that the C-terminal functional domain of the 51 proteins had six conserved motifs involved in methylation modification, indicating that the protein has a basic catalytic methylation function and the difference in the N-terminal regulatory domains of the 51 proteins divided the proteins into four classes, MET, CMT, DRM and DNMT2, in which DNMT2 lacks an N-terminal regulatory domain. Gene expression in cotton is not the same under different stress treatments. Different expression patterns of DNA MTases show the functional diversity of the cotton DNA methyltransferase gene family. VIGS silenced Gossypium hirsutum l. in the cotton seedling of DNMT2 family gene GhDMT6, after stress treatment the growth condition was better than the control. The distribution of DNA MTases varies among cotton species. Different DNA MTase family members have different genetic structures, and the expression level changes with different stresses, showing tissue specificity. Under salt and drought stress, G. hirsutum L. TM-1 increased the number of genes more than G. raimondii and G. arboreum L. Shixiya 1. The resistance of Gossypium hirsutum L.TM-1 to cold, drought and salt stress was increased after the plants were silenced with GhDMT6 gene.

genomics

A simple and effective method to isolate germ nuclei from C. elegans for genomic assays.

The wide variety of specialized permissive and repressive mechanisms by which germ cells regulate developmental gene expression are not well understood genome-wide. Isolation of germ cells with high integrity and purity from living animals is necessary to address these open questions, but no straightforward methods are currently available. Here we present an experimental paradigm that permits the isolation of the nuclei from C. elegans germ cells at quantities sufficient for genomic analyses. We demonstrate that these nuclei represent a very pure population and are suitable for both chromatin immunoprecipitation (ChIP-seq) and transcriptome (RNA-seq) analyses. The method does not require specialized transgenic strains or growth conditions and can be readily adopted by other researchers with minimal troubleshooting. This new capacity removes a major barrier in the field to dissect gene expression mechanisms in the germ line of C. elegans. Consequent discoveries using this technology will be relevant to conserved regulatory mechanisms across species.

genomics

PANDA: A comprehensive and flexible tool for proteomics data quantitative analysis

SummaryAs the experiment techniques and strategies in quantitative proteomics are improving rapidly, the corresponding algorithms and tools for protein quantification with high accuracy and precision are continuously required to be proposed. Here, we present a comprehensive and flexible tool named PANDA for proteomics data quantification. PANDA, which supports both label-free and labeled quantifications, is compatible with existing peptide identification tools and pipelines with considerable flexibility. Compared with MaxQuant on two complex da-tasets, PANDA was proved to be more accurate and precise with less computation time. Additionally, PANDA is an easy-to-use desktop ap-plication tool with user-friendly interfaces.\n\nAvailabilityPANDA is freely available for download at https://sourceforge.net/projects/panda-tools/.\n\nContact1987ccpacer@163.com and zhuyunping@gmail.com

bioinformatics

Holmes-ITS2: Consolidated ITS2 resources and search engines for plant DNA-based marker analyses

Plants are valuable resources for a variety of products in modern societies. Plant species identification is an integral part of research and practical application on plants. In parallel with high-throughput sequencing technology, the high-throughput screening of species is in high demand. Highly accurate and efficient DNA-based marker identification is essential for the effective analysis of plant species or biological constituents of a mixture of plants as well. Therefore, it is of general interests and significance to generate a comprehensive and accurate DNA-based marker sequence resource, as well as to build efficient sequence search engines, for the accurate and fast identification of plant species.\n\nIn this work, we have firstly established a high-quality ITS2 sequence database of plant species containing more than 150,000 entries, through the systematical collection and manually collation of the published ITS2 sequencing data of plant species, data quality control, as well as representative sequence refinement based on clustering method. Secondly, an accurate and efficient plant species identification system based on ITS2 sequence was constructed, which is the proper combination of sequence search algorithms including BLAST and Kraken. Through the deployment of high-performance and frequently updated web service, its expected to serve for a wide range of researchers involving the taxonomy classification of plant species, as well as for deciphering of plant mixed systems including herbal materials in TCM preparations.\n\nThe Holmes-ITS2 web service is freely accessible at: http://its2.tcm.microbioinformatics.org/. The input of this web service could be multiple sequences in a single fasta format, to search for matching ITS2 biomarker sequences already annotated in the database. This sequence-based search is based on two engines: BLAST, and k-mer based Kraken. Alternatively, users can directly search for species name for the corresponding ITS2 biomarker sequences. The web service has been put to the test by more than 50 experts from China, Denmark and US, and the average running time for the search ranges from 3-30 seconds for up to 100 sequences as a batch query.

bioinformatics

Data-mining of Antibiotic Resistance Genes Provides Insight into the Community Structure of Ocean Microbiome

BackgroundAntibiotics have been spread widely in environments, asserting profound effects on environmental microbes as well as antibiotic resistance genes (ARGs) within these microbes. Therefore, investigating the associations between ARGs and bacterial communities become an important issue for environment protection. Ocean microbiomes are potentially large ARG reservoirs, but the marine ARG distribution and its associations with bacterial communities remain unclear.\n\nMethodswe have utilized the big-data mining techniques on ocean microbiome data to analysis the marine ARGs and bacterial distribution on a global scale, and applied comprehensive statistical analysis to unveil the associations between ARG contents, ocean microbial community structures, and environmental factors by reanalyzing 132 metagenomic samples from the Tara Oceans project.\n\nResultsWe identified in total 1,926 unique ARGs and found that: firstly, ARGs are more abundant and diverse in the mesopelagic zone than other water layers. Additionally, ARG-enriched genera are closely connected in co-occurrence network. We also found that ARG-enriched genera are often more abundant than their ARG-less neighbors. Furthermore, we found that samples from the Mediterranean that is surrounded by human activities often contain more ARGs.\n\nConclusionOur research for investigating the marine ARG distribution and revealing the association between ARG and bacterial communities provide a deeper insight into the marine bacterial communities. We found that ARG-enriched genera were often more abundant than their ARG-less neighbors in the same environment, indicating that genera enriched with ARGs might possess an advantage over others in the competition for survival in the oceanic microbial communities.

microbiology

Agricultural Pollution Risks Influence Microbial Ecology in Honghu Lake

BackgroundAgricultural activities, such as stock-farming, planting industry, and fish aquaculture, can influence the physicochemistry and biology of freshwater lakes. However, the extent to which these agricultural activities, especially those that result in eutrophication and antibiotic pollution, effect water and sediment-associated microbial ecology, remains unclear.\n\nMethodsWe performed a geospatial analysis of water and sediment associated microbial community structure, as well as physicochemical parameters and antibiotic pollution, across 18 sites in Honghu lake, which range from impacted to less-impacted by agricultural pollution. Furthermore, the co-occurrence network of water and sediment were built and compared accorded to the agricultural activities.\n\nResultsPhysicochemical properties including TN, TP, NO3--N, and NO2--N were correlated with microbial compositional differences in water samples. Likewise, in sediment samples, Sed-OM and Sed-TN correlated with microbial diversity. Oxytetracycline and tetracycline concentration described the majority of the variance in taxonomic and predicted functional diversity between impacted and less-impacted sites in water and sediment samples, respectively. Finally, the structure of microbial co-associations was influenced by the eutrophication and antibiotic pollution.\n\nConclusionThese analyses of the composition and structure of water and sediment microbial communities in anthropologically-impacted lakes are imperative for effective environmental pollution monitoring. Likewise, the exploration of the associations between environmental variables (e.g. physicochemical properties, and antibiotics) and community structure is important in the assessment of lake water quality and its ability to sustain agriculture. These results show agricultural practices can negatively influence not only the physicochemical properties, but also the biodiversity of microbial communities associated with the Honghu lake ecosystem. And these results provide compelling evidence that the microbial community can be used as a sentinel of eutrophication and antibiotics pollution risk associated with agricultural activity; and that proper monitoring of this environment is vital to maintain a sustainable environment in Honghu lake.

microbiology

Harnessing the Cross-talk between Tumor Cells and Tumor-associated Macrophages with a Nano-drug for modulation of Glioblastoma Immune Microenvironment

Glioblastoma (GBM) is the most frequent and malignant brain tumor with a high mortality rate. The presence of a large population of macrophages (M{varphi}) in the tumor microenvironment is a prominent feature of GBM and these so-called tumor-associated M{varphi} (TAM) closely interact with the GBM cells to promote the survival, progression and therapy resistance of the GBM. Various therapeutic strategies have been devised either targeting the GBM cells or the TAM but few have addressed the cross-talks between the two cell populations. The present study was carried out to explore the possibility of exploiting the cross-talks between the GBM cells (GC) and TAM for modulation of the GBM microenvironment through using Nano-DOX, a drug composite based on nanodiamonds bearing doxorubicin. In the in vitro work on human cell models, Nano-DOX-loaded TAM were first shown to be viable and able to infiltrate three-dimensional GC spheroids and release cargo drug therein. GC were then demonstrated to encourage Nano-DOX-loaded TAM to unload Nano-DOX back into GC which consequently emitted damage-associated molecular patterns (DAMPs) that are powerful immunostimulatory agents as well as indicators of cell damage. Nano-DOX was next proven to be a more potent inducer of GC DAMPs emission than doxorubicin. As a result, Nano-DOX-damaged GC exhibited an enhanced ability to attract both TAM and Nano-DOX-loaded TAM. Most remarkably, Nano-DOX-damaged GC reprogrammed the TAM from a pro-GBM phenotype to an anti-GBM phenotype that suppressed GC growth. Finally, the in vivo relevance of the in vitro findings was tested in animal study. Mice bearing orthotopic human GBM xenografts were intravenously injected with Nano-DOX-loaded mouse TAM which were found releasing drug in the GBM xenografts 24 h after injection. GC damage was evidenced by the induction of DAMPs emission within the xenografts and a shift of TAM phenotype was detected as well. Taken together, our results demonstrate a novel way with therapeutic potential to harness the cross-talk between GBM cells and TAM for modulation of the tumor immune microenvironment.\n\nAbbreviationsATP, adenosine triphosphate; BBB, blood-brain barrier; BCA, bicinchoninic acid; BMDM, bone marrow derived macrophages; CD, cluster of differentiation; CFSE, 5(6)-carboxyfluorescein diacetate, succinimidyl ester; CM, conditioned culture medium; CNS, central nervous system; CRT, calreticulin; DAMPs, damage-associated molecular patterns; DAB, diaminobenzidine; DOX, doxorubicin; ECL, enhanced chemiluminescence; ELISA, enzyme-linked immunosorbent assay; HMGB1, high mobility group protein B1; HSP90, heat shock protein 90; FACS, flow cytometry; GBM, glioblastoma; Guanylate Binding Protein 5 (GBP5); GC, glioblastoma cells; IHC, immunohistochemical; IL, interleukin; M{varphi}, macrophages; mBMDM, mouse BMDM; mBMDM2, Type-2 mBMDM; M1, Type-1 Mo; M2, Type-2 Mo; Nano-DOX, ND-PG-RGD-DOX; ND, nanodiamonds; Nano-DOX-mBMDM, Nano-DOX-loaded mouse BMDM; NGCM, Nano-DOX-treated-GC-conditioned medium; PBS, phosphate buffered saline; PG, polyglycerol; PMA, phorbol 12-myristate 13-acetate; PVDF, polyvinylidene fluoride; RGD, tripeptide of L-arginine, glycine and L-aspartic acid; RM, regular culture medium; SD, standard deviation; TAM, tumor-associated M{varphi}; TBST, Tris Buffered Saline with Tween(R) 20.\n\nGraphic abstract\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=153 SRC=\"FIGDIR/small/170282_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (111K):\norg.highwire.dtl.DTLVardef@180aea2org.highwire.dtl.DTLVardef@14922f7org.highwire.dtl.DTLVardef@96a696org.highwire.dtl.DTLVardef@92f050_HPS_FORMAT_FIGEXP M_FIG C_FIG

pharmacology and toxicology