Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

An interlaboratory study of complex variant detection

Next-generation sequencing (NGS) is widely used and cost-effective. Depending on the specific methods, NGS can have limitations detecting certain technically challenging variant types even though they are both prevalent in patients and medically important. These types are underrepresented in validation studies, hindering the uniform assessment of test methodologies by laboratory directors and clinicians. Specimens containing such variants can be difficult to obtain; thus, we evaluated a novel solution to this problem in which a diverse set of technically challenging variants was synthesized and introduced into a known genomic background. This specimen was sequenced by 7 laboratories using 10 different NGS workflows. The specimen was compatible with all 10 workflows and presented biochemical and bioinformatic challenges similar to those of patient specimens. Only 10 of 22 challenging variants were correctly identified by all 10 workflows, and only 3 workflows detected all 22. Many, but not all, of the sensitivity limitations were bioinformatic in nature. We conclude that Synthetic controls can provide an efficient and informative mechanism to augment studies with technically challenging variants that are difficult to obtain otherwise. Data from such specimens can facilitate inter-laboratory methodologic comparisons and can help establish standards that improve communication between clinicians and laboratories.

genetics

A de novo approach to disentangle partner identity and function in holobiont systems

BackgroundStudy of meta-transcriptomic datasets involving non-model organisms represents bioinformatic challenges. The production of chimeric sequences and our inability to distinguish the taxonomic origins of the sequences produced are inherent and recurrent difficulties in de novo assembly analyses. The study of holobiont transcriptomes shares similarities with meta-transcriptomic, and hence, is also affected by challenges invoked above. Here we propose an innovative approach to tackle such difficulties which was applied to the study of marine holobiont models as a proof of concept.\n\nResultsWe considered three holobionts models, of which two transcriptomes were previously assembled and published, and a yet unpublished transcriptome, to analyze their raw reads and assign them to the host and/or to the symbiont(s) using Short Read Connector, a k-mer based similarity method. We were able to define four distinct categories of reads for each holobiont transcriptome: host reads, symbiont reads, shared reads and unassigned reads. The result of the independent assemblies for each category within a transcriptome led to a significant diminution of de novo assembled chimeras compared to classical assembly methods. Combining independent functional and taxonomic annotations of each partners transcriptome is particularly convenient to explore the functional diversity of an holobiont. Finally, our strategy allowed to propose new functional annotations for two well-studied holobionts and a first transcriptome from a planktonic Radiolaria-Dinophyta system forming widespread symbiotic association for which our knowledge is limited. Conclusions\n\nIn contrast to classical assembly approaches, our bioinformatic strategy not only allows biologists to studying separately host and symbiont data from a holobiont mixture, but also generates improved transcriptome assemblies. The use of Short Read Connector has proven to be an effective way to tackle meta-transcriptomic challenges to study holobiont systems composed of either well-studied or poorly characterized symbiotic lineages such as the newly sequenced marine plankton Radiolaria-Dinophyta symbiosis and ultimately expand our knowledge about these marine symbiotic associations.

genomics

Identification, expression analysis and molecular modeling of Iron deficiency specific clone 3 (Ids3) like gene in hexaploid wheat

Graminaceous plants secrete iron (Fe) chelators called mugineic acid family phytosiderophores (MAs) from their roots for solubilisation and mobilization of unavailable ferric (Fe3+) ions from the soil. The hydroxylated forms of these phytosiderophores have been found more efficient in chelation and subsequent uptake of minerals from soil which are available in very small quantities. The genes responsible for hydroxylation of phytosiderophores have been recognized as iron deficiency-specific clone 2 (Ids2) and iron deficiency-specific clone 3 (Ids3) in barley but their presence is not reported earlier in hexaploid wheat. Hence, the present investigation was done with the aim:(i) to search for the putative Hordeum vulgare Ids3 (HvIds3) ortholog in hexaploid wheat, (ii) physical mapping of HvIds3 ortholog on wheat chromosome using cytogenetic stocks developed in the background of wheat cultivar Chinese Spring and (iii) to analyze the effect of iron starvation on the expression pattern of this ortholog at transcription level. In the present investigation, a putative ortholog of HvIds3 gene was identified in hexaploid wheat using different bioinformatics tools. Further, protein structure of TaIDS3 was modelled using homology modeling and also evaluated modelled structure behavior on nanoseconds using molecular dynamics based approach. Additionally, the ProFunc results also predict the functional similarity between the proteins of HvIds3 and its wheat ortholog (TaIds3). The physical mapping study with the use of cytogenetic stocks confines TaIds3 in the telomeric region of chromosome 7AS which supports the results obtained by bioinformatics analysis. The relative expression analysis of TaIds3 indicated that the detectable expression of TaIds3 induces after 5th day of Fe-starvation and increases gradually up to 15th day and thereafter decreases till 35th day of Fe-starvation. This reflects that Fe deficiency directly regulates the induction of TaIds3 in the roots of hexaploid wheat.

genomics

Best Practices for Benchmarking Germline Small Variant Calls in Human Genomes

Assessing accuracy of NGS variant calling is immensely facilitated by a robust benchmarking strategy and tools to carry it out in a standard way. Benchmarking variant calls requires careful attention to definitions of performance metrics, sophisticated comparison approaches, and stratification by variant type and genome context. The Global Alliance for Genomics and Health (GA4GH) Benchmarking Team has developed standardized performance metrics and tools for benchmarking germline small variant calls. This team includes representatives from sequencing technology developers, government agencies, academic bioinformatics researchers, clinical laboratories, and commercial technology and bioinformatics developers for whom benchmarking variant calls is essential to their work. Benchmarking variant calls is a challenging problem for many reasons:\n\nO_LIEvaluating variant calls requires complex matching algorithms and standardized counting because the same variant may be represented differently in truth and query callsets.\nC_LIO_LIDefining and interpreting resulting metrics such as precision (aka positive predictive value = TP/(TP+FP)) and recall (aka sensitivity = TP/(TP+FN)) requires standardization to draw robust conclusions about comparative performance for different variant calling methods.\nC_LIO_LIPerformance of NGS methods can vary depending on variant types and genome context; and as a result understanding performance requires meaningful stratification.\nC_LIO_LIHigh-confidence variant calls and regions that can be used as \"truth\" to accurately identify false positives and negatives are difficult to define, and reliable calls for the most challenging regions and variants remain out of reach.\nC_LI\n\nWe have made significant progress on standardizing comparison methods, metric definitions and reporting, as well as developing and using truth sets. Our methods are publicly available on GitHub (https://github.com/ga4gh/benchmarking-tools) and in a web-based app on precisionFDA, which allow users to compare their variant calls against truth sets and to obtain a standardized report on their variant calling performance. Our methods have been piloted in the precisionFDA variant calling challenges to identify the best-in-class variant calling methods within high-confidence regions. Finally, we recommend a set of best practices for using our tools and critically evaluating the results.

genomics

Comprehensive analysis of small RNA profile by massive parallel sequencing in HTLV-1 asymptomatic subjects with monoclonal and polyclonal rearrangement of the T-cell antigen receptor γ-chain

IntroductionIn this study, we used a massive parallel sequencing technology to investigate the cellular small RNA (sRNA) operating in peripheral blood mononuclear cells (PBMCs) of the Human T-lymphotropic virus type I (HTLV-I) infected asymptomatic subjects with a monoclonal and polyclonal rearrangement of the T-cell antigen receptor {gamma}-chain.\n\nMaterials and MethodsBlood samples from 15 HTLV-1 asymptomatic carriers who were tested for clonal TCR-{gamma} gene (seven and eight subjects presented monoclonal and polyclonal expansion of HTLV-1 infected cells, respectively), and were submitted to Illumina for small RNA library construction. sRNA libraries were prepared from cryoperserved PBMCs using TrueSeq Small RNA Library Preparation Kit (Illumina). The sRNA-Seq reads were aligned, annotated, and profiled by different bioinformatics tools.\n\nResultsThrough bioinformatics analysis, we identified a total of 494 known sRNAs and 120 putative novel sRNAs. Twenty-two known and 15 novel sRNA showed a different expression (>2-fold) between the asymptomatic monoclonal (ASM) and asymptomatic polyclonal carriers (ASP). The hsa-mir-196a-5p was the most abundantly upregulated micro RNA (miRNA) and the hsa-mir-133a followed by hsa-mir-509-3p were significantly downregulated miRNAs with more than a three-fold difference in the ASM than ASP group. The target genes predicted to be regulated by the differentially expressed miRNAs play essential roles in diverse biological processes including cell proliferation, differentiation, and/or apoptosis.\n\nDiscussionOur results provide an opportunity for a further understanding of sRNA regulation and function in HTLV-1 infected subjects with monoclonality evidence.

microbiology

Nach is a novel ancestral subfamily of the CNC-bZIP transcription factors selected during evolution from the marine bacteria to human

All living organisms have undergone the evolutionary selection under the changing natural environments to survive as diverse life forms. All life processes including normal homeostatic development and growth into organismic bodies with distinct cellular identifications, as well as their adaptive responses to various intracellular and environmental stresses, are tightly controlled by signaling of transcriptional networks towards regulation of cognate genes by many different transcription factors. Amongst them, one of the most conserved is the basic-region leucine zipper (bZIP) family. They play vital roles essential for cell proliferation, differentiation and maintenance in complex multicellular organisms. Notably, an unresolved divergence on the evolution of bZIP proteins is addressed here. By a combination of bioinformatics with genomics and molecular biology, we have demonstrated that two of the most ancestral family members classified into BATF and Jun subgroups are originated from viruses, albeit expansion and diversification of the bZIP superfamily occur in different vertebrates. Interestingly, a specific ancestral subfamily of bZIP proteins is identified and also designated Nach (Nrf and CNC homology) on account of their highly conservativity with NF-E2 p45 subunit-related factors Nrf1/2. Further experimental evidence reveals that Nach1/2 from the marine bacteria exerts distinctive functions from Nrf1/2 in the transcriptional ability to regulate antioxidant response element (ARE)-driven cytoprotective genes. Collectively, an insight into Nach/CNC-bZIP proteins provides a better understanding of distinct biological functions between these factors selected during evolution from the marine bacteria to human.\n\nSignificanceWe identified the novel ancestral subfamily (i.e. Nach) of CNC-bZIP transcription factors with highly conservativity from marine bacteria to human. Combination of bioinformatics with genomics and molecular biology demonstrated that two of the most ancestral family members classified into BATF and Jun subgroups are originated from viruses. The Jun and CNC subfamilies also share a common origin of these bZIP proteins. Further experimental evidence reveals that Nach1/2 from the marine bacteria exerts nuance functions from human Nrf1/2 in the transcriptional ability to regulate antioxidant response element (ARE)-driven genes, responsible for the host cytoprotection against inflammation and cancer. Overall, this study is of multidisciplinary interests to provide a better understanding of distinct biological functions between Nach/CNC-bZIPs selected during evolution.

ecology

Mono-homologous linear DNA recombination by the non-homologous end-joining pathway as a novel and simple gene inactivation method: a proof of concept study in Dietzia sp. DQ12-45-1b

Non-homologous end-joining (NHEJ) is critical for genome stability because of its roles in double-strand break repair. Ku and ligase D (LigD) are the crucial proteins in this process, and strains expressing Ku and LigD can cyclize linear DNA in vivo. Herein, we established a proof-of-concept mono-homologous linear DNA recombination for gene inactivation or genome editing by which cyclization of linear DNA in vivo by NHEJ could be used to generate non-replicable circular DNA and could allow allelic exchanges between the circular DNA and the chromosome. We achieved this approach in Dietzia sp. DQ12-45-1b, which expresses Ku and LigD homologs and presents NHEJ activity. By transforming the strain with a linear DNA mono homolog to the sequence in chromosome, we mutated the genome. This method did not require the screening of suitable plasmids and was easy and time-effective. Bioinformatic analysis showed that more than 20% prokaryotic organisms contain Ku and LigD, suggesting the wide distribution of NHEJ activities. Moreover, the Escherichia coli strain also showed NHEJ activity when the Ku and LigD of Dietzia sp. DQ12-45-1b were introduced and expressed in it. Therefore, this method may be a widely applicable genome editing tool for diverse prokaryotic organisms, especially for non-model microorganisms.\n\nIMPORTANCEThe non-model gram-positive bacteria lack efficient genetic manipulation systems, but they express genes encoding Ku and LigD. The NHEJ pathway in Dietzia sp. DQ12-45-1b was evaluated and was used to successfully knockout eleven genes in the genome. Since bioinformatic studies revealed that the putative genes encoding Ku and LigD ubiquitously exist in phylogenetically diverse bacteria and archaea, the mono-homologous linear DNA recombination by the NHEJ pathway could be a potentially applicable genetic manipulation method for diverse non-model prokaryotic organisms.

microbiology

FusoPortal: An interactive repository of hybrid MinION sequenced Fusobacterium genomes improves gene identification and characterization

Here we present FusoPortal, an interactive repository of Fusobacterium genomes that were sequenced using a hybrid MinION long-read sequencing pipeline, followed by assembly and annotation using a diverse portfolio of predominantly open-source software. Significant efforts were made to provide genomic and bioinformatic data as downloadable files, including raw sequencing reads, genome maps, gene annotations, protein functional analysis and classifications, and a custom BLAST server for FusoPortal genomes. FusoPortal has been initiated with eight complete genomes, of which seven were previously only drafts that varied from 24-67 contigs. We showcase that genomes in FusoPortal provide accurate open reading frame annotations, and have corrected a number of large genes (>3 kb) that were previously misannotated due to contig boundaries. In summary, FusoPortal (http://fusoportal.org) is the first database of MinION sequenced and completely assembled Fusobacterium genomes, and this central Fusobacterium genomic and bioinformatic resource will aid the scientific community in developing a deeper understanding of how this human pathogen contributes to an array of diseases including periodontitis and colorectal cancer.\n\nImportanceIn this study, we report a hybrid MinION whole genome sequencing pipeline, and describe the genomic characteristics of the first eight strains deposited in the FusoPortal database. This collection of highly accurate and complete genomes drastically improves upon previous multi-contig assemblies by correcting and newly identifying a significant number of open reading frames. We believe this resource will result in the discovery of proteins and molecular mechanisms used by an oral pathogen, with the potential to further our understanding of how F. nucleatum contributes to a repertoire of diseases including periodontitis, pre-term birth, and colorectal cancer

microbiology

In Silico Analysis Reveals a Shared Immune Signature in CASP8-Mutated Carcinomas with Varying Correlations to Prognosis.

BackgroundSequencing studies across multiple cancers continue to reveal the spectrum of mutations and genes involved in the pathobiology of these cancers. Exome sequencing of oral cancers, a subset of Head and Neck Squamous cell Carcinomas (HNSCs) common among tobacco-chewing populations, revealed that ~34% of the affected patients harbor mutations in the CASP8 gene. Uterine Corpus Endometrial Carcinoma (UCEC) is another cancer type where about 10% cases harbor CASP8 mutations. Caspase-8, the protease encoded by CASP8 gene, plays a dual role in programmed cell death, which in turn has an important role in tumor cell death and drug resistance. CASP8 is a protease required for the extrinsic pathway of apoptosis and is also a negative regulator of necroptosis. Using bioinformatics approaches to mine data in The Cancer Genome Atlas, we compared the molecular features and survival of these carcinomas with and without CASP8 mutations.\n\nResultsOur in silico analyses showed that HNSCs with CASP8 mutations displayed a prominent signature of genes involved in immune response and inflammation, and were rich in immune cell infiltrates. However, in contrast to Human Papilloma Virus-positive HNSCs, a subtype that exhibits high immune cell infiltration and better overall survival, HNSC patients with mutant-CASP8 tumors did not display any survival advantage. A similar bioinformatic analyses in UCECs revealed that while UCECs with CASP8 mutations also displayed an immune signature, they had better overall survival, in contrast to the HNSC scenario. On further examination, we found that there was significant up-regulation of neutrophils as well as the cytokine, IL33 in mutant-CASP8 HNSCs, both of which were not observed in mutant-CASP8 UCECs.\n\nConclusionsThese results suggested that carcinomas with mutant CASP8 have broadly similar immune signatures albeit with different effects on survival. We hypothesize that subtle tissue-dependent differences could influence survival by modifying the micro-environment of mutant-CASP8 carcinomas. High neutrophil numbers, which is a well-known negative prognosticator in HNSCs, and/or high IL33 levels may be some of the factors affecting survival of mutant-CASP8 cases.

cancer biology

Stabilized Independent Component Analysis outperforms other methods in finding reproducible signals in tumoral transcriptomes

MotivationMatrix factorization methods are widely exploited in order to reduce dimensionality of transcriptomic datasets to the action of few hidden factors (metagenes). Applying such methods to similar independent datasets should yield reproducible inter-series outputs, though it was never demonstrated yet.\n\nResultsWe systematically test state-of-art methods of matrix factorization on several transcriptomic datasets of the same cancer type. Inspired by concepts of evolutionary bioinformatics, we design a new framework based on Reciprocally Best Hit (RBH) graphs in order to benchmark the methods reproducibility. We show that a particular protocol of application of Independent Component Analysis (ICA), accompanied by a stabilisation procedure, leads to a significant increase in the inter-series output reproducibility. Moreover, we show that the signals detected through this method are systematically more interpretable than those of other state-of-art methods. We developed a user-friendly tool BIODICA for performing the Stabilized ICA-based RBH meta-analysis. We apply this methodology to the study of colorectal cancer (CRC) for which 14 independent publicly available transcriptomic datasets can be collected. The resulting RBH graph maps the landscape of interconnected factors that can be associated to biological processes or to technological artefacts. These factors can be used as clinical biomarkers or robust and tumor-type specific transcriptomic signatures of tumoral cells or tumoral microenvironment. Their intensities in different samples shed light on the mechanistic basis of CRC molecular subtyping.\n\nAvailabilityThe BIODICA tool is available from https://github.com/LabBandSB/BIODICA.\n\nContactlaura.cantini@curie.fr and andrei.zinovyev@curie.fr\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

High-throughput identification and marker development of perfect SSR for cultivated genus of passion fruit (Passiflora edulis)

Simple sequence repeat (SSR) markers are characterized by high polymorphism, good reproducibility and co-dominance etc. They can be easily applied to develop efficient, simple and practical molecular markers. In the present study, bioinformatics methods were applied to identify high-throughput perfect SSRs of cultivar Passiflora genome. A total of 13104 perfect SSRs were obtained. SSR core sequence structure is mainly 2-4 bases, the maximum numbers are TA, AT, TC and AG. The maximum numbers of repetitions were up to 20 times. A total of 12934 pairs of SSR markers were developed by using bioinformatics software, and 20 pairs of markers were selected for amplification specificity assessment of MTX and WJ10, and the polymorphism rate was as high as 60%. The large-scale development of the SSR markers of Passiflora cultivar has paved a foundation for the efficient utilization of the germplasm resources of passion fruit, genetic improvement of the varieties and molecular breeding.

genomics

Molecular analysis of long non-coding RNA GAS5 and microRNA-34a expression signature in common solid tumors: A pilot study

Accumulating evidence indicates that non-coding RNAs including microRNAs (miRs) and long non-coding RNAs (lncRNAs) are aberrantly expressed in cancer, providing promising biomarkers for diagnosis, prognosis and/or therapeutic targets. We aimed in the current work to quantify the expression profile of miR-34a and one of its bioinformatically selected partner lncRNA growth arrest-specific 5 (GAS5) in a sample of Egyptian cancer patients, including three prevalent types of cancer in our region; renal cell carcinoma (RCC), hepatocellular carcinoma (HCC) and glioblastoma (GB) as well as to correlate these expression profiles with the available clinicopathological data in an attempt to clarify their roles in cancer. Quantitative real-time polymerase chain reaction analysis was applied. Different bioinformatics databases were searched to confirm the potential miRNAs-lncRNA interactions of the selected ncRNAs in cancer pathogenesis. GAS5 was significantly under-expressed in the three types of cancer. However, levels of miR-34a greatly varied according to the tumor type; it displayed an increased expression in RCC [4.05 (1.003-22.69), p <0.001] and a decreased expression in GB [0.35 (0.04-0.95), p <0.001]. A weak negative correlation was observed between levels of GAS5 and miR-34a in GB [r = -0.39, p =0.006]. Univariate analyses revealed a correlation of GAS5 downregulation with poor disease-free survival (r = 0.31, p =0.018) and overall survival (r = 0.28, p =0.029) in RCC but not in GB, and a marginal significance correlation with a higher number of lesions in HCC. Hierarchical clustering analysis showed RCC patients among others, could be clustered by GAS5 and miR-34a co-expression profile. Our results confirm the tumor suppressor role of GAS5 in cancer and suggest its potential applicability to be a predictor of bad outcomes with other conventional markers for various types of cancer. Further functional validation studies are warranted to confirm miR-34a/GAS5 interplay in cancer.

molecular biology

Combining mathematical and statistical modeling to simulate time course bulk and single cell gene expression data in cancer with CancerInSilico

Bioinformatics techniques to analyze time course bulk and single cell omics data are advancing. The absence of a known ground truth of the dynamics of molecular changes challenges benchmarking their performance on real data. Realistic simulated time-course datasets are essential to assess the performance of time course bioinformatics algorithms. We develop an R/Bioconductor package, CancerInSilico, to simulate bulk and single cell transcriptional data from a known ground truth obtained from mathematical models of cellular systems. This package contains a general R infrastructure for running cell-based models and simulating gene expression data based on the model states. We show how to use this package to simulate a gene expression data set and consequently benchmark analysis methods on this data set with a known ground truth. The package is freely available via Bioconductor: http://bioconductor.org/packages/CancerInSilico/

systems biology

Identification of a Nocardia seriolae secreted protein targeting host cell mitochondria and inducing apoptosis in fathead minnow (FHM) cells

Nocardia seriolae, is a Gram-positive, partially acid-fast, aerobic, and filamentous bacterium. This bacterium is the main pathogen of fish nocardiosis. A bioinformatic analysis based on the genomic sequence of the N. seriolae strain ZJ0503 showed that ORF3141 encoded a secreted protein with a signal peptide at the N-terminate which may target the mitochondria in the host cell. However, the functions of this protein and its homologs remain unknown. In this study, we experimentally tested the bioinformatic prediction on this protein. Mass spectrometry analysis of the extracellular products from N. seriolae showed that ORF3141 was a secreted protein. Subcellular localization of the ORF3141-GFP fusion protein revealed that the green fluorescence protein co-localized with the mitochondria, while ORF3141{Delta}sig-GFP (with the signal peptide deleted) fusion protein was evenly distributed in the whole cell of fathead minnow (FHM) cells. Thus, the N-terminate signal peptide had a significant impact on mitochondrial targeting. Notably, the expression of ORF3141 protein changed the distribution of mitochondria from perinuclear halo into lumps in the transfected FHM cells. In addition, apoptotic features were found in the transfected FHM cells by overexpression of ORF3141 and ORF3141{Delta}sig proteins, respectively. Quantitative assays of mitochondrial membrane potential value, caspase-3 activity and apoptosis-related gene mRNA expression suggested that cell apoptosis was induced in the transfected FHM cells. In conclusion, the ORF3141 was a secreted protein of N. seriolae that targeted host cell mitochondria and induced apoptosis in FHM cells. This protein may participate in the cell apoptosis regulation and plays an important role in the pathogenesis of N. seriolae.\n\nAuthor summaryNocardia seriolae is the causative pathogen responsible for fish nocardiosis. This facultative intercellular bacterium, adapts to survive and colonize by evading intracellular killing after being engulfed with macrophages in the host. Despite considerable economic losses caused by N. seriolae in fish infection, the pathogenic mechanism and specific virulence factor of this bacterium remain ambiguous. In this study, the characteristic of ORF3141 protein function was investigated by subcellular localization and its possible contributions on the ability of N. seriolae to induce apoptosis in transfected fathead minnow (FHM) cells was investigated. Here, we confirmed that ORF3141 was a secreted protein that targeted host cell mitochondria and induced cell apoptosis in FHM cells. Interestingly, after deleting the signal peptide, ORF3141{Delta}sig protein was evenly distributed in the whole host cell and did not co-localize with the mitochondria which could also induce cell apoptosis. Thus, the N-terminate signal peptide played an important role in mitochondrial targeting, and the domain part without the signal peptide had a critical relationship with cell apoptosis. These results demonstrated that ORF3141 mays act as a potential virulence factor that induces apoptosis in fish cells. This protein is significant to elucidate the pathogenic mechanism of N. seriolae and this study mays provide beneficial insight to prevent and treat fish nocardiosis.

pathology

An efficient and improved laboratory workflow and tetrapod database for larger scale eDNA studies

BackgroundThe use of environmental DNA, eDNA, for species detection via metabarcoding is growing rapidly. We present a co-designed lab workflow and bioinformatic pipeline to mitigate the two most important risks of eDNA: sample contamination and taxonomic mis-assignment. These risks arise from the need for PCR amplification to detect the trace amounts of DNA combined with the necessity of using short target regions due to DNA degradation. FindingsOur high-throughput workflow minimises these risks via a four-step strategy: (1) technical replication with two PCR replicates and two extraction replicates; (2) using multi-markers (12S, 16S, CytB); (3) a twin-tagging, two-step PCR protocol;(4) use of the probabilistic taxonomic assignment method PROTAX, which can account for incomplete reference databases. As annotation errors in the reference sequences can result in taxonomic mis-assignment, we supply a protocol for curating sequence datasets. For some taxonomic groups and some markers, curation resulted in over 50% of sequences being deleted from public reference databases, due to (1) limited overlap between our target amplicon and reference sequences; (2) mislabelling of reference sequences; (3) redundancy. Finally, we provide a bioinformatic pipeline to process amplicons and conduct PROTAX assignment and tested it on an invertebrate derived DNA (iDNA) dataset from 1532 leeches from Sabah, Malaysia. Twin-tagging allowed us to detect and exclude sequences with non-matching tags. The smallest DNA fragment (16S) amplified most frequently for all samples, but was less powerful for discriminating at species rank. Using a stringent and lax acceptance criteria we found 162 (stringent) and 190 (lax) vertebrate detections of 95 (stringent) and 109 (lax) leech samples. ConclusionsOur metabarcoding workflow should help research groups increase the robustness of their results and therefore facilitate wider usage of e/iDNA, which is turning into a valuable source of ecological and conservation information on tetrapods.

molecular biology

Metagenomic screening of global microbiomes identifies pathogen-enriched environments

BackgroundHuman pathogens are widespread in the environment, and examination of pathogen-enriched environments in a rapid and high-throughput fashion is important for development of pathogen-risk precautionary measures.\n\nMethodsIn this study, a Local BLASTP procedure for metagenomic screening of pathogens in the environment was developed using a toxin-centered database. A total of 27 microbiomes derived from ocean water, freshwater, soil, feces, and wastewater were screened using the Local BLASTP procedure. Bioinformatic analysis and Canonical Correspondence Analysis were conducted to examine whether the toxins included in the database were taxonomically associated.\n\nResultsThe specificity of the Local BLASTP method was tested with known and unknown toxin sequences. Bioinformatic analysis indicated that most toxins were phylum-specific but not genus-specific. Canonical Correspondence Analysis implied that almost all of the toxins were associated with the phyla of Proteobacteria, Nitrospirae and Firmicutes. Local BLASTP screening of the global microbiomes showed that pore-forming RTX toxin and adenylate cyclase Cya were most prevalent globally in terms of relative abundance, while polluted water and feces samples were the most pathogen-enriched.\n\nConclusionsA Local BLASTP procedure was established for rapid detection of toxins in environmental samples. Screening of global microbiomes in this study provided a quantitative estimate of the most prevalent toxins and most pathogen-enriched environment.

ecology

artMAP: a user-friendly tool for mapping EMS-induced mutations in Arabidopsis

Mapping-by-sequencing is a rapid method for identifying both natural as well as induced variations in the genome. However, it requires extensive bioinformatics expertise along with the computational infrastructure to analyze the sequencing data and these requirements have limited its widespread adoption. In the current study, we develop an easy to use tool, artMAP, to discover ethyl methanesulfonate (EMS) induced mutations in the Arabidopsis genome. The artMAP pipeline consists of well-established tools including TrimGalore, BWA, BEDTools, SAMtools, and SnpEff which were integrated in a Docker container. artMAP provides a graphical user interface and can be run on a regular laptop and desktop, thereby limiting the bioinformatics expertise required. artMAP can process input sequencing files generated from single or paired-end sequencing. The results of the analysis are presented in interactive graphs which display the annotation details of each mutation. Due to its ease of use, artMAP made the identification of EMS-induced mutations in Arabidopsis possible with only a few mouse click. The source code of artMAP is available on Github (https://github.com/RihaLab/artMAP).

plant biology

Double-digest RAD-sequencing: do wet and dry protocol parameters impact biological results?

O_LINext-generation sequencing technologies have opened a new era of research in genomics. Among these, restriction enzyme-based techniques such as restriction-site associated DNA sequencing (RADseq) or double-digest RAD-sequencing (ddRADseq) are now widely used in many population genomics fields. From DNA sampling to SNP calling, both wet and dry protocols have been discussed in the literature to identify key parameters for an optimal loci reconstruction.\nC_LIO_LIThe impact of these parameters on downstream analyses and biological results drawn from RADseq or ddRADseq data has however not been fully explored yet. In this study, we tackled this issue by investigating the effects of ddRADseq laboratory (i.e. wet protocol) and bioinformatics (i.e. dry protocol) settings on loci reconstruction and inferred biological signal at two evolutionary scale using two systems: a complex of butterfly species (Coenonympha sp.) and populations of Common beech (Fagus sylvatica).\nC_LIO_LIResults suggest an impact of wet protocol parameters (DNA quantity, number of PCR cycles during library preparation) on the number of recovered reads and SNPs, the number of unique alleles and individual heterozygosity. We also found that bioinformatic settings (i.e. clustering and minimum coverage thresholds) impact loci reconstruction (e.g. number of loci, mean coverage) and SNP calling (e.g. number of SNPs, heterozygosity). We however do not detect an impact of parameter settings on three types of analysis performed with ddRADseq data: measure of genetic differentiation, estimation of individual admixture, and demographic inferences. In addition, our work demonstrates the high reproducibility and low rate of genotyping inconsistencies of the ddRADseq protocol.\nC_LIO_LIThus, our study highlights the impact of wet parameters on ddRADseq protocol with strong consequences on experimental success and biological conclusions. Dry parameters affects loci reconstruction and descriptive statistics but not biological conclusion for the two studied systems. Overall, this study illustrates, with others, the relevance of ddRADseq for population and evolutionary genomics at the inter- or intraspecific scales.\nC_LI

molecular biology