Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

SciPipe - A workflow library for agile development of complex and dynamic bioinformatics pipelines

BackgroundThe complex nature of biological data has driven the development of specialized software tools. Scientific workflow management systems simplify the assembly of such tools into pipelines, assist with job automation and aid reproducibility of analyses. Many contemporary workflow tools are specialized and not designed for highly complex workflows, such as with nested loops, dynamic scheduling and parametrization, which is common in e.g. machine learning.\n\nFindingsSciPipe is a workflow programming library implemented in the programming language Go, for managing complex and dynamic pipelines in bioinformatics, cheminformatics and other fields. SciPipe helps in particular with workflow constructs common in machine learning, such as extensive branching, parameter sweeps and dynamic scheduling and parametrization of downstream tasks. SciPipe builds on Flow-based programming principles to support agile development of workflows based on a library of self-contained, re-usable components. It supports running subsets of workflows for improved iterative development, and provides a data-centric audit logging feature that saves a full audit trace for every output file of a workflow, which can be converted to other formats such as HTML, TeX and PDF on-demand. The utility of SciPipe is demonstrated with a machine learning pipeline, a genomics, and a transcriptomics pipeline.\n\nConclusionsSciPipe provides a solution for agile development of complex and dynamic pipelines, especially in machine leaning, through a flexible programming API suitable for scientists used to programming or scripting.

bioinformatics

Bioinformatics workflows for genomic analysis of tumors from Patient Derived Xenografts (PDX): challenges and guidelines

Bioinformatics workflows for analyzing genomic data obtained from xenografted tumor (e.g., human tumors engrafted in a mouse host) must address several challenges, including separating mouse and human sequence reads and accurate identification of somatic mutations and copy number aberrations when paired normal DNA from the patient is not available. We report here data analysis workflows that address these challenges and result in reliable identification of somatic mutations, copy number alterations, and transcriptomic profiles of tumors from patient derived xenograft models. We validated our analytical approaches using simulated data and by assessing concordance of the genomic properties of xenograft tumors with data from primary human tumors in The Cancer Genome Atlas (TCGA). The commands and parameters for the workflows are available at https://github.com/TheJacksonLaboratory/PDX-Analysis-Workflows.

bioinformatics

Neutools: A collection of bioinformatics web tools for neurogenomic analysis

With a large and growing number of neurogenomic data, from epigenetic, transcriptomic to proteomic data deposited to the public domains, visualization and mining of these data alongside one own data have become extremely useful for identifying potential genetic targets and/or biological pathways to further validate and characterize in model systems. Here, a series of easy-to-use web tools (Neutools) were developed using Shiny/R codes for neuroscientists to perform basic bioinformatics data analysis and visualization. Specifically, NeuVenn calculates and plots overlap statistics for multiple input gene sets; NeuGene and NeuChIP visualize gene expression data and histone ChIP-Seq data generated from brain related tissue and cell culture, respectively. NeuVar annotates human brain GWAS variants and epigenetic features based on user-specified genes and regions. Neutools are freely available resources for all academic users to use.

bioinformatics

Aminoacyl tRNA Synthetases as Malarial Drug Targets: A Comparative Bioinformatics Study

Treatment of parasitic diseases has been challenging due to the development of drug resistance by parasites, and thus there is need to identify new class of drugs and drug targets. Protein translation is important for survival of plasmodium and the pathway is present in all the life cycle stages of the plasmodium parasite. Aminoacyl tRNA synthetases are primary enzymes in protein translation as they catalyse the first reaction where an amino acid is added to the cognate tRNA. Currently, there is limited research on comparative studies of aminoacyl tRNA synthetases as potential drug targets. The aim of this study is to understand differences between plasmodium and human aminoacyl tRNA synthetases through bioinformatics analysis. Plasmodium falciparum, P. fragile, P. vivax, P. ovale, P. knowlesi, P. bergei, P. malariae and human aminoacyl tRNA synthetase sequences were retrieved from UniProt database and grouped into 20 families based on amino acid specificity. Despite functional and structural conservation, multiple sequence analysis, motif discovery, pairwise sequence identity calculations and molecular phylogenetic analysis showed striking differences between parasite and human proteins. Prediction of alternate binding sites revealed potential druggable sites in PfArgRS, PfMetRS and PfProRS at regions that were weakly conserved when compared to the human homologues. These differences provide a basis for further exploration of plasmodium aminoacyl tRNA synthetases as potential drug targets.

bioinformatics

Probing the transcriptome response to shivering in skeletal muscle using a multilayered bioinformatics approach

Cold acclimation holds therapeutic potential for improving metabolic health. We previously demonstrated that repeated cold-induced shivering enhances insulin sensitivity in humans. However, the molecular pathways that underlie the skeletal muscle shivering response, and how these relate to beneficial physiological effects, remain poorly understood. In this study, we combined complementary bioinformatics approaches to allow in-depth analysis of the transcriptomic response of human skeletal muscle to repeated shivering. We identified a robust transcriptional signature and show a sex-specific component in the shivering skeletal muscle response, which seemed to diminish following cold adaptation. Our findings provide mechanistic insights into cold-induced muscle adaptations, shed light on potential interesting molecular targets for further investigation, and emphasize the importance of including both sexes in future cold acclimation studies.

bioinformatics

Alternative splicing detection workflow needs a careful combination of sample prep and bioinformatics analysis

BackgroundRNAseq provides remarkable power in the area of biomarkers discovery and disease stratification. The main technical steps affecting the results of RNAseq experiments are Library Sample Preparation (LSP) and Bioinformatics Analysis (BA). At the best of our knowledge, a comparative evaluation of the combined effect of LSP and BA was never considered and it might represent a valuable knowledge to optimize alternative splicing detection, which is a challenging task due to moderate fold change differences to be detected within a complex isoforms background.\n\nResultsDifferent LSPs (TruSeq unstranded/stranded, ScriptSeq, NuGEN) allow the detection of a large common set of isoforms. However, each LSP also detects a smaller set of isoforms, which are characterized both by lower coverage and lower FPKM than that observed for the common ones among LSPs. This characteristic is particularly critical in case of low input RNA NuGEN v2 LSP.\n\nThe effect on statistical detection of alternative splicing considering low input LSP (NuGEN v2) with respect to high input LSP (TruSeq) on statistical detection of alternative splicing was studied using a benchmark dataset, in which both synthetic reads and reads generated from high (TruSeq) and low input (NuGEN) LSPs were spiked-in. Statistical detection of alternative splicing (AltDE) was done using prototypes of BA for isoform-reconstruction (Cuffdiff) and exon-level analysis (DEXSeq). Exon-level analysis performs slightly better than isoform-reconstruction approach although at most only 50% of the spiked-in transcripts are detected. Both isoform-reconstruction and exon-level analysis performances improve by rising the number of input reads.\n\nConclusionData, derived from NuGEN v2, are not the ideal input for AltDE, specifically when exon-level approach is used. It is notable that ribosomal depletion, with respect to polyA+ selection, reduces the amount of coding mappable reads resulting detrimental in the case of AltDE. Furthermore, we observed that both isoform-reconstruction and exon-level analysis performances are strongly dependent on the number of input reads.

Genomics

A Critical Review on the Use of Support Values in Tree Viewers and Bioinformatics Toolkits

Phylogenetic trees are routinely visualized to present and interpret the evolutionary relationships of species. Virtually all empirical evolutionary data studies contain a visualization of the inferred tree with branch support values. Ambiguous semantics in tree file formats can lead to erroneous tree visualizations and therefore to incorrect interpretations of phylogenetic analyses.\n\nHere, we discuss problems that can and do arise when displaying branch values on trees after re-rooting. Branch values are typically stored as node labels in the widely-used Newick tree format. However, such values are attributes of branches. Storing them as node labels can therefore yield errors when re-rooting trees. This depends on the mostly implicit semantics that tools deploy to interpret node labels.\n\nWe reviewed 10 tree viewers and 10 bioinformatics toolkits that can display and re-root trees. We found that 14 out of 20 of these tools do not permit users to select the semantics of node labels. Thus, unaware users might obtain incorrect results when rooting trees inferred by common phylogenetic inference programs. We illustrate such incorrect mappings for several test cases and real examples taken from the literature. This review has already led to improvements and workarounds in 8 of the tested tools. We suggest tools should provide an option that explicitly forces users to define the semantics of node labels.

Evolutionary Biology

Omics and bioinformatics approaches to target boar taint

In livestock species, a rapid growth in high-throughput omics data has accelerated the pace of studies that target to dissect economically important traits to provide better quality animal products to consumers. In pig industries, young boars are generally castrated to remove boar taint, a phenotypic and inheritable trait well-known by an abnormally bad smell and taste in pork meat derived from some uncastrated male pigs. Existence of porcine reference genome made possible to catalogue genome-wide QTLs, candidate genes and biomarkers in associations with boar taint and other industrially significant traits in pigs. The aim of this paper to review the contribution of bioinformatics resources and omics technology in boar taint related studies. This paper also provides concise details about state-of-the-art sequencing technology.

Genomics

A bioinformatic panel to interrogate thousands of ExAC variants with minor reference allele that are missed by conventional variant calling

In variation sites with minor reference alleles, overlooking the detection of homozygous reference genotypes results in inadequate identification of potential disease variants. Current variant calling practices miss these clinically relevant alleles warranting new approaches. More than 26,000 Eome Aggregation Consortium (ExAC) variants have a minor reference allele including 44 variants with known ClinVar disease alleles. We demonstrated how the current variant calling standards miss homozygous reference disease variants in these sites. We developed a bioinformatic panel that can be used to screen these variants using commonly available variant callers. We provide here a simple strategy to screen potential disease-causing variants when present in homozygous reference state.

genomics

A comprehensive automated pipeline for human microbiome sampling, 16S rRNA gene sequencing and bioinformatics processing

The advent of affordable high-throughput DNA sequencing has opened up a golden age of studies in the human microbiome. In order to understand the role of the human microbiota, standardized methods for large-scale, population-level studies are needed to avoid underpowered or poorly designed studies. The biggest bottlenecks to population-level microbiomics are sample collection, storage and DNA extraction. Here, we describe a flexible automated approach to process intestinal biopsies, fecal samples and vaginal swabs from sample collection to OTU table. We have evaluated storage conditions, DNA extraction methods, PCR strategies and bioinformatic pipelines for these three sample types, and present here a set of guidelines and best practices for each of these steps.

microbiology

Molecular Detection of H.pylori Antibiotic-Resistant Genes and Bioinformatics Predictive Analysis

To explore the mutation characteristics of H.pylori resistance-related genes to antibiotics of clarithromycin, levofloxacin and metronidazole. 23S rRNA, gyrA, gyrB, rdxA and frxA genes were amplified and sequenced, respectively. Their structural alteration after mutation was predicted using bioinformatics software. In the clarithromycin-resistant strains, the mutation rate in site A2143G was 74.2% (n=23). The mutations in sites C1883T, C2131T and T2179G might cause structural alteration. In the levofloxacin-resistant strains, the mutation rates in 87 (N to K/I) and 91 (D to N/Y/G) of gyrA were 28.6% (n=16) and 12.5% (n =7), respectively. Meanwhile, one of the mutation strains in site 91 was accompanied by D99N variation. Additionally, a D143E mutation was found in one drug-resistant strain. Some changes of tertiary structure occurred after these mutations. The mutation types of RdxA protein consisted of protein truncation caused by premature stop codons (n=26, 33.3%), frameshift mutations (n=8, 10.3%), FMN-binding sites (n=16, 20.5%) and the others (n=11, 14.1%). Predictive analysis showed that mutations in the first three groups and the A118S of the last group could lead to structural alteration. Our study suggested the clarithromycin-resistant sites of H.pylori were mainly located in A2143G of 23S rRNA. C1883T, C2131T and T2179G might also be related to resistance. Levofloxacin resistance was mainly based on the amino acid changes in 87 and 91 sites of gyrA. The new sites D99N and D143E might also be associated with resistance. Metronidazole resistance was related to RdxA protein truncation, frameshift, and FMN binding. The new site A118S might also be linked to drug resistance.

microbiology

Rapid Therapeutic Recommendations in the Context of a Global Public Health Crisis using Translational Bioinformatics Approaches: A proof-of-concept study using Nipah Virus Infection

We live in a world of emerging new diseases and old diseases resurging in more aggressive forms. Drug development by pharmaceutical companies is a market-driven and costly endeavor, and thus it is often a challenge when drugs are needed for diseases endemic only to certain regions or which affect only a few patients. However, biomedical open data is accessible and reusable for reanalysis and generation of a new hypotheses and discovery. In this study, we leverage biomedical data and tools to analyze available data on Nipah Virus (NiV) infection. NiV infection is an emerging zoonosis that is transmissible to humans and is associated with high mortality rates. In this study, explored the application of computational drug repositioning and chemogenomic enrichment analyses using host transcriptome data to match drugs that could reverse the virus-induced gene signature. We performed analyses using two gene signatures: i) A previously published gene signature (n=34), and ii) a gene signature generated using the characteristic direction method (n= 5,533). Our predictive framework suggests that several drugs including FDA approved therapies like beclometasone, trihexyphenidyl, S-propranolol etc. could modulate the NiV infection induced gene signatures in endothelial cells. A target specific analysis of CXCL10 also suggests the potential application of Eldelumab, an investigative therapy for Crohns disease and ulcerative colitis, as a putative candidate for drug repositioning. To conclude, we also discuss challenges and opportunities in clinical trials (n-of-1 and adaptive trials) for repositioned drugs. Further follow-up studies including biochemical assays and clinical trials are required to identify effective therapies for clinical use. Our proof-of-concept study highlights that translational bioinformatics methods including gene expression analyses and computational drug repositioning could augment epidemiological investigations in the context of an emerging disease with no effective treatment.

microbiology

Targeting microbial arsenic resistance genes: a new bioinformatic toolkit informs arsenic ecology and evolution in soil genomes and metagenomes

Environmental resistomes include transferable microbial genes. One important resistome component is resistance to arsenic, a ubiquitous and toxic metalloid that can have negative and chronic consequences for human and animal health. The distribution of arsenic resistance and metabolism genes in the environment is not well understood. However, microbial communities and their resistomes mediate key transformations of arsenic that are expected to impact both biogeochemistry and local toxicity. We examined the phylogenetic diversity, genomic location (chromosome or plasmid), and biogeography of arsenic resistance and metabolism genes in 922 soil genomes and 38 metagenomes. To do so, we developed a bioinformatic toolkit that includes BLAST databases, hidden Markov models and resources for gene-targeted assembly of nine arsenic resistance and metabolism genes: acr3, aioA, arsB, arsC (grx), arsC (trx), arsD, arsM, arrA, and arxA. Though arsenic related genes were common, they were not universally detected, contradicting the common conjecture that all organisms have them. From major clades of arsenic related genes, we inferred their potential for horizontal and vertical transfer. Different types and proportions of genes were detected across soils, suggesting microbial community composition will, in part, determine local arsenic toxicity and biogeochemistry. While arsenic related genes were globally distributed, particular sequence variants were highly endemic (e.g., acr3), suggesting dispersal limitation. The gene encoding arsenic methylase arsM was unexpectedly abundant in soil metagenomes (median 48%), suggesting that it plays a prominent role in global arsenic biogeochemistry. Our analysis advances understanding of arsenic resistance, metabolism, and biogeochemistry, and our approach provides a roadmap for the ecological investigation of environmental resistomes.

microbiology

WMP: A novel comprehensive wheat miRNA database, including related bioinformatics software

MicroRNAs (miRNAs) are emerging as important post tran-scriptional regulators that may regulate key plant genes responsible for agronomic traits such as grain yield and stress tolerance. Several studies identified species and clades specific miRNA families associated with plant stress regulated genes. Here, we propose a novel resource that provides data related to the expression of abiotic stress responsive miRNAs in wheat, one of the most important staple food crops. This database allows the query of small RNA libraries, including in silico predicted wheat miRNA sequences and the expression profiles of small RNAs identified from those libraries. Our database also provides a direct access to online miRNA prediction software tuned to de novo miRNA detection in wheat, in monocotyledon clades, as well as in other plant species. These data and software will facilitate multiple comparative analyses and reproducible studies on small RNAs and miRNA families in plants. Our web-portal is available at: http://wheat.bioinfo.uqam.ca.

Bioinformatics

Understanding properties of the master effector of phage shock operon in Mycobacterium tuberculosis via bioinformatics approach

The phage shock protein (Psp) is a part of the Psp operon, which assists in safeguarding the survival of bacterium in stress and shields the cell against proton motif force challenge. It is strongly induced by bacterium allied phages, improperly localized mutant porins and various other stresses. Master effector of the operon, PspA has been modeled and simulated, illustrating how it undergoes significant conformational transition at the far end in Mycobacterium tuberculosis. Association of this key protein of the operon influences action of Psp system on the whole. We are further working on the impact of phosphorylation perturbation and changes in the structure of PspA during complex formation with other moieties of interest.

Bioinformatics

Exploratory bioinformatics analysis reveals importance of "junk" DNA in early embryo development

BackgroundInstead of testing predefined hypotheses, the goal of exploratory data analysis (EDA) is to find what data can tell us. Following this strategy, we re-analyzed a large body of genomic data to investigate how the early mouse embryos develop from fertilized eggs through a complex, poorly understood process.\n\nResultsStarting with a single-cell RNA-seq dataset of 259 mouse embryonic cells from zygote to blastocyst stages, we reconstructed the temporal and spatial dynamics of gene expression. Our analyses revealed similarities in the expression patterns of regular genes and those of retrotransposons, and the enrichment of transposable elements in the promoters of corresponding genes. Long Terminal Repeats (LTRs) are associated with transient, strong induction of many nearby genes at the 2-4 cell stages, probably by providing binding sites for Obox and other homeobox factors. The presence of B1 and B2 SINEs (Short Interspersed Nuclear Elements) in promoters is highly correlated with broad upregulation of intracellular genes in a dosage-and distance-dependent manner. Such enhancer-like effects are also found for human Alu and bovine tRNA SINEs. Promoters for genes specifically expressed in embryonic stem cells (ESCs) are rich in B1 and B2 SINEs, but low in CpG islands.\n\nConclusionsOur results provide evidence that transposable elements may play a significant role in establishing the expression landscape in early embryos and stem cells. This study also demonstrates that open-ended, exploratory analysis aimed at a broad understanding of a complex process can pinpoint specific mechanisms for further study.\n\nMajor findingO_LISingle-cell RNA-seq data enables estimation of retrotransposon expression during PD\nC_LIO_LISimilar expression dynamics of retrotransposons and regular genes during PD\nC_LIO_LILong terminal repeats may be essential for the 1st wave of gene expression\nC_LIO_LIObox homeobox factors are possible regulators of PD, upstream of Zscan4\nC_LIO_LISINE repeats predict expression of nearby genes in murine, human and bovine embryos\nC_LIO_LIExploratory analysis of large single-cell data pinpoints developmental pathways\nC_LI

Bioinformatics

Hot-starting software containers for bioinformatics analyses

Using software containers has become standard practice to reproducibly deploy and execute biomedical workflows on the cloud. We demonstrate that hot-starting, from containers that have been frozen after the application has already begun execution, reduces the costs of cloud computing by avoiding repetitive initialization steps. The method is widely applicable and can provide substantial savings both for small jobs and for large-scale deployments using automated schedulers.

bioinformatics

Bioinformatics Workflow Management With The Wobidisco Ecosystem

ReferencesTo conduct our computational experiments, our team developed a set of workflow-management-related projects: Ketrew, Biokepi, and Coclobas. The family of tools and libraries are designed with reliability and flexibility as main guiding principles. We describe the components of the software stack and explain the choices we made. Every piece of software is free and open-source; the umbrella documentation project is available at https://github.com/hammerlab/wobidisco.

bioinformatics