Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Bioinformatics analysis quantifies neighborhood preferences of cancer cells in Hodgkin lymphoma.

MotivationHodgkin lymphoma is a tumor of the lymphatic system and represents one of the most frequent lymphoma in the Western world. It is characterized by Hodgkin cells and Reed-Sternberg cells, which exhibit a broad morphological spectrum. The cells are visualized by immunohistochemical staining of tissue sections. In pathology, tissue images are mainly manually evaluated, relying on the expertise and experience of pathologists. Computational quantification methods become more and more essential to evaluate tissue images. In particular, the distribution of cancer cells is of great interest.\n\nResultsHere, we systematically quantified and investigated cancer cell properties and their spatial neighborhood relations by applying statistical analyses to whole slide images of Hodgkin lymphoma and lymphadenitis, which describes a non-cancerous inflammation of the lymph node. We differentiated cells by their morphology and studied the spatial neighborhood relation of more than 400,000 immunohistochemically stained cells. We found that, according to their morphological features, the cells exhibited significant preferences for and aversions to cells of specific profiles as nearest neighbor. We quantified differences between Hodgkin lymphoma and lymphadenitis concerning the neighborhood relations of cells and the sizes of cells. The approach can easily be applied to other cancer types.\n\nContactina.koch@bioinformatik.uni-frankfurt.de

bioinformatics

Rapid Identification and Validation of Novel Rheumatoid Arthritis Drug Treatments using an Integrative Bioinformatics Platform

The majority of drugs currently used to treat rheumatoid arthritis (RA) act on a small number of immunomodulatory targets. We applied an integrative biomedical-informatics-based approach and in vivo testing to identify new drug candidates and potential therapeutic targets that could form the basis for future drug development in RA. A computational model of RA was constructed by integrating patient gene expression data, molecular interactions, and clinical drug-disease associations. Drug candidates were scored based on their predicted efficacy across these data types. Ten high-scoring candidates were subsequently screened in a collagen-induced arthritis model of RA. Treatment with exenatide, olopatadine, and TXR-112 significantly improved multiple preclinical endpoints, including animal mobility which was measured using a novel digital platform. These three drug candidates do not act on common RA therapeutic targets; however, links between known candidate pharmacology and pathological processes involved in RA suggest hypothetical mechanisms contributing to the observed efficacy.

bioinformatics

Bioinformatic analysis of endogenous and exogenous small RNAs on lipoproteins

To comprehensively study extracellular small RNAs (sRNA) by sequencing (sRNA-seq), we developed a novel pipeline to overcome current limitations in analysis entitled, \"Tools for Integrative Genome analysis of Extracellular sRNAs (TIGER)\". To demonstrate the power of this tool, sRNA-seq was performed on mouse lipoproteins, bile, urine, and liver samples. A key advance for the TIGER pipeline is the ability to analyze both host and non-host sRNAs at genomic, parent RNA, and individual fragment levels. TIGER was able to identify approximately 60% of sRNAs on lipoproteins, and >85% of sRNAs in liver, bile, and urine, a significant advance compared to existing software. Results suggest that the majority of sRNAs on lipoproteins are non-host sRNAs derived from bacterial sources in the microbiome and environment, specifically rRNA-derived sRNAs from Proteobacteria. Collectively, TIGER facilitated novel discoveries of lipoprotein and biofluid sRNAs and has tremendous applicability for the field of extracellular RNA.

bioinformatics

HPCI: A Perl module for writing cluster-portable bioinformatics pipelines

BackgroundMost biocomputing pipelines are run on clusters of computers. Each type of cluster has its own API (application programming interface). That API defines how a program that is to run on the cluster must request the submission, content and monitoring of jobs to be run on the cluster. Sometimes, it is desirable to run the same pipeline on different types of cluster. This can happen in situations including when:\n\nO_LIdifferent labs are collaborating, but they do not use the same type of cluster\nC_LIO_LIa pipeline is released to other labs as open source or commercial software\nC_LIO_LIa lab has access to multiple types of cluster, and wants to choose between them for scaling, cost or other purposes\nC_LIO_LIa lab is migrating their infrastructure from one cluster type to another\nC_LIO_LIduring testing or travelling, it is often desired to run on a single computer\nC_LI\n\nHowever, since each type of cluster has its own API, code that runs jobs on one type of cluster needs to be re-written if it is desired to run that application on a different type of cluster. To resolve this problem, we created a software module to generalize the submission of pipelines across computing environments, including local compute, clouds and clusters.\n\nResultsHPCI (High Performance Computing Interface) is a Perl module that provides the interface to a standardized generic cluster.\n\nWhen the HPCI module is used, it accepts a parameter to specify the cluster type. The HPCI module uses this to load a driver HPCD:: . This is used to translate the abstract HPCI interface to the specific software interface.\n\nSimply by changing the cluster parameter, the same pipeline can be run on a different type of cluster with no other changes.\n\nConclusionThe HPCI module assists in writing Perl programs that can be run in different lab environments, with different site configuration requirements and different types of hardware clusters. Rather than having to re-write portions of the program, it is only necessary to change a configuration file.\n\nUsing HPCI, an application can manage collections of jobs to be runs, specify ordering dependencies, detect success or failure of jobs run and allow automatic retry of failed jobs (allowing for the possibility of a changed configuration such as when the original attempt specified an inadequate memory allotment).

bioinformatics

Identification of potential microRNAs associated with Herpesvirus family based on bioinformatic analysis

MicroRNAs (miRNAs) are known key regulators of gene expression on posttranscriptional level in many organisms encoded in mammals, plants and also several viral families. To date, no homologous gene of a virus-originated miRNA is known in other organisms. To date, only a few homologous miRNA between two different viruses are known, however, no gene of a virus-originated miRNA is known in any organism of other kingdoms. This can be attributed to the fact that classical miRNA detection approaches such as homology-based predictions fail at viruses due to their highly diverse genomes and their high mutation rate.\n\nHere, we applied the virus-derived precursor miRNA (pre-miRNA) prediction pipeline ViMiFi, which combines information about sequence conservation and machine learning-based approaches, on Human Herpesvirus 7 (HHV7) and Epstein-Barr virus (EBV). ViMiFi was able to predict 61 candidates in EBV, which has 25 known pre-miRNAs. From these 25, ViMiFi identified 20. It was further able to predict 18 candidates in the HHV7 genome, in which no miRNA had been described yet. We also studied the undescribed candidates of both viruses for potential functions and found similarities with human snRNAs and miRNAs from mammals and plants.

bioinformatics

Single nucleotide polymorphisms of the c-MYC gene’s relationship with formation of Burkitt’s lymphoma using bioinformatics analysis

Burkitts lymphoma (BL) is an aggressive form of non-Hodgkin lymphoma, originates from germinal center B cells, MYC gene (MIM ID 190080) is an important proto-oncogene transcriptional factor encoding a nuclear phosphoprotein for central cellular processes. Dysregulated expression or function of c-MYC is one of the most common abnormalities in BL. This study focused on the investigation of the possible role of single nucleotide polymorphisms (SNPs) in MYC gene associated with formation of BL.\n\nMYC SNPs were obtained from NCBI database. SNPs in the coding region that are non-synonymous (nsSNPs) were analysed by multiple programs such as SIFT, Polyphen2, SNPs&GO, PHD-SNP and I-mutant. In this study, a total of 286 Homo sapiens SNPs were found. Roughly, forty-eight of them were deleterious and were furtherly investigated.\n\nEight SNPs were considered most disease causing [rs4645959 (N26S), rs4645959 (N25S), rs141095253 (P396L), rs141095253 (P397L), rs150308400 (C233Y), rs150308400 (C147Y), rs150308400 (C147Y), rs150308400 (C148Y)] according to the four softwares used. Two of which have not been reported previously [rs4645959 (N25S), rs141095253 (P396L)]. SNPs analysis helps is a diagnostic marker which helps in diagnosing and consequently, finding therapeutics for clinical diseases. This is through SNPs genotyping arrays and other techniques. Thus, it is highly recommended to confirm the findings in this study in vivo and in vitro.

bioinformatics

Novel bioinformatics approach to investigate quantitative phenotype-genotype associations in neuroimaging studies

Imaging genetics is an emerging field in which the association between genes and neuroimaging-based quantitative phenotypes are used to explore the functional role of genes in neuroanatomy and neurophysiology in the context of healthy function and neuropsychiatric disorders. The main obstacle for researchers in the field is the high dimensionality of the data in both the imaging phenotypes and the genetic variants commonly typed. In this article, we develop a novel method that utilizes Gene Ontology, an online database, to select and prioritize certain genes, employing a stratified false discovery rate (sFDR) approach to investigate their associations with imaging phenotypes. sFDR has the potential to increase power in genome wide association studies (GWAS), and is quickly gaining traction as a method for multiple testing correction. Our novel approach addresses both the pressing need in genetic research to move beyond candidate gene studies, while not being overburdened with a loss of power due to multiple testing. As an example of our methodology, we perform a GWAS of hippocampal volume using the Alzheimers Disease Neuroimaging Initiative sample.

Genetics

Laboratory cultivation of acidophilic nanoorganisms. Physiological and bioinformatic dissection of a stable laboratory co-culture.

This study describes the laboratory cultivation of ARMAN (Archaeal Richmond Mine Acidophilic Nanoorganisms). After 2.5 years of successive transfers in an anoxic medium containing ferric sulfate as an electron acceptor, a consortium was attained that is comprised of two members of the order Thermoplasmatales, a member of a proposed ARMAN group, as well as a fungus. The 16S rRNA of one archaeon is only 91.6% identical to Thermogymnomonas acidicola as most closely related isolate. Hence, this organism is the first member of a new genus. The enrichment culture is dominated by this microorganism and the ARMAN. The third archaeon in the community seems to be present in minor quantities and has a 100% 16S rRNA identity to the recently isolated Cuniculiplasma divulgatum. The enriched ARMAN species is most probably incapable of sugar metabolism because the key genes for sugar catabolism and anabolism could not be identified in the metagenome. Metatranscriptomic analysis suggests that the TCA cycle funneled with amino acids is the main metabolic pathway used by the archaea of the community. Microscopic analysis revealed that growth of the ARMAN is supported by the formation of cell aggregates. These might enable cross feeding by other community members to the ARMAN.

microbiology

Bioinformatic characterisation of the effector repertoire of the strawberry pathogen Phytophthora cactorum

The oomycete pathogen Phytophthora cactorum causes crown rot, a major disease of cultivated strawberry. We report the draft genome of P. cactorum isolate 10300, cultured from symptomatic Fragaria x ananassa tissue. Our analysis revealed that there are a large number of genes encoding putative secreted effectors in the genome, including nearly 200 RxLR domain containing effectors, 77 Crinklers (CRN) grouped into 38 families and numerous apoplastic effectors, such as phytotoxins (PcF proteins) and necrosis inducing proteins. As in other Phytophthora species, the genomic environment of many RxLR and CRN genes differed from core eukaryotic genes, a hallmark of the two-speed genome. We found genes homologous to known Phytophthora infestans avirulence genes including Avr1, Avr3b, Avr4, Avrblb1 and AvrSmira2 indicating effector sequence conservation between Phytophthora species of Clade 1A and 1C. The reported P. cactorum genome sequence and associated annotations represent a comprehensive resource for avirulence gene discovery in other Phytophthora species from Clade 1 and will facilitate effector informed breeding strategies in other crops.

pathology

Functional phenomics: An emerging field integrating high-throughput phenotyping, physiology, and bioinformatics

HighlightFunctional phenomics is an emerging field in plant biology that relies on high-throughput phenotyping, big data analytics, controlled manipulative experiments, and simulation modelling to increase understanding of plant physiology.\n\nAbstractThe emergence of functional phenomics signifies the rebirth of physiology as a 21st century science through the use of advanced sensing technologies and big data analytics. Functional phenomics highlights the importance of phenotyping beyond only identifying genetic regions because a significant knowledge gap remains in understanding which plant properties will influence ecosystem services beneficial to human welfare. Here, a general approach for the theory and practice of functional phenomics is outlined including exploring the phene concept as a unit of phenotype. The functional phenomics pipeline is proposed as a general method for conceptualizing, measuring, and validating utility of plant phenes. The functional phenomics pipeline begins with ideotype development. Second, a phenotyping platform is developed to maximize the throughput of phene measurements. A mapping population is screened measuring target phenes and indicators of plant performance such as yield and nutrient composition. Traditional forward genetics allows genetic mapping, while functional phenomics links phenes to plant performance. Based on these data, genotypes with contrasting phenotypes can be selected for smaller yet more intensive experiments to understand phene-environment interactions in depth. Simulation modeling can be used to understand the phenotypes and all stages of the pipeline feed back to ideotype and phenotyping platform development.

plant biology

Expression profile analysis of circular RNAs in essential hypertension by microarray and bioinformatics.

Circular RNAs (circRNAs), widely found in human cells, are involved in disease and play an important role in progression. To determine whether circRNAs are related in essential hypertension (EH), we analyzed the expression profile of circRNAs and miRNAs in 5 EH and 5 healthy controls cases which were screened by microarray. Through microarray data and public data analysis, differently expressed transcripts were divided into modules, and circRNAs were functionally annotated by miRNAs. The expression of two circRNAs, has_circ_0037909 and has_circ_0105015, were validated in EH by qRT-PCR, which may be associated with EH. Further analysis showed that two circRNAs might through immune system by up-regulation circRNAs and down-regulation expression. These circRNAs biological functions need to be further validated.

genetics

Characterization of emetic and diarrheal Bacillus cereus strains from a 2016 foodborne outbreak using whole-genome sequencing: addressing the microbiological, epidemiological, and bioinformatic challenges

The Bacillus cereus group comprises multiple species capable of causing emetic or diarrheal foodborne illness. Despite being responsible for tens of thousands of illnesses each year in the U.S. alone, whole-genome sequencing (WGS) has not been routinely employed to characterize B. cereus group isolates from foodborne outbreaks. Here, we describe the first WGS-based characterization of isolates linked to an outbreak caused by members of the B. cereus group. In conjunction with a 2016 outbreak traced to a supplier of refried beans served by a fast food restaurant chain in upstate New York, a total of 33 B. cereus group strains were obtained from human cases (n =7) and food samples (n = 26). Emetic (n = 30) and diarrheal (n = 3) isolates were most closely related to B. paranthracis (clade III) and B. cereus sensu stricto (clade IV), respectively. WGS indicated that the 30 emetic isolates (24 and 6 from food and humans, respectively) were closely-related and formed a well-supported clade relative to publicly-available emetic clade III genomes with an identical sequence type (ST 26). When compared to publicly-available emetic clade III ST 26 B. cereus group genomes, the 30 emetic clade III isolates from this outbreak differed from each other by a mean of 8.3 to 11.9 core single nucleotide polymorphisms (SNPs), while differing from publicly-available genomes by a mean of 301.7 to 528.0 core SNPs, depending on the SNP calling methodology used. Using a WST-1 cell proliferation assay, the strains isolated from this outbreak had only mild detrimental effects on HeLa cell metabolic activity compared to reference diarrheal strain B. cereus ATCC 14579. Based on both WGS and epidemiological data, we hypothesize that the outbreak was a single source outbreak caused by emetic clade III B. cereus belonging to the B. paranthracis species. In addition to showcasing how WGS can be used to characterize B. cereus group strains linked to a foodborne outbreak, we also discuss potential microbiological and epidemiological challenges presented by B. cereus group outbreaks, and we offer recommendations for analyzing WGS data from the isolates associated with them.

microbiology

Bioinformatic analysis reveals mechanisms underlying parasitoid venom evolution and function

Introduction Introduction Methods Results and Discussion Conclusions References Parasitoid wasps are a numerous and diverse group of insects that obligately infect other arthropod species. These wasps lay their eggs either on the surface or within the body cavity of their hosts, and the resulting offspring exploit the hosts resources to complete their development [1]. Many parasitoids also introduce venom gland derived proteins or polydnaviruses into the host during infection. These factors act through a variety of mechanisms to manipulate host biology in order to increase the fitness of the developing parasitoid offspring [2-4]. Many hosts mount immune responses to parasitoid infection, and accordingly, parasitoids have evolved multiple venom genes that encode immunom ...

genomics

Lessons Learned: Recommendations for Establishing Critical Periodic Scientific Benchmarking

The dependence of life scientists on software has steadily grown in recent years. For many tasks, researchers have to decide which of the available bioinformatics software are more suitable for their specific needs. Additionally researchers should be able to objectively select the software that provides the highest accuracy, the best efficiency and the highest level of reproducibility when integrated in their research projects.\n\nCritical benchmarking of bioinformatics methods, tools and web services is therefore an essential community service, as well as a critical component of reproducibility efforts. Unbiased and objective evaluations are challenging to set up and can only be effective when built and implemented around community driven efforts, as demonstrated by the many ongoing community challenges in bioinformatics that followed the success of CASP. Community challenges bring the combined benefits of intense collaboration, transparency and standard harmonization. Only open systems for the continuous evaluation of methods offer a perfect complement to community challenges, offering to larger communities of users that could extend far beyond the community of developers, a window to the developments status that they can use for their specific projects. We understand by continuous evaluation systems as those services which are always available and periodically update their data and/or metrics according to a predefined schedule keeping in mind that the performance has to be always seen in terms of each research domain.\n\nWe argue here that technology is now mature to bring community driven benchmarking efforts to a higher level that should allow effective interoperability of benchmarks across related methods. New technological developments allow overcoming the limitations of the first experiences on online benchmarking e.g. EVA. We therefore describe OpenEBench, a novel infra-structure designed to establish a continuous automated benchmarking system for bioinformatics methods, tools and web services.\n\nOpenEBench is being developed so as to cater for the needs of the bioinformatics community, especially software developers who need an objective and quantitative way to inform their decisions as well as the larger community of end-users, in their search for unbiased and up-to-date evaluation of bioinformatics methods. As such OpenEBench should soon become a central place for bioinformatics software developers, community-driven benchmarking initiatives, researchers using bioinformatics methods, and funders interested in the result of methods evaluation.

bioinformatics

BGDMdocker: an workflow base on Docker for analysis and visualization pan-genome and biosynthetic gene clusters of Bacterial

MotivationAt present Docker technology has received increasing level of attention throughout the bioinformatics community. However, its implementation details have not yet been mastered by most biologists and applied widely in biological researches. In order to popularizing this technology in the bioinformatics and sufficiently use plenty of public resources of bioinformatics tools (Dockerfile and image of scommunity, officially and privately) in Docker Hub Registry and other Docker sources based on Docker, we introduced full and accurate instance of a bioinformatics workflow based on Docker to analyse and visualize pan-genome and biosynthetic gene clusters of a bacteria in this article, provided the solutions for mining bioinformatics big data from various public biology databases. You could be guided step-by-step through the workflow process from docker file to build up your own images and run an container fast creating an workflow.\n\nResultsWe presented a BGDMdocker (bacterial genome data mining docker-based) workflow based on docker. The workflow consists of three integrated toolkits, Prokka v1.11, panX, and antiSMASH3.0. The dependencies were all written in Dockerfile, to build docker image and run container for analysing pan-genome of total 44 Bacillus amyloliquefaciens strains, which were retrieved from public? database. The pan-genome totally includes 172,432 gene, 2,306 Core gene cluster. The visualized pan-genomic data such as alignment, phylogenetic trees, maps mutations within that cluster to the branches of the tree, infers loss and gain of genes on the core-genome phylogeny for each gene cluster were presented. Besides, 997 known (MIBiG database) and 553 unknown (antiSMASH-predicted clusters and Pfam database) genes of biosynthesis gene clusters types and orthologous groups were mined in all strains. This workflow could also be used for other species pan-genome analysis and visualization. The display of visual data can completely duplicated as well as done in this paper. All result data and relevant tools and files can be downloaded from our website with no need to register. The pan-genome and biosynthetic gene clusters analysis and visualization can be fully reusable immediately in different computing platforms (Linux, Windows, Mac and deployed in the cloud), achieved cross platform deployment flexibility, rapid development integrated software package.\n\nAvailability and implementationBGDMdocker is available at http://42.96.173.25/bapgd/ and the source code under GPL license is available at https://github.com/cgwyx/debian_prokka_panx_antismash_biodocker.\n\nContactchenggongwyx@foxmail.com\n\nSupplementary informationSupplementary data are available at biorxiv online.

bioinformatics

Assessing 16S marker gene survey data analysis methods using mixtures of human stool sample DNA extracts.

BackgroundAnalysis of 16S rRNA marker-gene surveys, used to characterize prokaryotic microbial communities, may be performed by numerous bioinformatic pipelines and downstream analysis methods. However, there is limited guidance on how to decide between methods, appropriate data sets and statistics for assessing these methods are needed. We developed a mixture dataset with real data complexity and an expected value for assessing 16S rRNA bioinformatic pipelines and downstream analysis methods. We generate an assessment dataset using a two-sample titration mixture design. The sequencing data were processed using multiple bioinformatic pipelines, i) DADA2 a sequence inference method, ii) Mothur a de novo clustering method, and iii) QIIME with open-reference clustering. The mixture dataset was used to qualitatively and quantitatively assess count tables generated using the pipelines.\n\nResultsThe qualitative assessment was used to evalute features only present in unmixed samples and titrations. The abundance of Mothur and QIIME features specific to unmixed samples and titrations were explained by sampling alone. However, for DADA2 over a third of the unmixed sample and titration specific feature abundance could not be explained by sampling alone. The quantitative assessment evaluated pipeline performance by comparing observed to expected relative and differential abundance values. Overall the observed relative abundance and differential abundance values were consistent with the expected values. Though outlier features were observed across all pipelines.\n\nConclusionsUsing a novel mixture dataset and assessment methods we quantitatively and qualitatively evaluated count tables generated using three bioinformatic pipelines. The dataset and methods developed for this study will serve as a valuable community resource for assessing 16S rRNA marker-gene survey bioinformatic methods.

bioinformatics

Fasta-O-Matic: a tool to sanity check and if needed reformat FASTA files

BackgroundAs the sheer volume of bioinformatic sequence data increases, the only way to take advantage of this content is to more completely automate robust analysis workflows. Analysis bottlenecks are often mundane and overlooked processing steps. Idiosyncrasies in reading and/or writing bioinformatics file formats can halt or impair analysis workflows by interfering with the transfer of data from one informatics tools to another.\n\nResultsFasta-O-Matic automates handling of common but minor format issues that otherwise may halt pipelines. The need for automation must be balanced by the need for manual confirmation that any formatting error is actually minor rather than indicative of a corrupt data file. To that end Fasta-O-Matic reports any issues detected to the user with optionally color coded and quiet or verbose logs.\n\nFasta-O-Matic can be used as a general pre-processing tool in bioinformatics workflows (e.g. to automatically wrap FASTA files so that they can be read by BioPerl). It was also developed as a sanity check for bioinformatic core facilities that tend to repeat common analysis steps on FASTA files received from disparate sources. Fasta-O-Matic can be set with format requirements specific to downstream tools as a first step in a larger analysis workflow.\n\nAvailabilityFasta-O-Matic is available free of charge to academic and non-profit institutions on GitHub.

Bioinformatics

MathIOmica: An Integrative Platform for Dynamic Omics

Multiple omics data are rapidly becoming available, necessitating the use of new methods to integrate different technologies and interpret the results arising from multimodal assaying. The MathIOmica package for Mathematica provides one of the first extensive introductions to the use of the Wolfram Language to tackle such problems in bioinformatics. The package particularly addresses the necessity to integrate multiple omics information arising from dynamic profiling in a personalized medicine approach. It provides multiple tools to facilitate bioinformatics analysis, including importing data, annotating datasets, tracking missing values, normalizing data, clustering and visualizing the classification of data, carrying out annotation and enumeration of ontology memberships and pathway analysis. We anticipate MathIOmica to not only help in the creation of new bioinformatics tools, but also in promoting interdisciplinary investigations, particularly from researchers in mathematical, physical science and engineering fields transitioning into genomics, bioinformatics and omics data integration.

Bioinformatics