Search bioRxivSearch

Biology subjects

Diallo, A. B.

Publications and source records attributed to Diallo, A. B..

3 recordsLinked to original sources

CASTOR: A machine learning platform for reproducible viral genome classification

MotivationAdvances in cloning and sequencing technology yielded a massive number of genome of virus strains. The classification and annotation of these genomes constitute important assets in the discovery of genomic variability, taxonomic characteristics and disease mechanisms. Existing classification methods are often designed for a well-studied virus. Thus, the viral comparative genomic studies could benefit from more generic, fast and accurate tools for classifying and typing newly sequenced strains of diverse virus families.\n\nResultsHere, we introduce a fast, accurate and generic virus classification platform, CASTOR, based on a machine learning approach. CASTOR is inspired by a well-known technique in molecular biology: Restriction Fragment Length Polymorphism (RFLP). It simulates the restriction digestion of genomic material by different enzymes into fragments in-silico. It uses two metrics to construct feature vectors for machine learning algorithms in the classification step. We benchmark CASTOR for the classification of distinct datasets of Human Papillomaviruses (HPV), Hepatitis B Viruses (HBV) and Human Immunodeficiency viruses (HIV). Results reveal true positive rates of 99%, 99% and 98% for HPV Alpha species, HBV genotyping and HIV M group subtyping respectively. Furthermore, CASTOR shows a competitive performance compare to well-known HIV-specific classifier REGA and COMET on whole genome and pol fragments. With such prediction rates, genericity and robustness, as well as rapidity, such approach could constitute a reference in large-scale virus studies. Finally, we developed the CASTOR web platform for open access and reproducible viral machine learning classifiers.\n\nAvailabilityhttp://castor.bioinfo.uqam.ca\n\nContactdiallo.abdoulaye@uqam.ca

bioinformatics

Towards an ontology-based recommender system for relevant bioinformatics workflows

BackgroundWith the large and diverse type of biological data, bioinformatic solutions are being more complex and computationally intensive. New specialized data skills need to be acquired by researchers in order to follow this development. Workflow Management Systems rise as an efficient way to automate tasks through abstract models in order to assist users during their problem solving tasks. However, current solutions could have several problems in reusing the developed models for given tasks. The large amount of heterogenous data and the lack of knowledge in using bioinformatics tools could mislead the users during their analyses. To tackle this issue, we propose an ontology-based workflow-mining framework generating semantic models of bioinformatic best practices in order to assist scientists. To this end, concrete workflows are extracted from scientific articles and then mined using a rich domain ontology.\n\nResultsIn this study, we explore the specific topics of phylogenetic analyses. We annotated more than 300 recent articles using different ontological concepts and relations. Relative supports (frequencies) of discovered workflow components in texts show interesting results of relevant resources currently used in the different phylogenetic analysis steps. Mining concrete workflows from texts lead us to discover abstract but relevant patterns of the best combinations of tools, parameters and input data for specific phylogenetic problems.\n\nConclusionsExtracted patterns would make workflows more intuitive and easy to be reused in similar situations. This could provide a stepping-stone into the identification of best practices and pave the road to a recommender system.

bioinformatics

Ontology-based workflow extraction from texts using word sense disambiguation

This paper introduces a method for automatic workflow extraction from texts using Process-Oriented Case-Based Reasoning (POCBR). While the current workflow management systems implement mostly different complicated graphical tasks based on advanced distributed solutions (e.g. cloud computing and grid computation), workflow knowledge acquisition from texts using case-based reasoning represents more expressive and semantic cases representations. We propose in this context, an ontology-based workflow extraction framework to acquire processual knowledge from texts. Our methodology extends classic NLP techniques to extract and disambiguate tasks in texts. Using a graph-based representation of workflows and a domain ontology, our extraction process uses a context-based approach to recognize workflow components : data and control flows. We applied our framework in a technical domain in bioinformatics : i.e. phylogenetic analyses. An evaluation based on workflow semantic similarities on a gold standard proves that our approach provides promising results in the process extraction domain. Both data and implementation of our framework are available in : http://labo.bioinfo.uqam.ca/tgrowler.

bioinformatics