Search bioRxivSearch

Biology subjects

Hoang, Q. M.

Publications and source records attributed to Hoang, Q. M..

1 recordsLinked to original sources

HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata

We present HO_SCPLOWARVESTMANC_SCPLOW, a method that takes advantage of hierarchical relationships among the possible biological interpretations and representations of genomic variants to perform automatic feature learning, feature selection, and model building. We demonstrate that HO_SCPLOWARVESTMANC_SCPLOW scales to thousands of genomes comprising more than 84 million variants by processing phase 3 data from the 1000 Genomes Project, the largest publicly available collection of whole genome sequences. Next, using breast cancer data from The Cancer Genome Atlas, we show that HO_SCPLOWARVESTMANC_SCPLOW selects a rich combination of representations that are adapted to the learning task, and performs better than a binary representation of SNPs alone. Finally, we compare HO_SCPLOWARVESTMANC_SCPLOW to existing feature selection methods and demonstrate that our method selects smaller and less redundant feature subsets, while maintaining accuracy of the resulting classifier. The data used is available through either the 1000 Genomes Project or The Cancer Genome Atlas. Access to TCGA data requires the completion of a Data Access Request through the Database of Genotypes and Phenotypes (dbGaP). Binary releases of HO_SCPLOWARVESTMANC_SCPLOW compatible with Linux, Windows, and Mac are available for download at https://github.com/cmlh-gp/Harvestman-public/releases

bioinformatics