bioRxiv · 10.1101/702902
VariantSpark, A Random Forest Machine Learning Implementation for Ultra High Dimensional Data
Abstract
The demands on machine learning methods to cater for ultra high dimensional datasets, datasets with millions of features, have been increasing in domains like life sciences and the Internet of Things (IoT). While Random Forests are suitable for \"wide\" datasets, current implementations such as Googles PLANET lack the ability to scale to such dimensions. Recent improvements by Yggdrasil begin to address these limitations but do not extend to Random Forest. This paper introduces CursedForest, a novel Random Forest implementation on top of Apache Spark and part of the VariantSpark platform, which parallelises processing of all nodes over the entire forest. CursedForest is 9 and up to 89 times faster than Googles PLANET and Yggdrasil, respectively, and is the first method capable of scaling to millions of features.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Bayat, A., Szul, P., O'Brien, A., Dunne, R., Luo, O., Jain, Y., Hosking, B., Bauer, D.. 2019-07-15. VariantSpark, A Random Forest Machine Learning Implementation for Ultra High Dimensional Data. https://doi.org/10.1101/702902
Cite the original work for its findings. Save a collection to share your selection of sources.