Search bioRxivSearch

Biology subjects

Michael Snyder

Publications and source records attributed to Michael Snyder.

3 recordsLinked to original sources

Interactive Analytics for Very Large Scale Genomic Data

Large scale genomic sequencing is now widely used to decipher questions in diverse realms such as biological function, human diseases, evolution, ecosystems, and agriculture. With the quantity and diversity these data harbor, a robust and scalable data handling and analysis solution is desired. Here we present interactive analytics using public cloud infrastructure and distributed computing database Dremel and developed according to the standards of Global Alliance for Genomics and Health, to perform information compression, comprehensive quality controls, and biological information retrieval in large volumes of genomic data. We demonstrate that such computing paradigms can provide orders of magnitude faster turnaround for common analyses, transforming long-running batch jobs submitted via a Linux shell into questions that can be asked from a web browser in seconds.

Preprint

Practical Guidelines for Secure Cloud Computing using Genomic Data

Cloud security challenges Cloud security challenges Cloud security guidelines Security and Privacy Publications: Large scale genomics studies involving thousands of whole genome or exome sequences are underway1 on Cloud. While Cloud provides many conveniences for genomics research, it also raises concerns regarding large scale hacking, bad press, potential loss of patient privacy and the resulting loss of patient trust. Cloud providers argue that they have significant investments and expertise in security and, therefore, Cloud is equally secure, if not more so, compared to on-premise infrastructure. This gap in assessment of Cloud security is, in part, due to a fast evolving and largely unfamiliar technology stack for genomics data owners.\n\nWhat makes the Cloud security landscape discussion challenging is that security reco ...

Bioinformatics

Distance from Sub-Saharan Africa Predicts Mutational Load in Diverse Human Genomes

The Out-of-Africa (OOA) dispersal ~50,000 years ago is characterized by a series of founder events as modern humans expanded into multiple continents. Population genetics theory predicts an increase of mutational load in populations undergoing serial founder effects during range expansions. To test this hypothesis, we have sequenced full genomes and high-coverage exomes from 7 geographically divergent human populations from Namibia, Congo, Algeria, Pakistan, Cambodia, Siberia and Mexico. We find that individual genomes vary modestly in the overall number of predicted deleterious alleles. We show via spatially explicit simulations that the observed distribution of deleterious allele frequencies is consistent with the OOA dispersal, particularly under a model where deleterious mutations are recessive. We conclude that there is a strong signal of purifying selection at conserved genomic positions within Africa, but that many predicted deleterious mutations have evolved as if they were neutral during the expansion out of Africa. Under a model where selection is inversely related to dominance, we show that OOA populations are likely to have a higher mutation load due to increased allele frequencies of nearly neutral variants that are recessive or partially recessive.

Genomics