Search bioRxivSearch

Biology subjects

Carmi, S.

Publications and source records attributed to Carmi, S..

5 recordsLinked to original sources

FactorialHMM: Fast and exact inference in factorial hidden Markov models

MotivationHidden Markov models (HMMs) are powerful tools for modeling processes along the genome. In a standard genomic HMM, observations are drawn, at each genomic position, from a distribution whose parameters depend on a hidden state; the hidden states evolve along the genome as a Markov chain. Often, the hidden state is the Cartesian product of multiple processes, each evolving independently along the genome. Inference in these so-called Factorial HMMs has a naive running time that scales as the square of the number of possible states, which by itself increases exponentially with the number of subchains; such a running time scaling is impractical for many applications. While faster algorithms exist, there is no available implementation suitable for developing bioinformatics applications.\n\nResultsWe developed FactorialHMM, a Python package for fast exact inference in Factorial HMMs. Our package allows simulating either directly from the model or from the posterior distribution of states given the observations. Additionally, we allow the inference of all key quantities related to HMMs: (1) the (Viterbi) sequence of states with the highest posterior probability; (2) the likelihood of the data; and (3) the posterior probability (given all observations) of the marginal and pairwise state probabilities. The running time and space requirement of all procedures is linearithmic in the number of possible states. Our package is highly modular, providing the user with maximal flexibility for developing downstream applications.\n\nAvailabilityhttps://github.com/regevs/factorialhmm

bioinformatics

Re-identification of genomic data using long range familial searches

Consumer genomics databases reached the scale of millions of individuals. Recently, law enforcement investigators have started to exploit some of these databases to find distant familial relatives, which can lead to a complete re-identification. Here, we leveraged genomic data of 600,000 individuals tested with consumer genomics to investigate the power of such long-range familial searches. We project that half of the searches with European-descent individuals will result with a third cousin or closer match and will provide a search space small enough to permit re-identification using common demographic identifiers. Moreover, in the near future, virtually any European-descent US person could be implicated by this technique. We propose a potential mitigation strategy based on cryptographic signature that can resolve the issue and discuss policy implications to human subject research.

genomics

A study of Kibbutzim in Israel reveals risk factors for cardiometabolictraits and subtle population structure

Genetic studies in isolated populations have provided increased power for identifying loci associated with complex diseases and traits. We present here the Kibbutzim Family Study (KFS), initiated for investigating environmental and genetic determinants of cardiometabolic traits in extended Israeli families living in communes characterized by long-term social stability and homogeneous environment. Extensive information on cardiometabolic traits, as well as genome-wide genetic data, was collected on 901 individuals, making this study, to the best of our knowledge, the largest of its kind in Israel. We have thoroughly characterized the KFS genetic structure, observing that most participants were of Ashkenazi Jewish (AJ) origin, and confirming a recent severe bottleneck in their recent history (point estimates: effective size {approx}450 individuals, 23 generations ago). Focusing on genetic variants enriched in KFS compared with non-Finnish Europeans, we demonstrated that AJ-specific variants are largely involved in cancer-related pathways. Using linear mixed models, we conducted an association study of these enriched variants with 16 cardiometabolic traits. We found 24 variants to be significantly associated with cardiometabolic traits. The strongest association, which we also replicated, was between a variant upstream of the MSRA gene, {approx}200-fold enriched in KFS, and weight (P=3.6{middle dot}10-8). In summary, the KFS is a valuable resource for the study of the population genetics of Israel as well as the genetics of cardiometabolic traits in a homogeneous environment.

genetics

High-depth whole genome sequencing of a large population-specific reference panel: Enhancing sensitivity, accuracy, and imputation

BackgroundWhile increasingly large reference panels for genome-wide imputation have been recently made available, the degree to which imputation accuracy can be enhanced by population-specific reference panels remains an open question. In the present study, we sequenced at full-depth ([&ge;]30x) a moderately large (n=738) cohort of samples drawn from the Ashkenazi Jewish population across two platforms (Illumina X Ten and Complete Genomics, Inc.). We developed and refined a series of quality control steps to optimize sensitivity, specificity, and comprehensiveness of variant calls in the reference panel, and then tested the accuracy of imputation against target cohorts drawn from the same population.\n\nResultsFor samples sequenced on the Illumina X Ten platform, quality thresholds were identified that permitted highly accurate calling of single nucleotide variants across 94% of the genome. The Complete Genomics, Inc. platform was more conservative (fewer variants called) compared to the Illumina platform, but also demonstrated relatively greater numbers of false positives that needed to be filtered. Quality control procedures also permitted detection of novel genome reads that are not mapped to current reference or alternate assemblies. After stringent quality control, the population-specific reference panel produced more accurate and comprehensive imputation results relative to publicly available, large cosmopolitan reference panels. The population-specific reference panel also permitted enhanced filtering of clinically irrelevant variants from personal genomes.\n\nConclusionsOur primary results demonstrate enhanced accuracy of a population-specific imputation panel relative to cosmopolitan panels, especially in the range of infrequent (<5% non-reference allele frequency) and rare (<1% non-reference allele frequency) variants that may be most critical to further progress in mapping of complex phenotypes.

genomics

Environmental factors dominate over host genetics in shaping human gut microbiota composition

Human gut microbiome composition is shaped by multiple host intrinsic and extrinsic factors, but the relative contribution of host genetic compared to environmental factors remains elusive. Here, we genotyped a cohort of 696 healthy individuals from several distinct ancestral origins and a relatively common environment, and demonstrate that there is no statistically significant association between microbiome composition and ethnicity, single nucleotide polymorphisms (SNPs), or overall genetic similarity, and that only 5 of 211 (2.4%) previously reported microbiome-SNP associations replicate in our cohort. In contrast, we find similarities in the microbiome composition of genetically unrelated individuals who share a household. We define the term biome-explainability as the variance of a host phenotype explained by the microbiome after accounting for the contribution of human genetics. Consistent with our finding that microbiome and host genetics are largely independent, we find significant biome-explainability levels of 16-33% for body mass index (BMI), fasting glucose, high-density lipoprotein (HDL) cholesterol, waist circumference, waist-hip ratio (WHR), and lactose consumption. We further show that several human phenotypes can be predicted substantially more accurately when adding microbiome data to host genetics data, and that the contribution of both data sources to prediction accuracy is largely additive. Overall, our results suggest that human microbiome composition is dominated by environmental factors rather than by host genetics.

genetics