Search bioRxivSearch

Biology subjects

Andreu-Sanchez, S.

Publications and source records attributed to Andreu-Sanchez, S..

5 recordsLinked to original sources

A comparison of tools for copy-number variation detection in germline whole exome and whole genome sequencing data

BackgroundCopy-number variations (CNVs) have important clinical implications for several diseases and cancers. The clinically relevant CNVs are hard to detect because CNVs are common structural variations that define large parts of the normal human genome. CNV calling from short-read sequencing data has the potential to leverage available cohort studies and allow full genomic profiling in the clinic without the need for additional data modalities. Questions regarding performance of CNV calling tools for clinical use and suitable sequencing protocols remain poorly addressed, mainly because of the lack of good reference data sets. MethodsWe reviewed 50 popular CNV calling tools and included 11 tools for benchmarking in a unique reference cohort encompassing 39 whole genome sequencing (WGS) samples paired with analysis by the current clinical standard --SNP-array based CNV calling. Additionally, for nine of these samples we performed whole exome sequencing (WES) performed, in order to address the effect of sequencing protocol on CNV calling. Furthermore, we included Gold Standard reference sample NA12878, and tested 12 samples with CNVs confirmed by multiplex ligation-dependent probe amplification (MLPA). ResultsTool performance varied greatly in the number of called CNVs and bias for CNV lengths. Some tools had near-perfect recall of CNVs from arrays for some samples, but poor precision. Filtering output by CNV ranks from tools did not salvage precision. Several tools had better performance patterns for NA12878, and we hypothesize that this is the result of overfitting during the tool development. ConclusionsWe suggest combining tools with the best recall: GATK gCNV, Lumpy, DELLY, and cn.MOPS. These tools also capture different CNVs. Further improvements in precision requires additional development of tools, reference data sets, and annotation of CNVs, potentially assisted by the use of background panels for filtering of frequently called variants.

bioinformatics

Gut microbial structural variations as determinants of human bile acid metabolism

Bile acids (BAs) facilitate intestinal fat absorption and act as important signaling molecules in host{square}gut microbiota crosstalk. BA-metabolizing pathways in the microbial community have been identified, but how the highly variable genomes of gut bacteria interact with host BA metabolism remains largely unknown. We characterized 8,282 structural variants (SVs) of 55 bacterial species in the gut microbiomes of 1,437 individuals from two Dutch cohorts and performed a systematic association study with 39 plasma BA parameters. Both variations in SV-based continuous genetic makeup and discrete subspecies showed correlations with BA metabolism. Metagenome-wide association analysis identified 797 replicable associations between bacterial SVs and BAs and SV regulators that mediate the effects of lifestyle factors on BA metabolism. This is the first large-scale microbial genetic association analysis to demonstrate the impact of bacterial SVs on human BA composition, and highlights the potential of targeting gut microbiota to regulate BA metabolism through lifestyle intervention.

microbiology

Easymap: a user-friendly software package for rapid mapping by sequencing of point mutations and large insertions

Mapping-by-sequencing strategies combine next-generation sequencing (NGS) with classical linkage analysis, allowing rapid identification of the causal mutations of the phenotypes exhibited by mutants isolated in a genetic screen. Computer programs that analyze NGS data obtained from a mapping population of individuals derived from a mutant of interest in order to identify a causal mutation are available; however, the installation and usage of such programs requires bioinformatic skills, modifying or combining pieces of existing software, or purchasing licenses. To ease this process, we developed Easymap, an open-source program that simplifies the data analysis workflows from raw NGS reads to candidate mutations. Easymap can perform bulked segregant mapping of point mutations induced by ethyl methanesulfonate (EMS) with DNA-seq or RNA-seq datasets, as well as tagged-sequence mapping for large insertions, such as transposons or T-DNAs. The mapping analyses implemented in Easymap have been validated with experimental and simulated datasets from different plant and animal model species. Easymap was designed to be accessible to all users regardless of their bioinformatics skills by implementing a user-friendly graphical interface, a simple universal installation script, and detailed mapping reports, including informative images and complementary data for assessment of the mapping results. One sentence summaryEasymap is a versatile user-friendly software tool that facilitates mapping-by-sequencing of large insertions and point mutations in plant and animal genomes

plant biology

Effect of host genetics on the gut microbiome in 7,738 participants of the Dutch Microbiome Project

Host genetics are known to influence the gut microbiome, yet their role remains poorly understood. To robustly characterize these effects, we performed a genome-wide association study of 207 taxa and 205 pathways representing microbial composition and function within the Dutch Microbiome Project, a population cohort of 7,738 individuals from the northern Netherlands. Two robust, study-wide significant (p<1.89x10-10) signals near the LCT and ABO genes were found to affect multiple microbial taxa and pathways, and were replicated in two independent cohorts. The LCT locus associations were modulated by lactose intake, while those at ABO reflected participant secretor status determined by FUT2 genotype. Eighteen other loci showed suggestive evidence (p<5x10-8) of association with microbial taxa and pathways. At a more lenient threshold, the number of loci identified strongly correlated with trait heritability, suggesting that much larger sample sizes are needed to elucidate the remaining effects of host genetics on the gut microbiome.

genetics

The Dutch Microbiome Project defines factors that shape the healthy gut microbiome

The gut microbiome is associated with diverse diseases, but the universal signature of an (un)healthy microbiome remains elusive and there is a need to understand how genetics, exposome, lifestyle and diet shape the microbiome in health and disease. To fill this gap, we profiled bacterial composition, function, antibiotic resistance and virulence factors in the gut microbiomes of 8,208 Dutch individuals from a three-generational cohort comprising 2,756 families. We then correlated this to 241 host and environmental factors, including physical and mental health, medication use, diet, socioeconomic factors and childhood and current exposome. We identify that the microbiome is primarily shaped by environment and cohousing. Only [~]13% of taxa are heritable, which are enriched with highly prevalent and health-associated bacteria. By identifying 2,856 associations between microbiome and health, we find that seemingly unrelated diseases share a common signature that is independent of comorbidities. Furthermore, we identify 7,519 associations between microbiome features and diet, socioeconomics and early life and current exposome, of which numerous early-life and current factors are particularly linked to the microbiome. Overall, this study provides a comprehensive overview of gut microbiome and the underlying impact of heritability and exposures that will facilitate future development of microbiome-targeted therapies.

microbiology