Search bioRxiv⌕ Search

Biology subjects

Marroni, F.

Publications and source records attributed to Marroni, F..

2 recordsLinked to original sources

Power Calculator for Detecting Allelic Imbalance Using Hierarchical Bayesian Model

Allelic imbalance (AI) is the differential expression of the two alleles in a diploid. AI can vary between tissues, treatments, and environments. Statistical methods for testing in this area exist, with impacts of explosive type I error in the presence of bias well understood. However, for study design, the more important and understudied problem is the type II error and power. As the biological questions for this type of study explode, and the costs of the technology plummet, what is more important: reads or replicates? How small of an interaction can be detected while keeping the type I error at bay? Here we present a simulation study that demonstrates that the proper model can control type I error below 5% for most scenarios. We find that a minimum of 2400, 480, and 240 allele specific reads divided equally among 12, 5, and 3 replicates is needed to detect a 10%, 20%, and 30%, respectively, deviation from allelic balance in a condition with power >80%. A minimum of 960 and 240 allele specific reads is needed to detect a 20% or 30% difference in AI between conditions with comparable power but these reads need to be divided amongst 8 replicates. Higher numbers of replicates increase power more than adding coverage without affecting type I error. We provide a Python package that enables simulation of AI scenarios and enables individuals to estimate type I error and power in detecting AI and differences in AI between conditions tailored to their own specific study needs.

bioinformatics↗

Testcrosses are an efficient strategy for identifying cis regulatory variation: Bayesian analysis of allele specific expression (BASE)

Allelic imbalance (AI) occurs when alleles in a diploid individual are differentially expressed and indicates cis acting regulatory variation. What is the distribution of allelic effects in a natural population? Are all alleles the same? Are all alleles distinct? Tests of allelic effect are performed by crossing individuals and comparing expression between alleles directly in the F1. However, a crossing scheme that compares alleles pairwise is a prohibitive cost for more than a handful of alleles as the number of crosses is at least (n2-n)/2 where n is the number of alleles. We show here that a testcross design followed by a hypothesis test of AI between testcrosses can be used to infer differences between non-tester alleles, allowing n alleles to be compared with n crosses. Using a mouse dataset where both testcrosses and direct comparisons have been performed, we show that [~]75% of the predicted differences between non-tester alleles are validated in a background of [~]10% differences in AI. The testing for AI involves several complex bioinformatics steps. BASE is a complete bioinformatics pipeline that incorporates state-of-the-art error reduction techniques and a flexible Bayesian approach to estimating AI and formally comparing levels of AI between conditions. The modular structure of BASE has been packaged in Galaxy, made available in Nextflow and sbatch. (https://github.com/McIntyre-Lab/BASE_2020). In the mouse data, the direct test identifies more cis effects than the testcross. Cis-by-trans interactions with trans-acting factors on the X contributing to observed cis effects in autosomal genes in the direct cross remains a possible explanation for the discrepancy.

bioinformatics↗