Search bioRxivSearch

Biology subjects

Blanchette, M.

Publications and source records attributed to Blanchette, M..

5 recordsLinked to original sources

Estimating DNA-DNA interaction frequency from Hi-C data at restriction-fragment resolution

Hi-C is a popular technique to map three-dimensional chromosome conformation by capturing the frequency of physical contacts between pairs of genomic regions in cell populations. Although the resolution of Hi-C data is in principle only limited by the size of restriction fragments (300 bp - 4 kb), stochastic noise caused by the limited sequencing coverage forces researchers to artificially reduce the resolution of Hi-C matrices by binning the genome into 5-100 kb regions, resulting in a loss of information and biological interpretability. Here, we present the Hi-C Interaction Frequency Inference (HIFI) algorithms, a family of computational approaches that takes advantage of dependencies between neighboring restriction fragments to estimate restriction-fragment resolution interaction frequency matrices from Hi-C data. HIFI is shown to be superior to existing fixed-binning and state-of-the-art approaches via cross-validation experiments on Hi-C data and comparisons to 5C data. It also greatly improves the delineation of enhancer-promoter contacts. Finally, the high resolution afforded by HIFI reveals a new role for active regulatory regions in structuring topologically associating domains (TADs) and subTADs. By operating upstream of many Hi-C data analysis tools (e.g., normalization tools, as well as loop, TAD, and compartment predictors), HIFI will be easily inserted into a number of Hi-C data analysis pipelines, enabling a variety of high-resolution genomic organization analyses.\n\nAvailabilitygithub.com/BlanchetteLab/HIFI

bioinformatics

RobusTAD: A Tool for Robust Annotation ofTopologically Associating Domain Boundaries

MotivationTopologically Associating Domains (TADs) are chromatin structures that can be identified by analysis of Hi-C data. Tools currently available for TAD identification are sensitive to experimental conditions such as coverage, resolution and noise level.\n\nResultsHere, we present RobusTAD, a tool to score TAD boundaries in a manner that is robust to these parameters. In doing so, RobusTAD eases comparative analysis of TAD structures across multiple heterogeneous samples.\n\nAvailabilityRobusTAD is implemented in R and released under a GPL license. RobusTAD can be downloaded from https://github.com/rdali/RobusTAD and runs on any standard desktop computer.\n\nContactrola.dali@mail.mcgill.ca, blanchem@cs.mcgill.ca\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Structural variation detection by proximity ligation from FFPE tumor tissue

The clinical management and therapy of many solid tumor malignancies is dependent on detection of medically actionable or diagnostically relevant genetic variation. However, a principal challenge for genetic assays from tumors is the fragmented and chemically damaged state of DNA in formalin-fixed paraffin-embedded (FFPE) samples. From highly fragmented DNA and RNA there is no current technology for generating long-range DNA sequence data as is required to detect genomic structural variation or long-range genotype phasing. We have developed a high-throughput chromosome conformation capture approach for FFPE samples that we call \"Fix-C\", which is similar in concept to Hi-C. Fix-C enables structural variation detection from fresh and archival FFPE samples. We applied this method to 15 clinical adenocarcinoma and sarcoma specimens spanning a broad range of tumor purities. In this panel, Fix-C analysis achieves a 90% concordance rate with FISH assays - the current clinical gold standard. Additionally, we are able to identify novel structural variation undetected by other methods and recover long-range chromatin configuration information from these FFPE samples harboring highly degraded DNA. This powerful approach will enable detailed resolution of global genome rearrangement events during cancer progression from FFPE material, and inform the development of targeted molecular diagnostic assays for patient care.

genomics

Efficient Homology Directed Repair by Cas9:DNA Localization and Cationic Polymeric Transfection in Mammalian Cells

Homology directed repair (HDR) induced by site specific DNA double strand breaks (DSB) with CRISPR/Cas9 is a precision gene editing approach that occurs at low frequency in comparison to indel forming non homologous end joining (NHEJ). In order to obtain high HDR percentages in mammalian cells, we engineered Cas9 protein fused to a high-affinity monoavidin domain to deliver biotinylated donor DNA to a DSB site. In addition, we used the cationic polymer, polyethylenimine, to deliver Cas9 RNP-donor DNA complex into the cell. Combining these strategies improved HDR percentages of up to 90% in three tested loci (CXCR4, EMX1, and TLR) in standard HEK293 cells. Our approach offers a cost effective, simple and broadly applicable gene editing method, thereby expanding the CRISPR/Cas9 genome editing toolbox.\n\nSummaryPrecision gene editing occurs at a low percentage in mammalian cells using Cas9. Colocalization of donor with Cas9MAV and PEI delivery raises HDR occurrence.

biochemistry

An analytic approach for interpretable predictive models in high dimensional data, in the presence of interactions with exposures

Predicting a phenotype and understanding which variables improve that prediction are two very challenging and overlapping problems in analysis of high-dimensional data such as those arising from genomic and brain imaging studies. It is often believed that the number of truly important predictors is small relative to the total number of variables, making computational approaches to variable selection and dimension reduction extremely important. To reduce dimensionality, commonly-used two-step methods first cluster the data in some way, and build models using cluster summaries to predict the phenotype.\n\nIt is known that important exposure variables can alter correlation patterns between clusters of high-dimensional variables, i.e., alter network properties of the variables. However, it is not well understood whether such altered clustering is informative in prediction. Here, assuming there is a binary exposure with such network-altering effects, we explore whether use of exposure-dependent clustering relationships in dimension reduction can improve predictive modelling in a two-step framework. Hence, we propose a modelling framework called ECLUST to test this hypothesis, and evaluate its performance through extensive simulations.\n\nWith ECLUST, we found improved prediction and variable selection performance compared to methods that do not consider the environment in the clustering step, or to methods that use the original data as features. We further illustrate this modelling framework through the analysis of three data sets from very different fields, each with high dimensional data, a binary exposure, and a phenotype of interest. Our method is available in the eclust CRAN package.

genomics