Search bioRxiv⌕ Search

Biology subjects

Durrant, M. G.

Publications and source records attributed to Durrant, M. G..

4 recordsLinked to original sources

Large-scale discovery of recombinases for integrating DNA into the human genome

Recent microbial genome sequencing efforts have revealed a vast reservoir of mobile genetic elements containing integrases that could be useful genome engineering tools. Large serine recombinases (LSRs), such as Bxb1 and PhiC31, are bacteriophage-encoded integrases that can facilitate the insertion of phage DNA into bacterial genomes. However, only a few LSRs have been previously characterized and they have limited efficiency in human cells. Here, we developed a systematic computational discovery workflow that identifies thousands of new LSRs and their cognate DNA attachment sites by. We validate this approach via experimental characterization of LSRs in human cells, leading to three classes of LSRs distinguished from one another by their efficiency and specificity. We identify landing pad LSRs that efficiently integrate into synthetically installed attachment sites orthogonal to the human genome, human genome-targeting LSRs with computationally predictable pseudosites, and multi-targeting LSRs that can unidirectionally integrate cargos at with similar efficiency and superior specificity to commonly used transposases. LSRs from each category were functionally characterized in human cells, overall achieving up to 7-fold higher plasmid recombination than Bxb1 and genome insertion efficiencies of 40-70% with cargo sizes over 7 kb. Overall, we establish a paradigm for large-scale discovery of microbial recombinases and reconstruction of their target sites directly from microbial sequencing data. This strategy provides a rich resource of over 60 experimentally characterized LSRs that can function in human cells and thousands of additional candidates for large-payload genome editing without exposed DNA double-stranded breaks.

bioengineering↗

Chromatin accessibility changes induced by the microbial metabolite butyrate reveal possible mechanisms of anti-cancer effects

Butyrate is a four-carbon fatty acid produced in large quantities by bacteria found in the human gut. It is the major source of colonic epithelial cell energy, can bind to and agonize short-chain fatty acid G-protein coupled receptors and functions as a histone deacetylase (HDAC) inhibitor. Anti-cancer effects of butyrate are attributed to a global increase in histone acetylation in colon cancer cells; however, the role that corresponding chromatin remodeling plays in this effect is not fully understood. We used longitudinal paired ATAC-seq and RNA-seq on HCT-116 colon cancer cells to determine how butyrate-related chromatin changes functionally associate with cancer. We detected distinct temporal changes in chromatin accessibility in response to butyrate with less accessible regions enriched in transcription factor binding motifs and distal enhancers. These regions significantly overlapped with regions maintained by the SWI/SNF chromatin remodeler, and were further enriched amongst chromatin regions that are associated with ARID1A/B synthetic lethality. Finally, we found that butyrate-induced chromatin regions were enriched for both colorectal cancer GWAS loci and somatic mutations in cancer. These results demonstrate the convergence of both somatic mutations and GWAS risk variants for colon cancer within butyrate-responsive chromatin regions, providing a molecular map of the mechanisms by which this microbial metabolite might confer anti-cancer properties. HighlightsO_LIChromatin accessibility changes longitudinally upon butyrate exposure in colon cancer cells. C_LIO_LIChromatin regions that close in response to butyrate are enriched among distal enhancers. C_LIO_LIThere is strong overlap between butyrate-induced peaks and peaks associated with SWI/SNF synthetic lethality. C_LIO_LIButyrate-induced peaks are enriched for colorectal cancer GWAS loci and somatic variation in colorectal cancer. C_LI

genomics↗

An integrated approach to identify environmental modulators of genetic risk factors for complex traits

Complex traits and diseases can be influenced by both genetics and environment. However, given the large number of environmental stimuli and power challenges for gene-by-environment testing, it remains a critical challenge to identify and prioritize specific disease-relevant environmental exposures. We propose a novel framework for leveraging signals from transcriptional responses to environmental perturbations to identify disease-relevant perturbations that can modulate genetic risk for complex traits and inform the functions of genetic variants associated with complex traits. We perturbed human skeletal muscle, fat, and liver relevant cell lines with 21 perturbations affecting insulin resistance, glucose homeostasis, and metabolic regulation in humans and identified thousands of environmentally responsive genes. By combining these data with GWAS from 31 distinct polygenic traits, we show that heritability of multiple traits is enriched in regions surrounding genes responsive to specific perturbations and, further, that environmentally responsive genes are enriched for associations with specific diseases and phenotypes from the GWAS catalogue. Overall, we demonstrate the advantages of large-scale characterization of transcriptional changes in diversely stimulated and pathologically relevant cells to identify disease-relevant perturbations.

genomics↗

Automated prediction and annotation of small proteins in microbial genomes

Recent work performed by Sberro et al. (2019) revealed a vast unexplored space of small proteins existing within the human microbiome. At present, these small open reading frames (smORFs) are unannotated in existing reference genomes and standard genome annotation tools are not able to accurately predict them. In this study, we introduce an annotation tool named SmORFinder that predicts small proteins based on those identified by Sberro et al. This tool combines profile Hidden Markov models (pHMMs) of each small protein family and deep learning models that may better generalize to smORF families not seen in the training set. We find that combining predictions of both pHMM and deep learning models leads to more precise smORF predictions and that these predicted smORFs are enriched for Ribo-Seq or MetaRibo-Seq translation signals. Feature importance analysis reveals that the deep learning models learned to identify Shine-Dalgarno sequences, deprioritize the wobble position in each codon, and group codons in a way that strongly corresponds to the codon synonyms found in the codon table. We perform a core genome analysis of 26 bacterial species and identify many core smORFs of unknown function. We pre-compute small protein annotations for thousands of RefSeq isolate genomes and HMP metagenomes, and we make these data available through a web portal along with other useful tools for small protein annotation and analysis. The systematic identification and annotation of those important small proteins will help researchers to expand our understanding of this exciting field of biology.

bioinformatics↗