Search bioRxiv⌕ Search

Biology subjects

Stojanovska, F.

Publications and source records attributed to Stojanovska, F..

2 recordsLinked to original sources

Contextualising transcription factor binding during embryogenesis using natural sequence variation

Understanding how genetic variation impacts transcription factor (TF) binding remains a major challenge, limiting our ability to model disease-associated variants. Here, we used a highly controlled system of F1 crosses with extensive genetic diversity to profile allele-specific binding of four TFs at several embryonic time-points, using Drosophila as a model. Using a combined haplotype test, we identified 9-18% of TF bound regions impacted by genetic variation. By expanding WASP (a tool for allele-specific read mapping) to examine INDELs, we increased detection of allele imbalanced (AI) peaks by 30-50%. This fine-grained mutagenesis could reconstruct functionalized binding motifs of all factors. To prioritise potential causal variants, we trained a convolutional neural network (Basenji) to predict TF binding from DNA sequence. The model could accurately predict experimental AI for strong effect variants, providing a mechanistic interpretation for how genetic variation impacted TF binding. This revealed unexpected relationships between TFs, including potential cooperative pairs, and mechanisms of tissue specific recruitment of the ubiquitous factor CTCF.

genetics↗

Convolutional networks for supervised mining of molecular patterns within cellular context

Cryo-electron tomograms capture a wealth of structural information on the molecular constituents of cells and tissues. We present DeePiCt (Deep Picker in Context), an open-source deep-learning framework for supervised structure segmentation and macromolecular complex localization in cellular cryo-electron tomography. To train and benchmark DeePiCt on experimental data, we comprehensively annotated 20 tomograms of Schizosaccharomyces pombe for ribosomes, fatty acid synthases, membranes, nuclear pore complexes, organelles and cytosol. By comparing our method to state-of-the-art approaches on this dataset, we show its unique ability to identify low-abundance and low-density complexes. We use DeePiCt to study compositionally-distinct subpopulations of cellular ribosomes, with emphasis on their contextual association with mitochondria and the endoplasmic reticulum. Finally, by applying pre-trained networks to a HeLa cell dataset, we demonstrate that DeePiCt achieves high-quality predictions in unseen datasets from different biological species in a matter of minutes. The comprehensively annotated experimental data and pre-trained networks are provided for immediate exploitation by the community.

bioinformatics↗