Search bioRxiv⌕ Search

Biology subjects

Levitsky, V. G.

Publications and source records attributed to Levitsky, V. G..

4 recordsLinked to original sources

MIMOSA: A model-independent framework for transcription factor binding site motif similarity assessment

Transcription factors (TFs) regulate gene expression by binding specific DNA sequences, called transcription factor binding sites (TFBSs), and motifs summarize the sequence specificity of these interactions. Although the position weight matrix (PWM) remains the most widely used motif model, alternative models can capture dependencies between nucleotide positions. Available tools for motif comparison are designed only for PWM motifs, and converting a motif from an alternative model into a PWM often leads to a loss of information. We propose MIMoSA (Model-Independent Motif Similarity Assessment), a tool that compares motif models independently of their representation. MIMoSA compares recognition profiles produced by different motifs on the same DNA sequence set rather than their internal parameters. Comparison of MIMoSA with PWM-based tools TomTom and MACRO-APE with the HOCOMOCO motif collection ensured comparable performance of all tools. A case study of a ChIPseq dataset for ATF3 TF further supported the reliability of MIMoSA application. The tool is available at \url{https://github.com/ubercomrade/mimosa}.

bioinformatics↗

NOVEL PRINCIPLES OF MOLECULAR GENETIC MAPPING OF THE INTERPHASE GENOME OF Drosophila melanogaster

Genome and chromosome maps have played a great role in the development of molecular genetics and biology. Gene activation, expression, and inactivation directly depend on the chromosomal (protein, nucleosomal) environment of the gene and formation of specific protein complexes on various gene structures. Polytene chromosomes are the only object in which interphase chromosomes can be analyzed, but the known Drosophila genome maps provide only an abstract view of gene distribution on a physical map, and there is no connection between these genes and the structures of interphase polytene chromosomes. A combination of bioinformatic methods was applied in this study; the data on interphase distribution of chromatin, H3K36me3 histone modifications, and the insulator protein Chriz, as well as the ChIP-seq, FAIRE-seq, and FISH methods, were used to investigate the genome-wide localization of key marker proteins. Having combined these mapping techniques for the small region 1AF of the X chromosome, we developed a novel method for analyzing the mutual arrangement of developmental and housekeeping genes, their promoters, and various types of proteins in the interphase genome and chromosome structures: compacted black bands, interbands, and gray bands, as well as sites of localization of exons and introns of housekeeping genes. Mapping was based on three consecutive stages: the 4 state Hidden Markov Model (hereinafter referred to as 4HMM) and data distribution for H3K36me3 and Chriz from the cells in which polytene chromosomes had been formed were used to localize interbands; FISH probes were then prepared from interband DNA, and blocks of developmental and housekeeping genes in interphase chromosome structures were mapped. The elaborated mapping methods can be further used to build similar maps for the entire interphase genome of Drosophila. Comparison of the map of bands and interbands in the region 1AF revealed full matching of the boundaries of black bands (developmental genes) and TADs. Within the regions where the housekeeping genes are located (the groups of interbands and gray bands), the TADs is formed on the basis of a gene cluster (the interband-gray band complex).

genomics↗

EBSn, a robust synthetic reporter for monitoring ethylene responses in plants

Ethylene is a gaseous plant hormone that controls a wide array of physiologically relevant processes, including plant responses to biotic and abiotic stress, and induces ripening in climacteric fruits. To monitor ethylene in plants, analytical methods, phenotypic assays, gene expression analysis, and transcriptional or translational reporters are typically employed. In the model plant Arabidopsis, two ethylene-sensitive synthetic transcriptional reporters have been described, 5xEBS:GUS and 10x2EBS-S10:GUS. These reporters harbor a different type, arrangement, and number of homotypic cis-elements in their promoters and thus may recruit the ethylene master regulator EIN3 in the context of alternative transcriptional complexes. Accordingly, the patterns of GUS activity in these transgenic lines differ and neither of them encompasses all plant tissues even in the presence of saturating levels of exogenous ethylene. Herein, we set out to develop and test a more sensitive version of the ethylene-inducible promoter that we refer to as EBSnew (abbreviated as EBSn). EBSn leverages a tandem of ten non-identical, natural copies of a novel, dual, everted, 11bp-long EIN3-binding site, 2EBS(-1). We show that in Arabidopsis, EBSn outperforms its predecessors in terms of its ethylene sensitivity, having the capacity to monitor endogenous levels of ethylene and displaying more ubiquitous expression in response to the exogenous hormone. We demonstrate that the EBSn promoter is also functional in tomato, opening new avenues to manipulating ethylene-regulated processes, such as ripening and senescence, in crops.

plant biology↗

Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data

Efficient de novo motif discovery from the results of wide-genome mapping of transcription factor binding sites (ChIP-seq) is dependent on the choice of background nucleotide sequences. The foreground sequences (peaks) represent not only specific motifs of target transcription factors, but also the motifs overrepresented throughout the genome, such as simple sequence repeats. We performed a massive comparison of the synthetic and genomic approaches to generate background sequences for de novo motif discovery. The synthetic approach shuffled nucleotides in peaks, while in the genomic approach randomly selected sequences from the reference genome or only from gene promoters according to the fraction of A/T nucleotides in each sequence. We compiled the benchmark collections of ChIP-seq datasets for mammalian and Arabidopsis, and performed de novo motif discovery. We showed that the genomic approach has both more robust detection of the known motifs of target transcription factors and more stringent exclusion of the simple sequence repeats as possible non-specific motifs. The advantage of the genomic approach over the synthetic one was greater in plants compared to mammals. We developed the AntiNoise web service (https://denovosea.icgbio.ru/antinoise/) which implements a genomic approach to extract genomic background sequences for twelve eukaryotic genomes.

bioinformatics↗