Search bioRxiv⌕ Search

Biology subjects

Tsukanov, A. V.

Publications and source records attributed to Tsukanov, A. V..

3 recordsLinked to original sources

MIMOSA: A model-independent framework for transcription factor binding site motif similarity assessment

Transcription factors (TFs) regulate gene expression by binding specific DNA sequences, called transcription factor binding sites (TFBSs), and motifs summarize the sequence specificity of these interactions. Although the position weight matrix (PWM) remains the most widely used motif model, alternative models can capture dependencies between nucleotide positions. Available tools for motif comparison are designed only for PWM motifs, and converting a motif from an alternative model into a PWM often leads to a loss of information. We propose MIMoSA (Model-Independent Motif Similarity Assessment), a tool that compares motif models independently of their representation. MIMoSA compares recognition profiles produced by different motifs on the same DNA sequence set rather than their internal parameters. Comparison of MIMoSA with PWM-based tools TomTom and MACRO-APE with the HOCOMOCO motif collection ensured comparable performance of all tools. A case study of a ChIPseq dataset for ATF3 TF further supported the reliability of MIMoSA application. The tool is available at \url{https://github.com/ubercomrade/mimosa}.

bioinformatics↗

NOVEL PRINCIPLES OF MOLECULAR GENETIC MAPPING OF THE INTERPHASE GENOME OF Drosophila melanogaster

Genome and chromosome maps have played a great role in the development of molecular genetics and biology. Gene activation, expression, and inactivation directly depend on the chromosomal (protein, nucleosomal) environment of the gene and formation of specific protein complexes on various gene structures. Polytene chromosomes are the only object in which interphase chromosomes can be analyzed, but the known Drosophila genome maps provide only an abstract view of gene distribution on a physical map, and there is no connection between these genes and the structures of interphase polytene chromosomes. A combination of bioinformatic methods was applied in this study; the data on interphase distribution of chromatin, H3K36me3 histone modifications, and the insulator protein Chriz, as well as the ChIP-seq, FAIRE-seq, and FISH methods, were used to investigate the genome-wide localization of key marker proteins. Having combined these mapping techniques for the small region 1AF of the X chromosome, we developed a novel method for analyzing the mutual arrangement of developmental and housekeeping genes, their promoters, and various types of proteins in the interphase genome and chromosome structures: compacted black bands, interbands, and gray bands, as well as sites of localization of exons and introns of housekeeping genes. Mapping was based on three consecutive stages: the 4 state Hidden Markov Model (hereinafter referred to as 4HMM) and data distribution for H3K36me3 and Chriz from the cells in which polytene chromosomes had been formed were used to localize interbands; FISH probes were then prepared from interband DNA, and blocks of developmental and housekeeping genes in interphase chromosome structures were mapped. The elaborated mapping methods can be further used to build similar maps for the entire interphase genome of Drosophila. Comparison of the map of bands and interbands in the region 1AF revealed full matching of the boundaries of black bands (developmental genes) and TADs. Within the regions where the housekeeping genes are located (the groups of interbands and gray bands), the TADs is formed on the basis of a gene cluster (the interband-gray band complex).

genomics↗

Genomic background sequences systematically outperform synthetic ones in de novo motif discovery for ChIP-seq data

Efficient de novo motif discovery from the results of wide-genome mapping of transcription factor binding sites (ChIP-seq) is dependent on the choice of background nucleotide sequences. The foreground sequences (peaks) represent not only specific motifs of target transcription factors, but also the motifs overrepresented throughout the genome, such as simple sequence repeats. We performed a massive comparison of the synthetic and genomic approaches to generate background sequences for de novo motif discovery. The synthetic approach shuffled nucleotides in peaks, while in the genomic approach randomly selected sequences from the reference genome or only from gene promoters according to the fraction of A/T nucleotides in each sequence. We compiled the benchmark collections of ChIP-seq datasets for mammalian and Arabidopsis, and performed de novo motif discovery. We showed that the genomic approach has both more robust detection of the known motifs of target transcription factors and more stringent exclusion of the simple sequence repeats as possible non-specific motifs. The advantage of the genomic approach over the synthetic one was greater in plants compared to mammals. We developed the AntiNoise web service (https://denovosea.icgbio.ru/antinoise/) which implements a genomic approach to extract genomic background sequences for twelve eukaryotic genomes.

bioinformatics↗