Search bioRxiv⌕ Search

Biology subjects

Vorontsov, I.

Publications and source records attributed to Vorontsov, I..

2 recordsLinked to original sources

GAME: Genomic API for Model Evaluation

The rapid expansion of genomics datasets and the application of machine learning has produced sequence-to-activity genomics models with ever-expanding capabilities. However, benchmarking these models on practical applications has been challenging because individual projects evaluate their models in ad hoc ways, and there is substantial heterogeneity of both model architectures and benchmarking tasks. To address this challenge, we have created GAME, a system for large-scale, community-led standardized model benchmarking on user-defined evaluation tasks. We borrow concepts from the Application Programming Interface (API) paradigm to allow for seamless communication between pre-trained models and benchmarking tasks, ensuring consistent evaluation protocols. Because all models and benchmarks are inherently compatible in this framework, the continual addition of new models and new benchmarks is easy. We also developed a Matcher module powered by a large language model (LLM) to automate ambiguous task alignment between benchmarks and models. Containerization of these modules enhances reproducibility and facilitates the deployment of models and benchmarks across computing platforms. By focusing on predicting underlying biochemical phenomena (e.g. gene expression, open chromatin, DNA binding), we ensure that tasks remain technology-independent. We provide examples of benchmarks and models implementing this framework, and anticipate that the community will contribute their own, leading to an ever-expanding and evolving set of models and evaluation tasks. This resource will accelerate genomics research by illuminating the best models for a given task, motivating novel functional genomic benchmarks, and providing a more nuanced understanding of model abilities.

bioinformatics↗

Extensive binding of uncharacterized human transcription factors to genomic dark matter

The functional impact of a large portion of the human genome known as "dark matter DNA", which is composed mainly of repeat sequences, remains enigmatic. The genome also encodes hundreds of putative and poorly characterized transcription factors (TFs). Here, we determined genomic binding locations of 166 poorly characterized human TFs in living cells. Nearly half of them associate strongly with known regulatory regions such as promoters and enhancers, frequently co-localizing with each other at conserved motif matches. The other half often associate with genomic dark matter, however, at largely non-overlapping (i.e., unique) sites, via intrinsic sequence recognition. Fifty-four of the latter half, which we term "Dark TFs", mainly bind within regions of closed chromatin, with each recognizing a unique set of repeat sequences. The Dark TFs include many KZNFs, which are known to bind and silence TEs, and other TFs with apparent repressive functions. By contrast, some may be pioneers: we find that induction of TPRX1, a known regulator of zygotic preimplantation, leads to chromatin opening at many of its binding sites in the dark matter genome. Altogether, our results shed light on a large fraction of poorly characterized human TFs and simultaneously illuminate the diversity of function within the dark matter genome.

genomics↗