Search bioRxiv⌕ Search

Biology subjects

Kamenets, V.

Publications and source records attributed to Kamenets, V..

2 recordsLinked to original sources

EPIC: An open community challenge for sequence-based prediction of transcription initiation in five non-model metazoans

Predicting gene expression from DNA sequence is a central problem with critical biological and clinical implications. Recent sequence-to-function models are reported to achieve improved performance, yet it remains unclear how much of what they learn reflects genuine principles of eukaryotic transcription initiation and to what extent they are able to generalize beyond humans and other primary model species. A fair and blind benchmark has likewise been missing. Here we introduce EPIC, the Eukaryotic Promoter and transcription Initiation prediction Challenge. Teams receive strand-specific, single-nucleotide-resolution initiation profiles for 80-95% of the genome and are asked to predict the remainder from DNA sequence alone. To level the field, the challenge relies on understudied animals spanning three phyla: octopus, oyster, milkweed bug, Indian meal moth, and shark. EPIC is open to everyone and closes on December 31, 2026. All teams clearing the dinucleotide precision baseline are invited to join the consortium authorship of the post-challenge publication, and winners are invited for personal authorship.

bioinformatics↗

Inferring binding specificities of human transcription factors with the wisdom of crowds

DNA motif discovery and, particularly, computational modeling of transcription factor binding motifs, has been a mecca of algorithmic bioinformatics for several decades. Here, we report the results of the largest open community challenge in Inferring BInding Specificities (IBIS), where participants all over the world were invited to construct binding specificity models from multi-assay experimental data for poorly studied human transcription factors. The submissions were rigorously tested against a rich held-out dataset. Benchmarking demonstrated a consistent advantage of properly designed deep learning models over traditional positional weight matrices and other machine learning methods. Yet, the positional weight matrices displayed a surprisingly strong performance out of the box, being only slightly behind the best deep learning models. A post-challenge assessment of a selection of other deep learning methods further solidified this finding. IBIS highlights the power of benchmarking in finding adequate DNA motif representations, emphasizes the pros and cons of various machine learning methods applied to DNA motif modeling, and establishes a rich dataset, benchmarking protocols, and computational framework for a fair cross-platform evaluation of future models of transcription factor binding motifs in DNA sequences. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=175 SRC="FIGDIR/small/688692v1_ufig1.gif" ALT="Figure 1"> View larger version (64K): org.highwire.dtl.DTLVardef@1c6677corg.highwire.dtl.DTLVardef@b4124aorg.highwire.dtl.DTLVardef@1ce2b1org.highwire.dtl.DTLVardef@66e917_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗