Search bioRxiv⌕ Search

Biology subjects

Barbadilla-Martinez, L.

Publications and source records attributed to Barbadilla-Martinez, L..

2 recordsLinked to original sources

Decoding the Sequence Requirements for Translation Initiation

Accurate selection of start codons by ribosomes is a fundamental determinant of proteome composition. Although the Kozak sequence--an 8-nucleotide sequence flanking the start codon--has long been viewed as the primary determinant of initiation in eukaryotes, it fails to explain the large diversity of start codon usage across transcripts. Here we combine massively parallel reporter assays, bioinformatics, machine learning, single-molecule imaging and cryo-electron microscopy to define the extended translation initiation sequence (eTIS), an [~]80-nucleotide sequence surrounding the start codon that governs initiation efficiency. A deep-learning model trained on eTIS features accurately predicts translation initiation across transcripts. Unexpectedly, we find that the Kozak sequence is not optimal for initiation as is widely presumed, and we identify the origin of this discrepancy. eTIS nucleotides that promote efficient initiation are enriched in the human transcriptome and are evolutionarily conserved, underscoring their functional importance. Biophysical and structural analyses reveal that specific eTIS residues--including the key +6 position and residues in the mRNA entry and exit channel--engage ribosomal proteins, rRNA and initiation factors to promote start codon recognition by stabilizing the ribosome at the start codon and facilitating the structural transitions required for initiation. Finally, optimization of the eTIS markedly enhances translational fidelity and protein output from therapeutic mRNAs, highlighting its practical utility. Together, these findings redefine the sequence logic of translation initiation and establish a framework for precise control of protein expression.

molecular biology↗

The regulatory grammar of human promoters uncovered by MPRA-trained deep learning

One of the major challenges in genomics is to build computational models that accurately predict genome-wide gene expression from the sequences of regulatory elements. At the heart of gene regulation are promoters, yet their regulatory logic is still incompletely understood. Here, we report PARM, a cell-type specific deep learning model trained on specially designed massively parallel reporter assays that query human promoter sequences. PARM requires [~]1,000 times less computational power than state-of-the-art technology, and reliably predicts autonomous promoter activity throughout the genome from DNA sequence alone, in multiple cell types. PARM can even design purely synthetic strong promoters. We leveraged PARM to systematically identify binding sites of transcription factors (TFs) that are likely to contribute to the activity of each natural human promoter. We uncovered and experimentally confirmed striking positional preferences of TFs that differ between activating and repressive regulatory functions, as well as a complex grammar of motif-motif interactions. For example, many, but not all, TFs act as repressors when their binding motif is located near or just downstream of the transcription start site. Our approach lays the foundation towards a deep understanding of the regulation of human promoters by TFs. HighlightsO_LICausality-trained deep learning model PARM captures regulatory grammar of human promoters C_LIO_LIPARM is highly economical, both experimentally and computationally C_LIO_LITranscription factors have different preferred positions for their regulatory activity C_LIO_LIMany (but not all) transcription factors act as repressors when binding downstream of transcription start sites C_LI

genetics↗