Search bioRxiv⌕ Search

Biology subjects

Goncharov, A. O.

Publications and source records attributed to Goncharov, A. O..

3 recordsLinked to original sources

Massive proteogenomic reanalysis of publicly available proteomic datasets of human tissues in search for protein recoding via adenosine-to-inosine RNA editing

The proteogenomic search pipeline developed in this work has been applied for re-analysis of 40 publicly available shotgun proteomic datasets from various human tissues comprising more than 8,000 individual LC-MS/MS runs, of which 5442 .raw data files were processed in total. The scope of this re-analysis was focused on searching for ADAR-mediated RNA editing events, their clustering across samples of different origin, and classification. In total, 33 recoded protein sites were identified in 21 datasets. Of those, 18 sites were detected in at least two datasets representing the core human protein editome. In agreement with prior art works, neural and cancer tissues were found being enriched with recoded proteins. Quantitative analysis indicated that recoding of specific sites did not directly depend on the levels of ADAR enzymes or targeted proteins themselves, rather it was provided by differential and yet undescribed regulation of interaction of enzymes with mRNA. Nine recoding sites conservative between human and rodents were validated by targeted proteomics using stable isotope standards in murine brain cortex and cerebellum, and an additional one was validated in human cerebrospinal fluid. In addition to previous data of the same type from cancer proteomes, we provide a comprehensive catalog of recoding events caused by ADAR RNA editing in the human proteome.

bioinformatics↗

Validating amino acid variants in proteogenomics using sequence coverage by multiple reads

Mass spectrometry-based proteome analysis usually implies matching mass spectra of proteolytic peptides to amino acid sequences predicted from nucleic acid sequences. At the same time, due to the stochastic nature of the method when it comes to proteome-wide analysis, in which only a fraction of peptides are selected for sequencing, the completeness of protein sequence identification is undermined. Likewise, the reliability of peptide variant identification in proteogenomic studies is suffering. We propose a way to interpret shotgun proteomics results, specifically in data-dependent acquisition mode, as protein sequence coverage by multiple reads, just as it is done in the field of nucleic acid sequencing for the calling of single nucleotide variants. Multiple reads for each position in a sequence could be provided by overlapping distinct peptides, thus, confirming the presence of certain amino acid residues in the overlapping stretch with much lower false discovery rate than conventional 1%. The source of overlapping distinct peptides are, first, miscleaved tryptic peptides in combination with their properly cleaved counterparts, and, second, peptides generated by several proteases with different specificities after the same specimen is subject to parallel digestion and analyzed separately. We illustrate this approach using publicly available multiprotease proteomic datasets and our own data generated for HEK-293 cell line digests obtained using trypsin, LysC and GluC proteases. From 5000 to 8000 protein groups are identified for each digest corresponding to up to 30% of the whole proteome coverage. Most of this coverage was provided by a single read, while up to 7% of the observed protein sequences were covered two-fold and more. The proteogenomic analysis of HEK-293 cell line revealed 36 peptide variants associated with SNP, seven of which were supported by multiple reads. The efficiency of the multiple reads approach depends strongly on the depth of proteome analysis, the digesting features such as the level of miscleavages, and will increase with the number of different proteases used in parallel proteome digestion. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=75 SRC="FIGDIR/small/475497v1_ufig1.gif" ALT="Figure 1"> View larger version (14K): org.highwire.dtl.DTLVardef@1d6ee2org.highwire.dtl.DTLVardef@5ae8baorg.highwire.dtl.DTLVardef@652216org.highwire.dtl.DTLVardef@1a0d49b_HPS_FORMAT_FIGEXP M_FIG C_FIG

bioinformatics↗

Proteome-Wide Analysis of ADAR-mediated Messenger RNA Editing During Fruit Fly Ontogeny

Adenosine-to-inosine RNA editing is an enzymatic post-transcriptional modification which modulates immunity and neural transmission in multicellular organisms. Some of its functions are enforced through editing of mRNA codons with the resulting amino acid substitutions. We identified these sites originated from the RNA editing for developmental proteomes of Drosophila melanogaster at the protein level using available proteomic data for fifteen stages of fruit fly development from egg to imago and fourteen time points of embryogenesis. In total, 42 sites each belonging to a unique protein were found including four sites related to embryogenesis. The interactome analysis has revealed that most of the edited proteins are associated with synaptic vesicle trafficking and actomyosin organization. Quantitation data analysis suggested the existence of phase-specific RNA editing regulation by yet unknown mechanisms. These results support transcriptome analyses showing that a burst in RNA editing occurs during insect metamorphosis from pupa to imago. Further, targeted proteomics was employed to quantify edited and genomically encoded versions of five proteins in brains of larvae, pupae, and imago insects showing a clear trend towards an increase in editing rate for all of them. Our results may help to reveal the protein functions in physiological effects of RNA editing. SignificanceAdenosine-to-inosine RNA editing has multiple effects on body functions in many multicellular organisms from insects and molluscs to humans. Recent studies show that at least some of these effects are mediated by changes in protein sequences due to editing of codons in mRNA. However, it is not known how exactly the edited proteins can participate in RNA editing-mediated pathways. Moreover, most studies of edited proteins are based on the deduction of protein sequence changes from analysis of transcriptome without measurements of proteins themselves. Earlier, we explored for the first time the edited proteins of Drosophila melanogaster proteome. In this work, we continued the proteome-wide analysis of RNA editome using shotgun proteomic data of ontogeny phases of this model insect. It was found that non-synonymous RNA editing, which led to translation of changed proteins, is specific to the life cycle phase. Identification of tryptic peptides containing edited protein sites provides a basis for further direct and quantitative analysis of their editing rate by targeted proteomics. The latter was demonstrated in this study by multiple reaction monitoring experiments which were used to observe the dynamics of editing in selected brain proteins during developmental phases of fruit fly. HighlightsO_LIProteogenomic approach was applied to shotgun proteomics data of fruit fly ontogeny for identification of proteoforms originating from adenosine-to-inosine RNA editing. C_LIO_LIEdited proteins identified at all life cycle stages are enriched in annotated protein-protein interactions at statistically significant level with many of them associated with actomyosin and synaptic vesicle functions. C_LIO_LIProteome-wide RNA editing event profiles were found specific to life cycle phase and independent of the protein abundances. C_LIO_LIA majority of RNA editing events at the protein level was observed after metamorphosis in late pupae to adult insects, which was consistent with transcriptome data. C_LIO_LITargeted proteomic analysis of five selected edited sites and their genomic counterparts in brains for three phases of the fruit fly life cycle have demonstrated a clear increase in editing rate of up to 80% for the endophilin A protein in adult flies. C_LI

systems biology↗