Search bioRxiv⌕ Search

Biology subjects

Schein, C. H.

Publications and source records attributed to Schein, C. H..

3 recordsLinked to original sources

AllergenAI: a deep learning model predicting allergenicity based on protein sequence

Innovations in protein engineering can help redesign allergenic proteins to reduce adverse reactions in sensitive individuals. To accomplish this aim, a better knowledge of the molecular properties of allergenic proteins and the molecular features that make a protein allergenic is needed. We present a novel AI-based tool, AllergenAI, to quantify the allergenic potential of a given protein. Our approach is solely based on protein sequences, differentiating it from previous tools that use some knowledge of the allergens physicochemical and other properties in addition to sequence homology. We used the collected data on protein sequences of allergenic proteins as archived in the three well-established databases, SDAP 2.0, COMPARE, and AlgPred 2, to train a convolutional neural network and assessed its prediction performance by cross-validation. We then used Allergen AI to find novel potential proteins of the cupin family in date palm, spinach, maize, and red clover plants with a high allergenicity score that might have an adverse allergenic effect on sensitive individuals. By analyzing the feature importance scores (FIS) of vicilins, we identified a proline-alanine-rich (P-A) motif in the top 50% of FIS regions that overlapped with known IgE epitope regions of vicilin allergens. Furthermore, using[~] 1600 allergen structures in our SDAP database, we showed the potential to incorporate 3D information in a CNN model. Future, incorporating 3D information in training data should enhance the accuracy. AllergenAI is a novel foundation for identifying the critical features that distinguish allergenic proteins.

bioinformatics↗

Peptide and protein alphavirus antigens for broad spectrum vaccine design

Vaccines based on proteins and peptides may be safer and more broad-spectrum than other approaches Physicochemical property consensus (PCPcon) alphavirus antigens from the B-domain of the E2 envelope protein were designed and synthesized recombinantly. Those based on individual species (eastern or Venezuelan equine encephalitis (EEEVcon, VEEVcon), or chikungunya (CHIKVcon) viruses generated species-specific antibodies. Peptides designed to surface exposed areas of the E2-A-domain were added to the inocula to provide neutralizing antibodies against CHIKV. EVCcon, based on the three different alphavirus species, combined with E2-A-domain peptides from AllAV, a PCPcon of 24 diverse alphavirus, generated broad spectrum antibodies. The abs in the sera bound and neutralized diverse alphaviruses with less than 35% amino acid identity to each other. These included VEEV and its relative Mucambo virus, EEEV and the related Madariaga virus, and CHIKV strain 181/25. Further understanding of the role of coordinated mutations in the envelope proteins may yield a single, protein and peptide vaccine against all alphaviruses.

bioengineering↗

D-graph clusters flaviviruses and β-coronaviruses according to their hosts, disease type and human cell receptors

MotivationThere is a need for rapid and easy to use, alignment free methods to cluster large groups of protein sequence data. Commonly used phylogenetic trees based on alignments can be used to visualize only a limited number of protein sequences. DGraph, introduced here, is a dynamic programming application developed to generate 2D-maps based on similarity scores for sequences. The program automatically calculates and graphically displays property distance (PD) scores based on physico-chemical property (PCP) similarities from an unaligned list of FASTA files. Such "PD-graphs" show the interrelatedness of the sequences, whereby clusters can reveal deeper connectivities. ResultsPD-Graphs generated for flavivirus (FV), enterovirus (EV), and coronavirus (CoV) sequences from complete polyproteins or individual proteins are consistent with biological data on vector types, hosts, cellular receptors and disease phenotypes. PD-graphs separate the tick- from the mosquito-borne FV, clusters viruses that infect bats, camels, seabirds and humans separately and the clusters correlate with disease phenotype. The PD method segregates the {beta}-CoV spike proteins of SARS, SARS-CoV-2, and MERS sequences from other human pathogenic CoV, with clustering consistent with cellular receptor usage. The graphs also suggest evolutionary relationships that may be difficult to determine with conventional bootstrapping methods that require postulating an ancestral sequence. Availability and implementationDGraph is written in Java, compatible with the Java 5 runtime or newer. Source code and executable is available from the GitHub website (https://github.com/bjmnbraun/DGraph/releases). Documentation for installation and use of the software is available from the Readme.md file at (https://github.com/bjmnbraun/DGraph). Contactbjmnbraun@gmail.com or webraun@utmb.edu Supplementary informationSupplementary information Table S1 and Fig. S1 are online available.

bioinformatics↗