Search bioRxiv⌕ Search

Biology subjects

Zhao, Z.-G.

Publications and source records attributed to Zhao, Z.-G..

2 recordsLinked to original sources

From Randomness to Recognition: Modeling the Evolution of DNA Sequence Information During SELEX

Systematic Evolution of Ligands through Exponential Enrichment (SELEX) was used as a model system to explore the evolution of DNA sequence information and function during enrichment of molecular recognition to a series of related target molecules. Using a Natural Language Processing (NLP) based approach, a model was trained on an unlabeled mixture of oligonucleotide sequences sampled from both unenriched libraries and from six libraries that had been enriched for binding to either peptide or protein targets. This general model (pre-trained model) was then used to generate latent space representations of sequences that contained embedded binding information. Unsupervised clustering was used to create two dimensional maps that allowed comparison of the overlap between the latent space representations of sequences from different unenriched and enriched libraries. Replicate, independent enrichments to the same targets starting from completely unique random libraries gave rise to essentially indistinguishable latent space representations. However, similar representations between unenriched and enriched sequences, or between enriched sequences from different targets, resulted in distinct clustering patterns. The extent of overlap between the patterns from unenriched and enriched libraries was correlated with overall target binding by the enriched library. Further, the pre-trained model was fine-tuned to classify which target library each sequence belonged to, and the accuracy of classification between unenriched and enriched sequences was also correlated with the overall binding of the enriched library to its target. Finally, it was possible to distinguish libraries enriched for binding to two different targets differing by as few as 1 in 30 amino acids. Thus, using NLP-based approaches, it is possible to relate the DNA sequences present in an enriched library with the ability of the library to specifically bind its intended target, providing a tool for both guiding the enrichment process and more deeply understanding the structure-binding relationships that evolve.

biochemistry↗

Modeling the Sequence Dependence of Differential Antibody Binding in the Immune Response to Infectious Disease

Past studies have shown that incubation of human serum samples on high density peptide arrays followed by measurement of total antibody bound to each peptide sequence allows detection and discrimination of humoral immune responses to a wide variety of infectious disease agents. This is true even though these arrays consist of peptides with near-random amino acid sequences that were not designed to mimic biological antigens. Previously, this immune profiling approach or "immunosignature" has been implemented using a purely statistical evaluation of pattern binding, with no regard for information contained in the amino acid sequences themselves. Here, a neural network is trained on immunoglobulin G binding to 122,926 amino acid sequences selected quasi-randomly to represent a sparse sample of the entire combinatorial binding space in a peptide array using human serum samples from uninfected controls and 5 different infectious disease cohorts infected by either dengue virus, West Nile virus, hepatitis C virus, hepatitis B virus or Trypanosoma cruzi. This results in a sequence-binding relationship for each sample that contains the differential disease information. Processing array data using the neural network effectively aggregates the sequence-binding information, removing sequence-independent noise and improving the accuracy of array-based classification of disease compared to the raw binding data. Because the neural network model is trained on all samples simultaneously, the information common to all samples resides in the hidden layers of the model and the differential information between samples resides in the output layer of the model, one column of a few hundred values per sample. These column vectors themselves can be used to represent each sample for classification or unsupervised clustering applications such as human disease surveillance. Author SummaryPrevious work from Stephen Johnstons lab has shown that it is possible to use high density arrays of near-random peptide sequences as a general, disease agnostic approach to diagnosis by analyzing the pattern of antibody binding in serum to the array. The current approach replaces the purely statistical pattern recognition approach with a machine learning-based approach that substantially enhances the diagnostic power of these peptide array-based antibody profiles by incorporating the sequence information from each peptide with the measured antibody binding, in this case with regard to infectious diseases. This makes the array analysis much more robust to noise and provides a means of condensing the disease differentiating information from the array into a compact form that can be readily used for disease classification or population health monitoring.

immunology↗