Search bioRxivSearch

bioRxiv · 10.1101/2021.01.26.428348

Deep Directed Evolution of Solid Binding Peptides for Quantitative Big-Data Generation

Abstract

Proteins have evolved over millions of years to mediate and carry-out biological processes efficiently. Directed evolution approaches have been used to genetically engineer proteins with desirable functions such as catalysis, mineralization, and target-specific binding. Next-generation sequencing technology offers the capability to discover a massive combinatorial sequence space that is costly to sample experimentally through traditional approaches. Since the permutation space of protein sequence is virtually infinite, and evolution dynamics are poorly understood, experimental verifications have been limited. Recently, machine-learning approaches have been introduced to guide the evolution process that facilitates a deeper and denser search of the sequence-space. Despite these developments, however, frequently used high-fidelity models depend on massive amounts of properly labeled quality data, which so far has been largely lacking in the literature. Here, we provide a preliminary high-throughput peptide-selection protocol with functional scoring to enhance the quality of the data. Solid binding dodecapeptides have been selected against molybdenum disulfide substrate, a two-dimensional atomically thick semiconductor solid. The survival rate of the phage-clones, upon successively stringent washes, quantifies the binding affinity of the peptides onto the solid material. The method suggested here provides a fast generation of preliminary data-pool with [~]2 million unique peptides with 12 amino-acids per sequence by avoiding amplification. Our results demonstrate the importance of data-cleaning and proper conditioning of massive datasets in guiding experiments iteratively. The established extensive groundwork here provides unique opportunities to further iterate and modify the technique to suit a wide variety of needs and generate various peptide and protein datasets. Prospective statistical models developed on the datasets to efficiently explore the sequence-function space will guide towards the intelligent design of proteins and peptides through deep directed evolution. Technological applications of the future based on the peptide-single layer solid based bio/nano soft interfaces, such as biosensors, bioelectronics, and logic devices, is expected to benefit from the solid binding peptide dataset alone. Furthermore, protocols described herein will also benefit efforts in medical applications, such as vaccine development, that could significantly accelerate a global response to future pandemics.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yucesoy, D. T., Rath, S. S., Rodriguez, J. L., Francis-Landau, J. T., Nakano-Baker, O., Sarikaya, M.. 2021-01-27. Deep Directed Evolution of Solid Binding Peptides for Quantitative Big-Data Generation. https://doi.org/10.1101/2021.01.26.428348

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Trans-branching of polyubiquitin chains orchestrates the DNA replication stress response

Polyubiquitin chain geometry dictates functional consequences of ubiquitylation. Although branched polyubiquitin chains are abundant in cells, little is known about their functions. Here we show that branching on the DNA replication factor PCNA, mediated by the ubiquitin-conjugating enzyme UBE2K and involving lysines 63 and 48 of ubiquitin, orchestrates the sequence of events in response to replication stress. By inducing VCP-dependent extraction of PCNA from chromatin, branching promotes re-priming of stalled forks and necessitates a BRCA1-dependent pathway of daughter-strand gap repair. Our study identifies hyper-accumulation of daughter-strand gaps as the mechanistic basis underlying the toxicity of inhibitors of the PCNA-specific isopeptidase, USP1, in BRCA1-deficient cells. Moreover, an unexpected preference of UBE2K to operate in trans suggests a general timing mechanism to organize hierarchies amongst ubiquitin signals.

molecular biology

Impaired proteostasis is an early feature of the diabetic heart in humans and mice

Diabetes and obesity increase cardiac lipid levels leading to cardiomyopathy and heart failure. We hypothesized that intermittent fasting would reduce cardiac lipid levels. Surprisingly, intermittent fasting increased myocardial triglyceride content, but rescued mortality and attenuated cardiomyopathy in mice overexpressing cardiomyocyte acyl-CoA synthetase 1 (MHC-ACSL1). Lipid overload caused cardiomyocyte accumulation of polyubiquitinated protein aggregates containing desmin, a scaffolding intermediate filament protein, which intermittent fasting prevented. Furthermore, intermittent fasting reversed elevated myocardial C16:0 ceramide content, and knockdown of ceramide synthase CerS5 and CerS6 reduced palmitate-induced protein aggregation, highlighting a role for C16:0 ceramides in this pathology. Conversely, impairing aggrephagy with cardiomyocyte-specific p62 ablation induced heart failure in mice fed a high-fat diet, with paradoxically reduced cardiac lipid content. Crucially, non-failing diabetic human hearts also exhibited protein aggregate pathology. Taken together, these results demonstrate that impaired proteostasis characterizes cardiomyopathy from cardiac lipid overload and identify a promising new therapeutic target for this condition.

molecular biology

Spatial profiling and neurovascular communication in the developing and adolescent cortex following prenatal alcohol exposure

Fetal alcohol spectrum disorders (FASD) constitute a wide range of developmental, cognitive, and behavioral impairments caused by prenatal alcohol exposure (PAE). Although neuronal and vascular consequences of PAE have been studied, how alcohol affects the cerebrovasculature within the framework of the neurovascular unit (NVU) across development remains poorly understood. At minimum, the NVU comprises neurons, astrocyte endfeet, and endothelial cells (ECs), which coordinate to maintain brain homeostasis. Here, we used the NanoString Digital Spatial Profiling platform to characterize spatial transcriptomic data from neurons, astrocytes, and ECs from PAE and saccharin (SAC) control cortices at embryonic day 18 (E18) and postnatal day 28 (P28). Differentially expressed genes were then used for Ingenuity Pathway Analysis (IPA) to identify altered biological pathways and perform comparison analyses across developmental time points, while CellChat was used to infer cell cell communication networks. We uncovered thousands of differentially expressed genes and numerous altered pathways and biological processes in PAE cortices across development. Both IPA and CellChat analyses implicated dysregulation of vascular and extracellular matrix (ECM) remodeling, cell adhesion, and neuroinflammatory signaling. CellChat further predicted the loss of several key bidirectional relationships and altered ligand-receptor interactions among neurovascular cell types at E18 and P28. Overall, these findings identify PAE associated alterations in neurovascular gene expression and intercellular signaling across development, providing potential mechanisms by which PAE may disrupt neurodevelopment.

molecular biology