Search bioRxiv⌕ Search

Biology subjects

Mukherjee, V.

Publications and source records attributed to Mukherjee, V..

4 recordsLinked to original sources

Predicting Epitope Candidates for SARS-CoV-2

Epitopes are short amino acid sequences that define the antigen signature to which an antibody binds. In light of the current pandemic, epitope analysis and prediction is paramount to improving serological testing and developing vaccines. In this paper, we leverage known epitope sequences from SARS-CoV, SARS-CoV-2 and other Coronaviridae and use those known epitopes to identify additional antigen regions in 62k SARS-CoV-2 genomes. Additionally, we present epitope distribution across SARS-CoV-2 genomes, locate the most commonly found epitopes, discuss where epitopes are located on proteins, and how epitopes can be grouped into classes. We also discuss the mutation density of different regions on proteins using a big data approach. We find that there are many conserved epitopes between SARS-CoV-2 and SARS-CoV, with more diverse sequences found in Nucleoprotein and Spike Glycoprotein.

bioinformatics↗

Semi-supervised identification of SARS-CoV-2 molecular targets

SARS-CoV-2 genomic sequencing efforts have scaled dramatically to address the current global pandemic and aid public health. In this work, we analyzed a corpus of 66,000 SARS-CoV-2 genome sequences. We developed a novel semi-supervised pipeline for automated gene, protein, and functional domain annotation of SARS-CoV-2 genomes that differentiates itself by not relying on use of a single reference genome and by overcoming atypical genome traits. Using this method, we identified the comprehensive set of known proteins with 98.5% set membership accuracy and 99.1% accuracy in length prediction compared to proteome references including Replicase polyprotein 1ab (with its transcriptional slippage site). Compared to other published tools such as Prokka (base) and VAPiD, we yielded an 6.4- and 1.8-fold increase in protein annotations. Our method generated 13,000,000 molecular target sequences-- some conserved across time and geography while others represent emerging variants. We observed 3,362 non-redundant sequences per protein on average within this corpus and describe key D614G and N501Y variants spatiotemporally. For spike glycoprotein domains, we achieved greater than 97.9% sequence identity to references and characterized Receptor Binding Domain variants. Here, we comprehensively present the molecular targets to refine biomedical interventions for SARS-CoV-2 with a scalable high-accuracy method to analyze newly sequenced infections.

bioinformatics↗

A CRISPRi screen of essential genes reveals that proteasome regulation dictates acetic acid tolerance in Saccharomyces cerevisiae

CRISPR interference (CRISPRi) is a powerful tool to study cellular physiology under different growth conditions and this technology provides a means for screening changed expression of essential genes. In this study, a Saccharomyces cerevisiae CRISPRi library was screened for growth in medium supplemented with acetic acid. Acetic acid is a growth inhibitor challenging the use of yeast for industrial conversion of lignocellulosic biomasses. Tolerance towards acetic acid that is released during biomass hydrolysis is crucial for cell factories to be used in biorefineries. The CRISPRi library screened consists of >9,000 strains, where >98% of all essential and respiratory growth-essential genes were targeted with multiple gRNAs. The screen was performed using the high-throughput, high-resolution Scan-o-matic platform, where each strain is analyzed separately. Our study identified that CRISPRi targeting of genes involved in vesicle formation or organelle transport processes led to severe growth inhibition during acetic acid stress, emphasizing the importance of these intracellular membrane structures in maintaining cell vitality. In contrast, strains in which genes encoding subunits of the 19S regulatory particle of the 26S proteasome were downregulated had increased tolerance to acetic acid, which we hypothesize is due to ATP-salvage through an increased abundance of the 20S core particle that performs ATP-independent protein degradation. This is the first study where a high-resolution CRISPRi library screening paves the way to understand and bioengineer the robustness of yeast against acetic acid stress. IMPORTANCEAcetic acid is inhibitory to the growth of the yeast Saccharomyces cerevisiae, causing ATP starvation and oxidative stress, which leads to sub-optimal production of fuels and chemicals from lignocellulosic biomass. In this study, where each strain of a CRISPRi library was characterized individually, many essential and respiratory growth essential genes that regulate tolerance to acetic acid were identified, providing new understanding on the stress response of yeast and new targets for the bioengineering of industrial yeast. Our findings on the fine-tuning of the expression of proteasomal genes leading to increased tolerance to acetic acid suggests that this could be a novel strategy for increasing stress tolerance, leading to improved strains for production of biobased chemicals.

synthetic biology↗

Analysis and Forecasting of Global of RT-PCR Primers for SARS-CoV-2

Rapid tests for active SARS-CoV-2 infections rely on reverse transcription polymerase chain reaction (RT-PCR). RT-PCR uses reverse transcription of RNA into complementary DNA (cDNA) and amplification of specific DNA (primer and probe) targets using polymerase chain reaction (PCR). The technology makes rapid and specific identification of the virus possible based on sequence homology of nucleic acid sequence and is much faster than tissue culture or animal cell models. However the technique can lose sensitivity over time as the virus evolves and the target sequences diverge from the selective primer sequences. Different primer sequences have been adopted in different geographic regions. As we rely on these existing RT-PCR primers to track and manage the spread of the Coronavirus, it is imperative to understand how SARS-CoV-2 mutations, over time and geographically, diverge from existing primers used today. In this study, we analyze the performance of the SARS-CoV-2 primers in use today by measuring the number of mismatches between primer sequence and genome targets over time and spatially. We find that there is a growing number of mismatches, an increase by 2% per month, as well as a high specificity of virus based on geographic location.

bioinformatics↗