Search bioRxivSearch

Biology subjects

Perez-Romero, C.

Publications and source records attributed to Perez-Romero, C..

2 recordsLinked to original sources

Design of Specific Primer Sets for the Detection of B.1.1.7, B.1.351 and P.1 SARS-CoV-2 Variants using Deep Learning

As the COVID-19 pandemic continues, new SARS-CoV-2 variants with potentially dangerous features have been identified by the scientific community. Variant B.1.1.7 lineage clade GR from Global Initiative on Sharing All Influenza Data (GISAID) was first detected in the UK, and it appears to possess an increased transmissibility. At the same time, South African authorities reported variant B.1.351, that shares several mutations with B.1.1.7, and might also present high transmissibility. Earlier this year, a variant labelled P.1 with 17 non-synonymous mutations was detected in Brazil. Recently the World Health Organization has raised concern for the variants B.1.617.2 mainly detected in India but now exported worldwide. It is paramount to rapidly develop specific molecular tests to uniquely identify new variants. Using a completely automated pipeline built around deep learning and evolutionary algorithms techniques, we designed primer sets specific to variants B.1.1.7, B.1.351, P.1 and respectively. Starting from sequences openly available in the GISAID repository, our pipeline was able to deliver the primer sets for each variant. In-silico tests show that the sequences in the primer sets present high accuracy and are based on 2 mutations or more. In addition, we present an analysis of key mutations for SARS-CoV-2 variants. Finally, we tested the designed primers for B.1.1.7 using RT-PCR. The presented methodology can be exploited to swiftly obtain primer sets for each new variant, that can later be a part of a multiplexed approach for the initial diagnosis of COVID-19 patients.

bioinformatics

Design of Specific Primer Set for Detection of B.1.1.7 SARS-CoV-2 Variant using Deep Learning

The SARS-CoV-2 variant B.1.1.7 lineage, also known as clade GR from Global Initiative on Sharing All Influenza Data (GISAID), Nextstrain clade 20B, or Variant Under Investigation in December 2020 (VUI - 202012/01), appears to have an increased transmissability in comparison to other variants. Thus, to contain and study this variant of the SARS-CoV-2 virus, it is necessary to develop a specific molecular test to uniquely identify it. Using a completely automated pipeline involving deep learning techniques, we designed a primer set which is specific to SARS-CoV-2 variant B.1.1.7 with >99% accuracy, starting from 8,923 sequences from GISAID. The resulting primer set is in the region of the synonymous mutation C16176T in the ORF1ab gene, using the canonical sequence of the variant B.1.1.7 as a reference. Further in-silico testing shows that the primer sets sequences do not appear in different viruses, using 20,571 virus samples from the National Center for Biotechnology Information (NCBI), nor in other coronaviruses, using 487 samples from National Genomics Data Center (NGDC). In conclusion, the presented primer set can be exploited as part of a multiplexed approach in the initial diagnosis of Covid-19 patients, or used as a second step of diagnosis in cases already positive to Covid-19, to identify individuals carrying the B.1.1.7 variant.

bioinformatics