Search bioRxivSearch

Biology subjects

Crossman, L. C.

Publications and source records attributed to Crossman, L. C..

2 recordsLinked to original sources

Leverging Deep Learning to Simulate Coronavirus Spike proteins has the potential to predict future Zoonotic sequences

MotivationCoronaviridae are a family of positive-sense RNA viruses capable of infecting humans and animals. These viruses usually cause a mild to moderate upper respiratory tract infection, however, they can also cause more severe symptoms, gastrointestinal and central nervous system diseases. These viruses are capable of flexibly adapting to new environments, hence health threats from coronavirus are constant and long-term. Immunogenic spike proteins are glyco-proteins found on the surface of Coronaviridae particles that mediate entry to host cells. The aim of this study was to train deep learning neural networks to produce simulated spike protein sequences, which may be able to aid in knowledge and/or vaccine design by creating alternative possible spike sequences that could arise from zoonotic sources in future. ResultsHere we have trained deep learning recurrent neural networks (RNN) to provide computer-simulated coronavirus spike protein sequences in the style of previously known sequences and examine their characteristics. Training used a dataset of alpha, beta, gamma and delta coronavirus spike sequences. In a test set of 100 simulated sequences, all 100 had most significant BLAST matches to Spike proteins in searches against NCBI non-redundant dataset (NR) and also possessed concomitant Pfam domain matches. ConclusionsSimulated sequences from the neural network may be able to guide us in future with prospective targets for vaccine discovery in advance of a potential novel zoonosis. We may effectively be able to fast-forward through evolution using neural networks to investigate sequences that could arise.

bioinformatics

Punchline: Identifying and comparing significant Pfam protein domain differences across draft whole genome sequences

MotivationShort-read draft paired-end Illumina assemblies can be fragmented, contain many contigs and be impacted on by repeat regions, caused by mobile element activity within the genome or inherently repetitive gene structure. Annotating such assemblies for function and analysing gene content can be challenging if predicted genes are fragmented across contigs. Such a case can often occur within specific families of genes such as longer genes with repeating domains, genes specifying several transmembrane domains and of unusual nucleotide content. These genes can often be virulence determinants, therefore losing these specific types of data can seriously impact downstream studies.\n\nResultsRather than studying the predicted gene content of draft genomes, we examined predicted protein content using the Pfam domain complements of predicted proteins. We produced a workflow, Punchline, to study the genetic content of draft contig assemblies by looking at the complement of short domains that are unlikely to be affected. We investigated a dataset of Bacteroides ovatus in terms of a grouping involving the vertebrate host from which the organism was isolated and identified potential host restricted functions and host restricted phylogenetic clustering.\n\nAvailabilityhttps://github.com/LCrossman\n\nContact: seq@sequenceanalysis.co.uk

genomics