Search bioRxiv⌕ Search

Biology subjects

House, T. A.

Publications and source records attributed to House, T. A..

2 recordsLinked to original sources

Analysis and comprehensive lineage identification for SARS-CoV-2 genomes through scalable learning methods

Since its emergence in late 2019, SARS-CoV-2 has diversified into a large number of lineages and globally caused multiple waves of infection. Novel lineages have the potential to spread rapidly and internationally if they have higher intrinsic transmissibility and/or can evade host immune responses, as has been seen with the Alpha, Delta, and Omicron variants of concern (VoC). They can also cause increased mortality and morbidity if they have increased virulence, as was seen for Alpha and Delta, but not Omicron. Phylogenetic methods provide the gold standard for representing the global diversity of SARS-CoV-2 and to identify newly emerging lineages. However, these methods are computationally expensive, struggle when datasets get too large, and require manual curation to designate new lineages. These challenges together with the increasing volumes of genomic data available provide a motivation to develop complementary methods that can incorporate all of the genetic data available, without down-sampling, to extract meaningful information rapidly and with minimal curation. Here, we demonstrate the utility of using algorithmic approaches based on word-statistics to represent whole sequences, bringing speed, scalability, and interpretability to the construction of genetic topologies, and while not serving as a substitute for current phylogenetic analyses the proposed methods can be used as a complementary approach to identify and confirm new emerging variants.

bioinformatics↗

Quantification of circadian interactions and protein abundance defines a mechanism for operational stability of the circadian clock

The mammalian circadian clock exerts substantial control of daily gene expression through cycles of DNA binding. Understanding of mechanisms driving the circadian clock is hampered by lack of quantitative data, without which predictive mathematical models cannot be developed. Here we develop a quantitative understanding of how a finite pool of BMAL1 protein can regulate thousands of target sites over daily time scales. We have used fluorescent correlation spectroscopy (FCS) to track dynamic changes in CRISPR-modified fluorophore-tagged proteins in time and space in single cells across SCN and peripheral tissues. We determine the contribution of multiple rhythmic processes in coordinating BMAL1 DNA binding, including the roles of cycling molecular abundance, binding affinities and two repressive modes of action. We find that nuclear BMAL1 protein numbers determine corresponding nuclear CLOCK concentrations through heterodimerization and define a DNA residence time of 2.6 seconds for this complex. Repression of CLOCK:BMAL1 is in part achieved through rhythmic changes to BMAL1:CRY1 affinity as well as a high affinity interaction between PER2:CRY1 which mediates CLOCK:BMAL1 displacement from DNA. Finally, stochastic modelling of these data reveals a dual role for PER:CRY complexes in which increasing concentrations of PER2:CRY1 promotes removal of BMAL1:CLOCK from genes consequently enhancing ability to move to new target sites.

systems biology↗