Search bioRxiv⌕ Search

Biology subjects

Maxwell, M. J.

Publications and source records attributed to Maxwell, M. J..

2 recordsLinked to original sources

Modelling group heteroscedasticity in single-cellRNA-seq pseudo-bulk data

Group heteroscedasticity is commonly observed in pseudo-bulk single-cell RNA-seq datasets and when not modelled appro-priately, its presence can hamper the detection of differentially expressed genes. Most bulk RNA-seq methods assume equal group variances which will under- and/or over-estimate the true variability in such datasets. We present two methods that account for heteroscedastic groups, namely voomByGroup and voomWithQualityWeights using a blocked design (voomQWB). Compared to current gold standard methods that do not account for heteroscedasticity, we show results from simulation studies and various experiments that demonstrate the superior performance of both voomByGroup and voomQWB in error control and power when group variances in pseudo-bulk scRNA-seq data are unequal. We recommend the use of either of these methods over established approaches, with voomByGroup having the advantage of accurate variance estimation since group variance trends can take on different "shapes", whilst voomQWB has the advantage of catering to complex study designs.

bioinformatics↗

ScrepYard: an online resource for disulfide-stabilised tandem repeat peptides

Receptor avidity through multivalency is a highly sought-after property of ligands. While readily available in nature in the form of bivalent antibodies, this property remains challenging to engineer in synthetic molecules. The discovery of several bivalent venom peptides containing two homologous and independently folded domains (in a tandem repeat arrangement) has provided a unique opportunity to better understand the underpinning design of multivalency in multimeric biomolecules, as well as how naturally occurring multivalent ligands can be identified. In previous work we classified these molecules as a larger class termed secreted cysteine-rich repeat-proteins (SCREPs). Here, we present an online resource; ScrepYard, designed to assist researchers in identification of SCREP sequences of interest and to aid in characterizing this emerging class of biomolecules. Analysis of sequences within the ScrepYard reveals that two-domain tandem repeats constitute the most abundant SCREP domain architecture, while the interdomain "linker" regions connecting the ordered domains are found to be abundant in amino acids with short or polar sidechains and contain an unusually high abundance of proline residues. Finally, we demonstrate the utility of ScrepYard as a virtual screening tool for discovery of putatively multivalent peptides, by using it as a resource to identify a previously uncharacterised serine protease inhibitor and confirm its predicated activity using an enzyme assay.

bioinformatics↗