Search bioRxivSearch

Biology subjects

Thomas, J.

Publications and source records attributed to Thomas, J..

8 recordsLinked to original sources

Genomic and transcriptomic determinants of therapy resistance and immune landscape evolution during anti-EGFR treatment in colorectal cancer

Anti-epidermal growth factor receptor (EGFR) antibodies (anti-EGFR-Ab) are effective in a subgroup of patients with metastatic colorectal cancer (CRC). We applied genomic and transcriptomic analyses to biopsies from 35 RAS wild-type CRCs treated with the anti-EGFR-Ab cetuximab in a prospective trial to interrogate the molecular resistance landscape. This validated transcriptomic CRC-subtypes as predictors of cetuximab benefit; identified novel associations of NF1-inactivation and non-canonical RAS/RAF-aberrations with primary progression; and of FGF10- and non-canonical BRAF-aberrations with AR. No genetic resistance drivers were detected in 64% of AR biopsies. The majority of these had switched from the cetuximab-sensitive CMS2-subtype pretreatment to the fibroblast- and growth factor-rich CMS4-subtype at progression. Fibroblast supernatant conferred cetuximab resistance in vitro, together supporting subtype-switching as a novel mechanism of AR. Cytotoxic immune infiltrates and immune-checkpoint expression increased following cetuximab responses, potentially providing opportunities to treat CRCs with molecularly heterogeneous AR with immunotherapy.

cancer biology

Diverse endogenous retroviruses generate structural variation between human genomes via LTR recombination

Human endogenous retroviruses (HERVs) occupy a substantial fraction of the genome and impact cellular function with both beneficial and deleterious consequences. The vast majority of HERV sequences descend from ancient retroviral families no longer capable of infection or genomic propagation. In fact, most are no longer represented by full-length proviruses but by solitary long terminal repeats (solo LTRs) that arose via non-allelic recombination events between the two LTRs of a proviral insertion. Because LTR-LTR recombination events may occur long after proviral insertion but are challenging to detect in resequencing data, we hypothesize that this mechanism produces an underappreciated amount of genomic variation in the human population. To test this idea, we develop a computational pipeline specifically designed to capture such dimorphic HERV alleles from short-read genome sequencing data. When applied to 279 individuals sequenced as part of the Simons Genome Diversity Project, the pipeline retrieves most of the dimorphic variants previously reported for the HERV-K(HML2) subfamily as well as dozens of additional candidates, including members of the HERV-H and HERV-W families. We experimentally validate several of these candidates, including the first reported instance of an unfixed HERV-W provirus. These data indicate that human proviral content exhibit more extensive interindividual variation than previously recognized. These findings have important implications for our understanding of the contribution of HERVs to human physiology and disease.

genetics

RBPMetaDB: A comprehensive annotation of mouse RNA-Seq datasets with perturbations of RNA-binding proteins

RNA-binding proteins may play a critical role in gene regulation in various diseases or biological processes by controlling post-transcriptional events such as polyadenylation, splicing, and mRNA stabilization via binding activities to RNA molecules. Due to the importance of RNA-binding proteins in gene regulation, a great number of studies have been conducted, resulting in a large amount of RNA-Seq datasets. However, these datasets usually do not have structured organization of metadata, which limits their potentially wide use. To bridge this gap, the metadata of a comprehensive set of publicly available mouse RNA-Seq datasets with perturbed RNA-binding proteins were collected and integrated into a database called RBPMetaDB. This database contains 278 mouse RNA-Seq datasets for a comprehensive list of 163 RNA-binding proteins. These RNA-binding proteins account for only [~]10% of all known RNA-binding proteins annotated in Gene Ontology, indicating that most are still unexplored using high-throughput sequencing. This negative information provides a great pool of candidate RNA-binding proteins for biologists to conduct future experimental studies. In addition, we found that DNA-binding activities are significantly enriched among RNA-binding proteins in RBPMetaDB, suggesting that prior studies of these DNA- and RNA-binding factors focus more on DNA-binding activities instead of RNA-binding activities. This result reveals the opportunity to efficiently reuse these data for investigation of the roles of their RNA-binding activities. A web application has also been implemented to enable easy access and wide use of RBPMetaDB. It is expected that RBPMetaDB will be a great resource for improving understanding of the biological roles of RNA-binding proteins.\n\nDatabase URL: http://rbpmetadb.yubiolab.org

bioinformatics

Animal models of chemotherapy-induced peripheral neuropathy: a machine-assisted systematic review and meta-analysis A comprehensive summary of the field to inform robust experimental design

Background and aimsChemotherapy-induced peripheral neuropathy (CIPN) can be a severely disabling side-effect of commonly used cancer chemotherapeutics, requiring cessation or dose reduction, impacting on survival and quality of life. Our aim was to conduct a systematic review and meta-analysis of research using animal models of CIPN to inform robust experimental design.\n\nMethodsWe systematically searched 5 online databases (PubMed, Web of Science, Citation Index, Biosis Previews and Embase (September 2012) to identify publications reporting in vivo CIPN modelling. Due to the number of publications and high accrual rate of new studies, we ran an updated search November 2015, using machine-learning and text mining to identify relevant studies.\n\nAll data were abstracted by two independent reviewers. For each comparison we calculated a standardised mean difference effect size then combined effects in a random effects meta- analysis. The impact of study design factors and reporting of measures to reduce the risk of bias was assessed. We ran power analysis for the most commonly reported behavioural tests.\n\nResults341 publications were included. The majority (84%) of studies reported using male animals to model CIPN; the most commonly reported strain was Sprague Dawley rat. In modelling experiments, Vincristine was associated with the greatest increase in pain-related behaviour (-3.22 SD [-3.88; -2.56], n=152, p=0). The most commonly reported outcome measure was evoked limb withdrawal to mechanical monofilaments. Pain-related complex behaviours were rarely reported. The number of animals required to obtain 80% power with a significance level of 0.05 varied substantially across behavioural tests. Overall, studies were at moderate risk of bias, with modest reporting of measures to reduce the risk of bias.\n\nConclusionsHere we provide a comprehensive summary of the field of animal models of CIPN and inform robust experimental design by highlighting measures to increase the internal and external validity of studies using animal models of CIPN. Power calculations and other factors, such as clinical relevance, should inform the choice of outcome measure in study design.

neuroscience

Automation of citation screening in pre-clinical systematic reviews

BackgroundThe amount of published in vivo studies and the speed researchers are publishing them make it virtually impossible to follow the recent development in the field. Systematic review emerged as a method to summarise and analyse the studies quantitatively and critically but it is often out-of-date due to its lengthy process. MethodWe invited five machine learning and text-mining groups to build classifiers for identifying publications relevant to neuropathic pain (33814 training publications). We kept 1188 publications for the assessment of the performance of different classifiers. Two groups participated in the next stage: testing their algorithm on datasets labeled for psychosis (11777/2944) and datasets labeled for Vitamin D in multiple sclerosis (train/text: 2038/510). ResultThe performances (sensitive/specificity) of the most promising classifier built for neuropathic pain are: 95%/84%. The performance for psychosis and Vitamin D in multiple sclerosis datasets are 95%/73% and 100%/45%. ConclusionsMachine learning can significantly reduce the irrelevant publications in a systematic review, and save the scientists time and money. Classifier algorithms built for one dataset can be reapplied on another dataset in different field. We are building a machine learning service at the back of Systematic Review & Meta-analysis Facility (SyRF).

scientific communication and education

The use of text-mining and machine learning algorithms in systematic reviews: reducing workload in preclinical biomedical sciences and reducing human screening error

BackgroundHere we outline a method of applying existing machine learning (ML) approaches to aid citation screening in an on-going broad and shallow systematic review of preclinical animal studies, with the aim of achieving a high performing algorithm comparable to human screening.\n\nMethodsWe applied ML approaches to a broad systematic review of animal models of depression at the citation screening stage. We tested two independently developed ML approaches which used different classification models and feature sets. We recorded the performance of the ML approaches on an unseen validation set of papers using sensitivity, specificity and accuracy. We aimed to achieve 95% sensitivity and to maximise specificity. The classification model providing the most accurate predictions was applied to the remaining unseen records in the dataset and will be used in the next stage of the preclinical biomedical sciences systematic review. We used a cross validation technique to assign ML inclusion likelihood scores to the human screened records, to identify potential errors made during the human screening process (error analysis).\n\nResultsML approaches reached 98.7% sensitivity based on learning from a training set of 5749 records, with an inclusion prevalence of 13.2%. The highest level of specificity reached was 86%. Performance was assessed on an independent validation dataset. Human errors in the training and validation sets were successfully identified using assigned the inclusion likelihood from the ML model to highlight discrepancies. Training the ML algorithm on the corrected dataset improved the specificity of the algorithm without compromising sensitivity. Error analysis correction leads to a 3% improvement in sensitivity and specificity, which increases precision and accuracy of the ML algorithm.\n\nConclusionsThis work has confirmed the performance and application of ML algorithms for screening in systematic reviews of preclinical animal studies. It has highlighted the novel use of ML algorithms to identify human error. This needs to be confirmed in other reviews, , but represents a promising approach to integrating human decisions and automation in systematic review methodology.

neuroscience

The Vibrio cholerae Type VI Secretion System Can Modulate Host Intestinal Mechanics to Displace Commensal Gut Bacteria

Host-associated microbiota help defend against bacterial pathogens; the mechanisms that pathogens possess to overcome this defense, however, remain largely unknown. We developed a zebrafish model and used live imaging to directly study how the human pathogen Vibrio cholerae invades the intestine. The gut microbiota of fish mono-colonized by commensal strain Aeromonas veronii was displaced by V. cholerae expressing its Type VI Secretion System (T6SS), a syringe-like apparatus that deploys effector proteins into target cells. Surprisingly, displacement was independent of T6SS-mediated killing of Aeromonas, driven instead by T6SS-induced enhancement of zebrafish intestinal movements that led to expulsion of the resident commensal by the host. Deleting an actin crosslinking domain from the T6SS apparatus returned intestinal motility to normal and thwarted expulsion, without weakening V. cholerae's ability to kill Aeromonas in vitro. Our finding that bacteria can manipulate host physiology to influence inter-microbial competition has implications for both pathogenesis and microbiome engineering.

microbiology

Horizontal Gene Transfer of Functional Type VI Killing Genes by Natural Transformation

Horizontal gene transfer can have profound effects on bacterial evolution by allowing individuals to rapidly acquire adaptive traits that shape their strategies for competition. One strategy for intermicrobial antagonism often used by Proteobacteria is the genetically-encoded contact-dependent Type VI secretion system (T6SS); a weapon used to kill heteroclonal neighbors by direct injection of toxic effectors. Here, we experimentally demonstrate that Vibrio cholerae can acquire new T6SS effector genes via horizontal transfer and utilize them to kill neighboring cells. Replacement of one or more parental alleles with novel effectors allows the recombinant strain to dramatically outcompete its parent. Through spatially-explicit simulation modeling, we show that the HGT is risky: transformation brings a cell into conflict with its former clonemates, but can be adaptive when superior T6SS alleles are acquired. More generally, we find that these costs and benefits are not symmetric, and that high rates of HGT can act as hedge against competitors with unpredictable T6SS efficacy. We conclude that antagonism and horizontal transfer drive successive rounds of weapons-optimization and selective sweeps, dynamically shaping the composition of microbial communities.

microbiology