Search bioRxivSearch

Biology subjects

Fenyo, D.

Publications and source records attributed to Fenyo, D..

12 recordsLinked to original sources

Quantitative mass spectrometry to interrogate proteomic heterogeneity in metastatic lung adenocarcinoma and validate a novel somatic mutation CDK12-G879V

Lung cancer is the leading cause of cancer death both in men and women. Tumor heterogeneity is an impediment to targeted treatment of all cancers, including lung cancer. Here, we sought to characterize changes in tumor proteome and phosphoproteome by longitudinal, prospective collection of tumor tissue of an exceptional responder lung adenocarcinoma patient who survived with metastatic lung adenocarcinoma for more than seven years with HER2-directed therapy in combination with chemotherapy. We employed \"Super-SILAC\" and TMT labeling strategies to quantify the proteome and phosphoproteome of a lung metastatic site and ten different metastatic progressive lymph nodes collected across a span of seven years, including five lymph nodes procured at autopsy. We identified specific signaling networks enriched in lung compared to the lymph node metastatic sites. We correlated the changes in protein abundance with changes in copy number alteration (CNA) and transcript expression. To further interrogate the mass spectrometry data, patient-specific database was built incorporating all the somatic variants identified by whole genome sequencing (WGS) of genomic DNA from the lung, one lymph node metastatic site and blood. An extensive validation pipeline was built for confirmation of variant peptides. We validated 360 spectra corresponding to 55 germline and 6 somatic variant peptides. Targeted MRM assays demonstrated expression of two novel variant somatic peptides, CDK12-G879V and FASN-R1439Q, with expression in lung and lymph node metastatic sites, respectively. CDK12 G879V mutation likely results in a nonfunctional CDK12 kinase and chemotherapy susceptibility in lung metastatic sites. Knockdown of CDK12 in lung adenocarcinoma cells results in increased chemotherapy sensitivity, explaining the complete resolution of the lung metastatic sites in this patient.

cancer biology

Transcription drives DNA replication initiation and termination in human cells

The locations of active DNA replication origins in the human genome, and the determinants of origin activation, remain controversial. Additionally, neither the predominant sites of replication termination nor the impact of transcription on replication-fork mobility have been defined. We demonstrate that replication initiation occurs preferentially in the immediate vicinity of the transcription start site of genes occupied by high levels of RNA polymerase II, ensuring co-directional replication of the most highly transcribed genes. Further, we demonstrate that dormant replication origin firing represents the global activation of pre-existing origins. We also show that DNA replication naturally terminates at the polyadenylation site of transcribed genes. During replication stress, termination is redistributed to gene bodies, generating a global reorientation of replication relative to transcription. Our analysis provides a unified model for the coupling of transcription with replication initiation and termination in human cells.

molecular biology

Integration and analysis of CPTAC proteomics data in the context of cancer genomics in the cBioPortal

The Clinical Proteomic Tumor Analysis Consortium (CPTAC) has produced extensive mass spectrometry based proteomics data for selected breast, colon and ovarian tumors from The Cancer Genome Atlas (TCGA). We have incorporated the CPTAC proteomics data into the cBioPotal to support easy exploration and integrative analysis of these proteomic datasets in the context of the clinical and genomics data from the same tumors. cBioPortal is an open source platform for exploring, visualizing, and analyzing multi-dimensional cancer genomics and clinical data. The public instance of the cBioPortal (http://cbioportal.org/) hosts more than 100 cancer genomics studies including all of the data from TCGA. Its biologist-friendly interface provides many rich analysis features, including a graphical summary of gene-level data across multiple platforms, correlation analysis between genes or other data types, survival analysis, and network visualization. Here, we present the integration of the CPTAC mass spectrometry based proteomics data into the cBioPortal, consisting of 77 breast, 95 colorectal, and 174 ovarian tumors that already have been profiled by TCGA for mutations, copy number alterations, gene expression, and DNA methylation. As a result, the CPTAC data can now be easily explored and analyzed in the cBioPortal in the context of clinical and genomics data. By integrating CPTAC data into cBioPortal, limitations of TCGA proteomics array data can be overcome while also providing a user-friendly web interface, a web API and an R client to query the mass spectrometry data together with genomic, epigenomic, and clinical data.

cancer biology

Classification and Mutation Prediction from Non-Small Cell Lung Cancer Histopathology Images using Deep Learning

Visual analysis of histopathology slides of lung cell tissues is one of the main methods used by pathologists to assess the stage, types and sub-types of lung cancers. Adenocarcinoma and squamous cell carcinoma are two most prevalent sub-types of lung cancer, but their distinction can be challenging and time-consuming even for the expert eye. In this study, we trained a deep learning convolutional neural network (CNN) model (inception v3) on histopathology images obtained from The Cancer Genome Atlas (TCGA) to accurately classify whole-slide pathology images into adenocarcinoma, squamous cell carcinoma or normal lung tissue. Our method slightly outperforms a human pathologist, achieving better sensitivity and specificity, with [~]0.97 average Area Under the Curve (AUC) on a held-out population of whole-slide scans. Furthermore, we trained the neural network to predict the ten most commonly mutated genes in lung adenocarcinoma. We found that six of these genes - STK11, EGFR, FAT1, SETBP1, KRAS and TP53 - can be predicted from pathology images with an accuracy ranging from 0.733 to 0.856, as measured by the AUC on the held-out population. These findings suggest that deep learning models can offer both specialists and patients a fast, accurate and inexpensive detection of cancer types or gene mutations, and thus have a significant impact on cancer treatment.

cancer biology

A Method for Quantifying Molecular Interactions Using Stochastic Modelling and Super-Resolution Microscopy

We introduce the Interaction Factor (IF), a measure for quantifying the interaction of molecular clusters in super-resolution microscopy images. The IF is robust in the sense that it is independent of cluster density, and it only depends on the extent of the pair-wise interaction between different types of molecular clusters in the image. The IF for a single or a collection of images is estimated by first using stochastic modelling where the locations of clusters in the images are repeatedly randomized to estimate the distribution of the overlaps between the clusters in the absence of interaction (IF=0). Second, an analytical form of the relationship between IF and the overlap (which has the random overlap as its only parameter) is used to estimate the IF for the experimentally observed overlap. The advantage of IF compared to conventional methods to quantify interaction in microscopy images is that it is insensitive to changing cluster density and is an absolute measure of interaction, making the interpretation of experiments easier. We validate the IF method by using both simulated and experimental data and provide an ImageJ plugin for determining the IF of an image.

bioinformatics

Pathway-Level Integration of Proteogenomic data in Breast Cancer Using Independent Component Analysis

Recent advances in the multi-omics characterization necessitate pathway-level abstraction and knowledge integration across different data types. In this study, we apply independent component analysis (ICA) to human breast cancer proteogenomics data to retrieve mechanistic information. We show that as an unsupervised feature extraction method, ICA was able to construct signatures with known biological relevance on both transcriptome and proteome levels. Moreover, proteome and transcriptome signatures can be associated by their respective correlation with patient clinical features, providing an integrated description of phenotype-related biological processes. Our results demonstrate that the application of ICA to proteogenomics data could lead to pathway-level knowledge discovery. Potential extension of this approach to other data and cancer types may contribute to pan-cancer integration of multi-omics information.

cancer biology

The largest SWI/SNF polyglutamine domain is a pH sensor

Polyglutamines are known to form aggregates in pathogenic contexts, such as in Huntingtons disease, however little is known about their role in normal biological processes. We found that a polyglutamine domain in the SNF5 subunit of the yeast SWI/SNF complex, histidines within this sequence, and transient intracellular acidification are required for efficient transcriptional regulation during carbon starvation. We hypothesized that a pH-driven oligomerization of the SNF5 polyglutamine region is required for transcriptional reprogramming. In support of this idea, we found that a synthetic spidroin domain from spider silk, which is soluble at pH 7 but oligomerizes at pH ~ 6.3, could partially complement the function of the SNF5 polyglutamine. These results suggest that the SNF5 polyglutamine domain acts as a pH-driven transcriptional regulator.

genetics

LINE-1 and the cell cycle: protein localization and functional dynamics

LINE-1/L1 retrotransposon sequences comprise 17% of the human genome. Among the many classes of mobile genetic elements, L1 is the only autonomous retrotransposon that still drives human genomic plasticity today. Through its co-evolution with the human genome, L1 has intertwined itself with host cell biology to aid its proliferation. However, a clear understanding of L1s lifecycle and the processes involved in restricting its insertion and its intragenomic spreading remains elusive. Here we identify modes of L1 proteins entrance into the nucleus, a necessary step for L1 proliferation. Using functional, biochemical, and imaging approaches, we also show a clear cell cycle bias for L1 retrotransposition that peaks during the S phase. Our observations provide a basis for novel interpretations about the nature of nuclear and cytoplasmic L1 ribonucleoproteins (RNPs) and the potential role of DNA replication in L1 retrotransposition.

cell biology

Dissection of purified LINE-1 reveals distinct nuclear and cytoplasmic intermediates

1.Long Interspersed Nuclear Element-1 (LINE-1, L1) is a mobile genetic element active in human genomes. L1-encoded ORF1 and ORF2 proteins bind L1 RNAs, forming ribonucleoproteins (RNPs). These RNPs interact with diverse host proteins, some repressive and others required for the L1 lifecycle. Using differential affinity purifications and quantitative mass spectrometry, we have characterized the proteins associated with distinctive L1 macromolecular complexes. Our findings support the presence of multiple L1-derived retrotransposition intermediates in vivo. Among them, we describe a cytoplasmic intermediate that we hypothesize to be the canonical ORF1p/ORF2p/L1-RNA-containing RNP, and we describe a nuclear population containing ORF2p, but lacking ORF1p, which likely contains host factors participating in template-primed reverse transcription.

molecular biology

The proBAM and proBed standard formats: enabling a seamless integration of genomics and proteomics data.

On behalf of The Human Proteome Organization (HUPO) Proteomics Standards Initiative (PSI), we are here introducing two novel standard data formats, proBAM and proBed, that have been developed to address the current challenges of integrating mass spectrometry based proteomics data with genomics and transcriptomics information in proteogenomics studies. proBAM and proBed are adaptations from the well-defined, widely used file formats SAM/BAM and BED respectively, and both have been extended to meet specific requirements entailed by proteomics data. Therefore, existing popular genomics tools such as SAMtools and Bedtools, and several very popular genome browsers, can be used to manipulate and visualize these formats already out-of-the-box. We also highlight that a number of specific additional software tools, properly supporting the proteomics information available in these formats, are now available providing functionalities such as file generation, file conversion, and data analysis. All the related documentation to the formats, including the detailed file format specifications, and example files are accessible at http://www.psidev.info/probam and http://www.psidev.info/probed.

bioinformatics

Human to yeast pathway transplantation: cross-species dissection of the adenine de novo pathway regulatory node

Pathway transplantation from one organism to another represents a means to a more complete understanding of a biochemical or regulatory process. The purine biosynthesis pathway, a core metabolic function, was transplanted from human to yeast. We replaced the entire Saccharomyces cerevisiae adenine de novo pathway with the cognate human pathway components. A yeast strain was \"humanized\" for the full pathway by deleting all relevant yeast genes completely and then providing the human pathway in trans using a neochromosome expressing the human protein coding regions under the transcriptional control of their cognate yeast promoters and terminators. The \"humanized\" yeast strain grows in the absence of adenine, indicating complementation of the yeast pathway by the full set of human proteins. While the strain with the neochromosome is indeed prototrophic, it grows slowly in the absence of adenine. Dissection of the phenotype revealed that the human ortholog of ADE4, PPAT, shows only partial complementation. We have used several strategies to understand this phenotype, that point to PPAT/ADE4 as the central regulatory node. Pathway metabolites are responsible for regulating PPATs protein abundance through transcription and proteolysis as well as its enzymatic activity by allosteric regulation in these yeast cells. Extensive phylogenetic analysis of PPATs from diverse organisms hints at adaptations of the enzyme-level regulation to the metabolite levels in the organism. Finally, we isolated specific mutations in PPAT as well as in other genes involved in the purine metabolic network that alleviate incomplete complementation by PPAT and provide further insight into the complex regulation of this critical metabolic pathway.

synthetic biology

A toolbox of immunoprecipitation-grade monoclonal antibodies against human transcription factors.

A key component to overcoming the reproducibility crisis in biomedical research is the development of readily available, rigorously validated and renewable protein affinity reagents. As part of the NIH Protein Capture Reagents Program (PCRP), we have generated a collection of 1406 highly validated, immunoprecipitation (IP) and/or immunoblotting (IB) grade, mouse monoclonal antibodies (mAbs) to 736 human transcription factors. We used HuProt human protein microarrays to identify mAbs that recognize their cognate targets with exceptional specificity. Using an integrated production and validation pipeline, we validated these mAbs in multiple experimental applications, and have distributed them to the Developmental Studies Hybridoma Bank (DSHB) and several commercial suppliers. This study allowed us to perform a meta-analysis that identified critical variables that contribute to the generation of high quality mAbs. We find that using full-length antigens for immunization, in combination with HuProt analysis, provides the highest overall success rates. The efficiencies built into this pipeline ensure substantial cost savings compared to current standard practices.

biochemistry