Search bioRxivSearch

Biology subjects

Brazma, A.

Publications and source records attributed to Brazma, A..

7 recordsLinked to original sources

Assessing the Gene Regulatory Landscape in 1,188 Human Tumors

Cancer is characterised by somatic genetic variation, but the effect of the majority of non-coding somatic variants and the interface with the germline genome are still unknown. We analysed the whole genome and RNA-Seq data from 1,188 human cancer patients as provided by the Pan-cancer Analysis of Whole Genomes (PCAWG) project to map cis expression quantitative trait loci of somatic and germline variation and to uncover the causes of allele-specific expression patterns in human cancers. The availability of the first large-scale dataset with both whole genome and gene expression data enabled us to uncover the effects of the non-coding variation on cancer. In addition to confirming known regulatory effects, we identified novel associations between somatic variation and expression dysregulation, in particular in distal regulatory elements. Finally, we uncovered links between somatic mutational signatures and gene expression changes, including TERT and LMO2, and we explained the inherited risk factors in APOBEC-related mutational processes. This work represents the first large-scale assessment of the effects of both germline and somatic genetic variation on gene expression in cancer and creates a valuable resource cataloguing these effects.

cancer biology

A pan cancer analysis of promoter activity highlights the regulatory role of alternative transcription start sites and their association with noncoding mutations

Most human protein-coding genes are regulated by multiple, distinct promoters, suggesting that the choice of promoter is as important as its level of transcriptional activity. While the role of promoters as driver elements in cancer has been recognized, the contribution of alternative promoters to regulation of the cancer transcriptome remains largely unexplored. Here we infer active promoters using RNA-Seq data from 1,188 cancer samples with matched whole genome sequencing data. We find that alternative promoters are a major contributor to context-specific regulation of isoform expression and that alternative promoters are frequently deregulated in cancer, affecting known cancer-genes and novel candidates. Our study suggests that a highly dynamic landscape of active promoters shapes the cancer transcriptome, opening many opportunities to further explore the interplay of regulatory mechanism and noncoding somatic mutations with transcriptional aberrations in cancer.

genomics

Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes

Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 to 6 MYA, but that are absent in the Hominidae. In fact, Hominidae show between four-and seven-fold lower rates of nucleotide change and feature turnover in both neutral and functional sequences suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. For example, recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli. This process resulted in thousands of novel, species-specific CTCF binding sites. Our results demonstrate that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.

genomics

Comprehensive genome and transcriptome analysis reveals genetic basis for gene fusions in cancer

Gene fusions are an important class of cancer-driving events with therapeutic and diagnostic values, yet their underlying genetic mechanisms have not been systematically characterized. Here by combining RNA and whole genome DNA sequencing data from 1188 donors across 27 cancer types we obtained a list of 3297 high-confidence tumour-specific gene fusions, 82% of which had structural variant (SV) support and 2372 of which were novel. Such a large collection of RNA and DNA alterations provides the first opportunity to systematically classify the gene fusions at a mechanistic level. While many could be explained by single SVs, numerous fusions involved series of structural rearrangements and thus are composite fusions. We discovered 75 fusions of a novel class of inter-chromosomal composite fusions, termed bridged fusions, in which a third genomic location bridged two different genes. In addition, we identified 522 fusions involving non-coding genes and 157 ORF-retaining fusions, in which the complete open reading frame of one gene was fused to the UTR region of another. Although only a small proportion (5%) of the discovered fusions were recurrent, we found a set of highly recurrent fusion partner genes, which exhibited strong 5 or 3 bias and were significantly enriched for cancer genes. Our findings broaden the view of the gene fusion landscape and reveal the general properties of genetic alterations underlying gene fusions for the first time.

genetics

An integrated genomic analysis of anaplastic meningioma identifies prognostic molecular signatures

Anaplastic meningioma is a rare and aggressive brain tumor characterised by intractable recurrences and dismal outcomes. Here, we present an integrated analysis of the whole genome, transcriptome and methylation profiles of primary and recurrent anaplastic meningioma. A key finding was the delineation of two distinct molecular subgroups that were associated with diametrically opposed survival outcomes. Relative to lower grade meningiomas, anaplastic tumors harbored frequent driver mutations in SWI/SNF complex genes, which were confined to the poor prognosis subgroup. Our analyses discern two biologically distinct variants of anaplastic meningioma with potential prognostic and therapeutic significance.

genomics

Genomic determinants of protein abundance variation in colorectal cancer cells

Assessing the extent to which genomic alterations compromise the integrity of the proteome is fundamental in identifying the mechanisms that shape cancer heterogeneity. We have used isobaric labelling and tribrid mass spectrometry to characterize the proteomic landscapes of 50 colorectal cancer cell lines and to decipher the relationships between genomic and proteomic variation. The robust quantification of 12,000 proteins and 27,000 phosphopeptides revealed how protein symbiosis translates to a co-variome which is subjected to a hierarchical order and exposes the collateral effects of somatic mutations on protein complexes. Targeted depletion of key chromatin modifiers confirmed the transmission of variation and the directionality as characteristics of protein interactions. Protein level variation was leveraged to build drug response predictive models towards a better understanding of pharmacoproteomic interactions in colorectal cancer. Overall, we provide a deep integrative view of the molecular structure underlying the variation of colorectal cancer cells.\n\nHighlightsO_LIThe cancer cell functional \"co-variome\" is a strong attribute of the proteome.\nC_LIO_LIMutations can have a direct impact on protein levels of chromatin modifiers.\nC_LIO_LITransmission of genomic variation is a characteristic of protein interactions.\nC_LIO_LIPharmacoproteomic models are strong predictors of response to DNA damaging agents.\nC_LI\n\nAbbreviations

systems biology

The Image Data Resource: A Scalable Platform for Biological Image Data Access, Integration, and Dissemination

Access to primary research data is vital for the advancement of science. To extend the data types supported by community repositories, we built a prototype Image Data Resource (IDR) that collects and integrates imaging data acquired across many different imaging modalities. IDR links high-content screening, super-resolution microscopy, time-lapse and digital pathology imaging experiments to public genetic or chemical databases, and to cell and tissue phenotypes expressed using controlled ontologies. Using this integration, IDR facilitates the analysis of gene networks and reveals functional interactions that are inaccessible to individual studies. To enable re-analysis, we also established a computational resource based on IPython notebooks that allows remote access to the entire IDR. IDR is also an open source platform that others can use to publish their own image data. Thus IDR provides both a novel on-line resource and a software infrastructure that promotes and extends publication and re-analysis of scientific image data.

bioinformatics