Search bioRxivSearch

Biology subjects

Jones, S. J.

Publications and source records attributed to Jones, S. J..

4 recordsLinked to original sources

CancerMine: A literature-mined resource for drivers, oncogenes and tumor suppressors in cancer

Understanding a mutation in cancer requires knowledge of the different roles that genes play in cancer as drivers, oncogenes and tumor suppressors. We present CancerMine, a high-quality text-mined knowledgebase that catalogues over 856 genes as drivers, 2,421 as oncogenes and 2,037 as tumor suppressors in 426 cancer types. We compile 3,485 genes that are not in the IntOGen resource of drivers and complement the Cancer Gene Census with 3,136 new genes identified as oncogenes and tumor suppressors. CancerMine provides a method for gene-centric clustering of cancer types illustrating genetic similarities between cancer types of different organs and was validated against data from the Cancer Genome Atlas (TCGA) project. Finally with 178 novel cancer gene mentions in publications each month, this resource will be updated monthly, pre-empting the need to manually curate the ever-increasing number of novel cancer associated genes. CancerMine is viewable through a web portal (http://bionlp.bcgsc.ca/cancermine/) and available for download (https://github.com/jakelever/cancermine).

bioinformatics

Tigmint: Correcting Assembly Errors Using Linked Reads From Large Molecules

Genome sequencing yields the sequence of many short snippets of DNA (reads) from a genome. Genome assembly attempts to reconstruct the original genome from which these reads were derived. This task is difficult due to gaps and errors in the sequencing data, repetitive sequence in the underlying genome, and heterozygosity, and assembly errors are common. These misassemblies may be identified by comparing the sequencing data to the assembly, and by looking for discrepancies between the two. Once identified, these misassemblies may be corrected, improving the quality of the assembly. Although tools exist to identify and correct misassemblies using Illumina pair-end and mate-pair sequencing, no such tool yet exists that makes use of the long distance information of the large molecules provided by linked reads, such as those offered by the 10x Genomics Chromium platform. We have developed the tool Tigmint for this purpose. To demonstrate the effectiveness of Tigmint, we corrected assemblies of a human genome using short reads assembled with ABySS 2.0 and other assemblers. Tigmint reduced the number of misassemblies identified by QUAST in the ABySS assembly by 216 (27%). While scaffolding with ARCS alone more than doubled the scaffold NGA50 of the assembly from 3 to 8 Mbp, the combination of Tigmint and ARCS improved the scaffold NGA50 of the assembly over five-fold to 16.4 Mbp. This notable improvement in contiguity highlights the utility of assembly correction in refining assemblies. We demonstrate its usefulness in correcting the assemblies of multiple tools, as well as in using Chromium reads to correct and scaffold assemblies of long single-molecule sequencing. The source code of Tigmint is available for download from https://github.com/bcgsc/tigmint, and is distributed under the GNU GPL v3.0 license.

genomics

Simulating Pedigrees Ascertained for Multiple Disease-Affected Relatives

Background: Studies that ascertain families containing multiple relatives affected by disease can be useful for identification of causal, rare variants from next-generation sequencing data.\n\nResults: We present the R package SimRVPedigree, which allows researchers to simulate pedigrees ascertained on the basis of multiple, affected relatives. By incorporating the ascertainment process in the simulation, SimRVPedigree allows researchers to better understand the within-family patterns of relationship amongst affected individuals and ages of disease onset.\n\nConclusions: Through simulation, we show that affected members of a family segregating a rare disease variant tend to be more numerous and cluster in relationships more closely than those for sporadic disease. We also show that the family ascertainment process can lead to apparent anticipation in the age of onset. Finally, we use simulation to gain insight into the limit on the proportion of ascertained families segregating a causal variant. SimRVPedigree should be useful to investigators seeking insight into the family-based study design through simulation.

genetics

Conserved roles of RECQ-like helicases Sgs1 and BLM in preventing R-loop induced genome instability

Sgs1 is a yeast DNA helicase functioning in DNA replication and repair, and is the orthologue of the human Blooms syndrome helicase BLM. Here we analyze the mutation signature associated with SGS1 deletion in yeast, and find frequent copy number changes flanked by regions of repetitive sequence and high R-loop forming potential. We show that loss of SGS1 increases R-loop accumulation and sensitizes cells to replication-transcription collisions. Accordingly, in sgs1{Delta} cells the genome-wide distribution of R-loops shifts to known sites of Sgs1 action, replication pausing regions, and to long genes. Depletion of the orthologous BLM helicase from human cancer cells also increases R-loop levels, and R-loop-associated genome instability. In support of a direct effect, BLM is found physically proximal to DNA:RNA hybrids in human cells, and can efficiently unwind R-loops in vitro. Together our data describe a conserved role for Sgs1/BLM in R-loop suppression and support an increasingly broad view of DNA repair and replication fork stabilizing proteins as modulators of R-loop mediated genome instability.

cell biology