Search bioRxivSearch

Biology subjects

McMurry, J.

Publications and source records attributed to McMurry, J..

4 recordsLinked to original sources

A harmonized meta-knowledgebase of clinical interpretations of cancer genomic variants

Precision oncology relies on the accurate discovery and interpretation of genomic variants to enable individualized diagnosis, prognosis, and therapy selection. We found that knowledgebases containing clinical interpretations of somatic cancer variants are highly disparate in interpretation content, structure, and supporting primary literature, impeding consensus when evaluating variants and their relevance in a clinical setting. With the cooperation of experts of the Global Alliance for Genomics and Health (GA4GH) and six prominent cancer variant knowledgebases, we developed a framework for aggregating and harmonizing variant interpretations to produce a meta-knowledgebase of 12,856 aggregate interpretations covering 3,437 unique variants in 415 genes, 357 diseases, and 791 drugs. We demonstrated large gains in overlap between resources across variants, diseases, and drugs as a result of this harmonization. We subsequently demonstrated improved matching between a patient cohort and harmonized interpretations of potential clinical significance, observing an increase from an average of 33% per individual knowledgebase to 56% in aggregate. Our analyses illuminate the need for open, interoperable sharing of variant interpretation data. We also provide an open and freely available web interface (search.cancervariants.org) for exploring the harmonized interpretations from these six knowledgebases.

bioinformatics

A Measure of Open Data: A Metric and Analysis of Reusable Data Practices in Biomedical Data Resources

Data is the foundation of science, and there is an increasing focus on how data can be reused and enhanced to drive scientific discoveries. However, most seemingly \"open data\" do not provide legal permissions for reuse and redistribution. Not being able to integrate and redistribute our collective data resources blocks innovation, and stymies the creation of life-improving diagnostic and drug selection tools. To help the biomedical research and research support communities (e.g. libraries, funders, repositories, etc.) understand and navigate the data licensing landscape, the (Re)usable Data Project (RDP) (http://reusabledata.org) assesses the licensing characteristics of data resources and how licensing behaviors impact reuse. We have created a ruleset to determine the reusability of data resources and have applied it to 56 scientific data resources (i.e. databases) to date. The results show significant reuse and interoperability barriers. Inspired by game-changing projects like Creative Commons, the Wikipedia Foundation, and the Free Software movement, we hope to engage the scientific community in the discussion regarding the legal use and reuse of scientific data, including the balance of openness and how to create sustainable data resources in an increasingly competitive environment.

scientific communication and education

Identifiers for the 21st century:How to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data

In many disciplines, data is highly decentralized across thousands of online databases (repositories, registries, and knowledgebases). Wringing value from such databases depends on the discipline of data science and on the humble bricks and mortar that make integration possible; identifiers are a core component of this integration infrastructure. Drawing on our experience and on work by other groups, we outline ten lessons we have learned about the identifier qualities and best practices that facilitate large-scale data integration. Specifically, we propose actions that identifier practitioners (database providers) should take in the design, provision and reuse of identifiers; we also outline important considerations for those referencing identifiers in various circumstances, including by authors and data generators. While the importance and relevance of each lesson will vary by context, there is a need for increased awareness about how to avoid and manage common identifier problems, especially those related to persistence and web-accessibility/resolvability. We focus strongly on web-based identifiers in the life sciences; however, the principles are broadly relevant to other disciplines.

bioinformatics

The Monarch Initiative: Insights across species reveal human disease mechanisms

The principles of genetics apply across the whole tree of life: on a cellular level, we share mechanisms with species from which we diverged millions or even billions of years ago. We can exploit this common ancestry at the level of sequences, but also in terms of observable outcomes (phenotypes), to learn more about health and disease for humans and all other species. Applying the range of available knowledge to solve challenging disease problems requires unified data relating genomics, phenotypes, and disease; it also requires computational tools that leverage these multimodal data to inform interpretations by geneticists and to suggest experiments. However, the distribution and heterogeneity of databases is a major impediment: databases tend to focus either on a single data type across species, or on single species across data types. Although each database provides rich, high-quality information, no single one provides unified data that is comprehensive across species, biological scales, and data types. Without a big-picture view of the data, many questions in genetics are difficult or impossible to answer. The Monarch Initiative (https://monarchinitiative.org) is an international consortium dedicated to providing computational tools that leverage a computational representation of phenotypic data for genotype-phenotype analysis, genomic diagnostics, and precision medicine on the basis of a large-scale platform of multimodal data that is deeply integrated across species and covering broad areas of disease.

evolutionary biology