Search bioRxivSearch

Biology subjects

Ohno-Machado, L.

Publications and source records attributed to Ohno-Machado, L..

3 recordsLinked to original sources

Sharing genetic admixture and diversity of public biomedical datasets

Genetic ancestry and admixture are critical co-factors to study phenotype-genotype associations using cohorts of human subjects. Most publically available molecular datasets - genomes, exomes or transcriptomes - are however missing this information or only share self-reported ancestry. This represents a limitation to identify and re-purpose datasets to investigate the contribution of race and ethnicity to diseases and traits. we propose an analytical framework to enrich the meta-data from publically available cohorts with admixture information and a resulting diversity score at continental resolution, calculated directly from the data. We illustrate the utility and versatility of the framework using The Cancer Genome Atlas datasets indexed and searched through the DataMed Data Discovery Index. Data repositories or data contributors can use this framework to provide, as metadata, admixture for controlled access datasets, minimizing the work involved in requesting a dataset that may ultimately prove inadequate for a researchers purpose. With the increasingly global scale of human genetics research, research on disease risk and susceptibility would benefit greatly from the adequate estimation and sharing of admixture data following a framework such as the one presented.

genetics

DATS: the data tag suite to enable discoverability of datasets

Todays science increasingly requires effective ways to find and access existing datasets that are distributed across a range of repositories. For researchers in the life sciences, discoverability of datasets may soon become as essential as identifying the latest publications via PubMed. Through an international collaborative effort funded by the National Institutes of Health (NIH)s Big Data to Knowledge (BD2K) initiative, we have designed and implemented the DAta Tag Suite (DATS) model to support the DataMed data discovery index. DataMeds goal is to be for data what PubMed has been for the scientific literature. Akin to the Journal Article Tag Suite (JATS) used in PubMed, the DATS model enables submission of metadata on datasets to DataMed. DATS has a core set of elements, which are generic and applicable to any type of datasets, and an extended set that can accommodate more specialized data types. DATS is a platform-independent model also available as a Schema.org annotated serialization to be used beyond DataMed, for example, in projects like DataCite.

bioinformatics

DataMed: Finding useful data across multiple biomedical data repositories

The value of broadening searches for data across multiple repositories has been identified by the biomedical research community. As part of the NIH Big Data to Knowledge initiative, we work with an international community of researchers, service providers and knowledge experts to develop and test a data index and search engine, which are based on metadata extracted from various datasets in a range of repositories. DataMed is designed to be, for data, what PubMed has been for the scientific literature. DataMed supports Findability and Accessibility of datasets. These characteristics - along with Interoperability and Reusability - compose the four FAIR principles to facilitate knowledge discovery in todays big data-intensive science landscape.

bioinformatics