Search bioRxivSearch

Biology subjects

Vasilevsky, N.

Publications and source records attributed to Vasilevsky, N..

2 recordsLinked to original sources

Teaching data science fundamentals through realistic synthetic clinical cardiovascular data

ObjectiveOur goal was to create a synthetic dataset and curricular materials to assist in teaching fundamentals of translational data science.\n\nMaterials and MethodsA literature review was conducted to extract current cardiovascular risk score logic, data elements, and population characteristics. Then, clinical data elements in the models were pulled from clinical data and transformed to the Observational Medical Outcomes Partnership (OMOP) common data model; genetic data elements were added based on population rates. A hybrid Bayesian network was used to create synthetic data from the logical elements of the risk scores and the underlying population frequencies of the clinical data.\n\nResultsA synthetic dataset of 446,000 patients was created. A two-day curriculum was created based on this synthetic data with exploratory data analysis and machine learning components. The curriculum was offered on two separate occasions; the two groups of learners were given the curriculum and data, and results were tallied, summarized, and compared. Students ability to complete the challenge was mixed; more experienced students achieved a range of 70%-85% in balanced accuracy, but many others did not perform better than the baseline model.\n\nDiscussionOverall, students enjoyed the course and dataset, but some struggled to consistently apply machine learning techniques. The curriculum, data set, techniques for generation, and results are available for others to use for their own training.\n\nConclusionA realistic synthetic data with clinical and genetic components helps students learn issues in cardiovascular risk scoring, practice data science skills, and compete in a challenge to improve identification of risk.

scientific communication and education

The Monarch Initiative: Insights across species reveal human disease mechanisms

The principles of genetics apply across the whole tree of life: on a cellular level, we share mechanisms with species from which we diverged millions or even billions of years ago. We can exploit this common ancestry at the level of sequences, but also in terms of observable outcomes (phenotypes), to learn more about health and disease for humans and all other species. Applying the range of available knowledge to solve challenging disease problems requires unified data relating genomics, phenotypes, and disease; it also requires computational tools that leverage these multimodal data to inform interpretations by geneticists and to suggest experiments. However, the distribution and heterogeneity of databases is a major impediment: databases tend to focus either on a single data type across species, or on single species across data types. Although each database provides rich, high-quality information, no single one provides unified data that is comprehensive across species, biological scales, and data types. Without a big-picture view of the data, many questions in genetics are difficult or impossible to answer. The Monarch Initiative (https://monarchinitiative.org) is an international consortium dedicated to providing computational tools that leverage a computational representation of phenotypic data for genotype-phenotype analysis, genomic diagnostics, and precision medicine on the basis of a large-scale platform of multimodal data that is deeply integrated across species and covering broad areas of disease.

evolutionary biology