Search bioRxiv⌕ Search

Biology subjects

de Souza, V.

Publications and source records attributed to de Souza, V..

4 recordsLinked to original sources

The Unified Phenotype Ontology (uPheno): A framework for cross-species integrative phenomics

Phenotypic data are critical for understanding biological mechanisms and consequences of genomic variation, and are pivotal for clinical use cases such as disease diagnostics and treatment development. For over a century, vast quantities of phenotype data have been collected in many different contexts covering a variety of organisms. The emerging field of phenomics focuses on integrating and interpreting these data to inform biological hypotheses. A major impediment in phenomics is the wide range of distinct and disconnected approaches to recording the observable characteristics of an organism. Phenotype data are collected and curated using free text, single terms or combinations of terms, using multiple vocabularies, terminologies, or ontologies. Integrating these heterogeneous and often siloed data enables the application of biological knowledge both within and across species. Existing integration efforts are typically limited to mappings between pairs of terminologies; a generic knowledge representation that captures the full range of cross-species phenomics data is much needed. We have developed the Unified Phenotype Ontology (uPheno) framework, a community effort to provide an integration layer over domain-specific phenotype ontologies, as a single, unified, logical representation. uPheno comprises (1) a system for consistent computational definition of phenotype terms using ontology design patterns, maintained as a community library; (2) a hierarchical vocabulary of species-neutral phenotype terms under which their species-specific counterparts are grouped; and (3) mapping tables between species-specific ontologies. This harmonized representation supports use cases such as cross-species integration of genotype-phenotype associations from different organisms and cross-species informed variant prioritization.

bioinformatics↗

Towards a standard benchmark for variant and gene prioritisation algorithms: PhEval - Phenotypic inference Evaluation framework

BackgroundComputational approaches to support rare disease diagnosis are challenging to build, requiring the integration of complex data types such as ontologies, gene-to-phenotype associations, and cross-species data into variant and gene prioritisation algorithms (VGPAs). However, the performance of VGPAs has been difficult to measure and is impacted by many factors, for example, ontology structure, annotation completeness or changes to the underlying algorithm. Assertions of the capabilities of VGPAs are often not reproducible, in part because there is no standardised, empirical framework and openly available patient data to assess the efficacy of VGPAs - ultimately hindering the development of effective prioritisation tools. ResultsIn this paper, we present our benchmarking tool, PhEval, which aims to provide a standardised and empirical framework to evaluate phenotype-driven VGPAs. The inclusion of standardised test corpora and test corpus generation tools in the PhEval suite of tools allows open benchmarking and comparison of methods on standardised data sets. ConclusionsPhEval and the standardised test corpora solve the issues of patient data availability and experimental tooling configuration when benchmarking and comparing rare disease VGPAs. By providing standardised data on patient cohorts from real-world case-reports and controlling the configuration of evaluated VGPAs, PhEval enables transparent, portable, comparable and reproducible benchmarking of VGPAs. As these tools are often a key component of many rare disease diagnostic pipelines, a thorough and standardised method of assessment is essential for improving patient diagnosis and care.

bioinformatics↗

GARSA: An integrative pipeline for genome wide association studies and polygenic risk score inference in admixed human populations

Genome-wide association studies (GWAS) and polygenic risk scores (PRS) are multistep analytical tools to identify genetic variants and to assess their contribution to phenotypes/diseases. These analyses are evolving and becoming instrumental to understand the genetic architecture of complex phenotypes/diseases. Nevertheless, to date, there is no single solution incorporating all major steps related to those analyses combined with robust populational bias correction. Here, we describe a semi-automated pipeline unifying steps involved in GWAS and PRS including widely used software. Our pipeline handles quality control (QC), GWAS, and PRS steps, managing different types of input/output files. Furthermore, it includes robust bias correction steps, such as inference of kinship matrix with correction for population structure, use of principal component analysis (PCA) with detection and removal of outlier variant followed by re-projection of related individuals (if desired), generation of PCA figures that assist in setting the best number of principal components (PCs) for association analysis, availability of mixed models, use of recommended software for GWAS based on population size, and a Markov chain Monte Carlo (MCMC) method to estimate best set of PRS parameters. Finally, we tested GARSA pipeline in a family-based Brazilian admixed population and demonstrated that the corrections implemented indeed mitigate bias in downstream analysis. The pipeline can be implemented on personal or server-side environments. AvailabilityThe development version (open-source) is available in https://github.com/LGCM-OpenSource/GARSA ContactFernando P. N. Rossi - fernando.rossi@hc.fm.usp.br; Jose S. L. Patane - jose.patane@hc.fm.usp.br Supplementary informationSupplementary tutorial.

bioinformatics↗

The Ontology of Biological Attributes (OBA) - Computational Traits for the Life Sciences

Existing phenotype ontologies were originally developed to represent phenotypes that manifest as a character state in relation to a wild-type or other reference. However, these do not include the phenotypic trait or attribute categories required for the annotation of genome-wide association studies (GWAS), Quantitative Trait Loci (QTL) mappings or any population-focused measurable trait data. Moreover, variations in gene expression in response to environmental disturbances even without any genetic alterations can also be associated with particular biological attributes. The integration of trait and biological attribute information with an ever increasing body of chemical, environmental and biological data greatly facilitates computational analyses and it is also highly relevant to biomedical and clinical applications. The Ontology of Biological Attributes (OBA) is a formalised, species-independent collection of interoperable phenotypic trait categories that is intended to fulfil a data integration role. OBA is a standardised representational framework for observable attributes that are characteristics of biological entities, organisms, or parts of organisms. OBA has a modular design which provides several benefits for users and data integrators, including an automated and meaningful classification of trait terms computed on the basis of logical inferences drawn from domain-specific ontologies for cells, anatomical and other relevant entities. The logical axioms in OBA also provide a previously missing bridge that can computationally link Mendelian phenotypes with GWAS and quantitative traits. The term components in OBA provide semantic links and enable knowledge and data integration across specialised research community boundaries, thereby breaking silos.

bioinformatics↗