Search bioRxiv⌕ Search

Biology subjects

Matos-Filipe, P.

Publications and source records attributed to Matos-Filipe, P..

2 recordsLinked to original sources

Integration of expression datasets to identify biomarkers for accurate Gleason scoring in Prostate Cancer

Prostate cancer is a significant global health issue with considerable mortality rates, emphasising the urgent need for advanced treatment options and improved diagnostic methods. Current diagnostic standards for prostate cancer, including PSA testing and digital rectal examination, often produce false positives, resulting in unnecessary biopsies for patients. This limitation highlights the critical requirement to incorporate more precise biomarkers to enhance diagnostic accuracy and reduce unnecessary procedures. This study aims to investigate biomarker candidates that can effectively determine prostate cancer aggressiveness. By integrating diverse prostate tissue expression datasets and employing machine-learning techniques, this approach seeks to refine diagnostics and provide insights into the molecular underpinnings of the disease, potentially transforming early detection and patient management strategies. Our proposed biomarkers achieve a minimum precision of 0.80, addressing the false positives limitations associated with classical prostate cancer biomarkers. Moreover, the ROC-AUC profiles of most of the candidates proposed in this study align with those exhibited by other innovative biomarkers recently proposed (ROC-AUC [≥] 0.70). We believe these biomarkers are promising candidates for further in vivo and in vitro investigation.

cancer biology↗

The usage of transcriptomics datasets as sources of Real-World Data for clinical trialling

BackgroundRandomised Clinical Trials (RCT) reflect results within their specific controlled settings, necessitating further studies to understand outcomes across all possible scenarios. The usage of Real-World Data (RWD) has been recently considered to be a viable alternative to overcome these issues and complement clinical conclusions. Molecular profiles of patients captured by high-throughput measures reflect their medical conditions. When this information is linked to clinical and demographical information, nuances in transcriptomics data can uncover subtle variations in disease pathways among distinct patient groups. This work focuses on the construction of a patient repository database with molecular and clinical information resulting from the integration of publicly available transcriptomics datasets. ResultsPatient data were integrated into the patient repository by using a novel post-processing technique allowing for the usage of samples originating from different/multiple Gene Expression Omnibus (GEO) datasets. Our post-processing technique, which we have named MicroArray Cross-plAtfoRm pOst-prOcessiNg (MACAROON), aims to standardise and integrate transcriptomics data (considering batch effects and possible processing-originated artefacts). This process was able to better reproduce the down streaming biological conclusions in a 45% improvement compared to other methods available. Furthermore, RWD was mined from GEO samples metadata and a clinical and demographical characterisation of the database was obtained. RWD mining was done through a manually curated synonym dictionary allowing for the correct assignment (95.33% median accuracy; 80.14% average) of medical conditions. ConclusionsOur strategy produced a repository, which includes molecular, clinical and demographical RWD by integrating multiple public datasets. The exploration of these data facilitates the discovery of clinical outcomes and molecular pathways specific to predetermined patient populations.

bioinformatics↗