Search bioRxiv⌕ Search

Biology subjects

Goh, L. L.

Publications and source records attributed to Goh, L. L..

2 recordsLinked to original sources

Polygenic Risk Scores Across Genomic Platforms for Reliable Breast Cancer Risk Stratification

PurposeWe evaluated differences in a 313-variant breast cancer polygenic risk score (PRS313) across genomic platforms and their impact on risk stratification. MethodsWe compared PRS313 derived from genotyping arrays (Global Screening Array [GSA], OncoArray-500K [OncoArray], Global Diversity Array [GDA], custom Axiom_PrecipV1 array [ThermoFisher]) and low-coverage genome sequencing (lc-WGS) in 2 cell lines and 92 individuals. Probes were designed for all variants on ThermoFisher (success rate: 259/313). Sanger sequencing was performed to profile indels. Concordance of high-risk classification (PRSscore>0.6) across platforms was assessed using Kappa statistics. ResultsPRS313-lc-WGS was identical in the 4 cell line repeats. In saliva samples, indel concordance with Sanger sequencing varied widely (Kappa: 0.007-1.000). PRS313-ThermoFisher was predictable from other platforms using linear models, despite systematic differences. Greater agreement was observed between arrays with high imputation overlap (e.g., GDA[~]GSA slope=0.986). Pre-calibration agreement in high-risk classification was moderate (Fleiss Kappa=0.552) and improved post-calibration (Kappa=0.650). Arrays with similar designs showed higher pre-calibration agreement (Kappa=0.745). Calibration narrowed high-risk proportions from 4-45% to 15-21% -28% were high-risk by any platform, while 8% were high-risk across all five. ConclusionPlatform-specific biases affect PRS interpretation. Calibration enhances consistency in identifying high-risk individuals. STATEMENT OF SIGNIFICANCEThis study compares the performance of a validated 313-variant breast cancer polygenic risk score across platforms, revealing systematic biases in risk stratification and raising concerns about including inconsistent indels in the model.

genetics↗

RAPTOR: A Five-Safes approach to a secure, cloud native and serverless genomics data repository

Genomic researchers are increasingly utilizing commercial cloud platforms (CCPs) to manage their data and analytics needs. Commercial clouds allow researchers to grow their storage and analytics capacity on demand, keeping pace with expanding project data footprints and enabling researchers to avoid large capital expenditures while paying only for IT capacity consumed by their project. Cloud computing also allows researchers to overcome common network and storage bottlenecks encountered when combining or re-analysing large datasets. However, cloud computing presents a new set of challenges. Without adequate security controls, the risk of unauthorised access may be higher for data stored on the cloud. In addition, regulators are increasingly mandating data access patterns and specific security protocols on the storage and use of genomic data to safeguard rights of the study participants. While CCPs provide tools for security and regulatory compliance, utilising these tools to build the necessary controls required for cloud solutions is not trivial as such skill sets are not commonly found in a genomics lab. The Research Assets Provisioning and Tracking Online Repository (RAPTOR) by the Genome Institute of Singapore is a cloud native genomics data repository and analytics platform focusing on security and regulatory compliance. Using a "five-safes" framework (Safe Purpose, Safe People, Safe Settings, Safe Data and Safe Output), RAPTOR provides security and governance controls to data contributors and users leveraging cloud computing for sharing and analysis of large genomic datasets without the risk of security breaches or running afoul of regulations. RAPTOR can also enable data federation with other genomic data repositories using GA4GH community-defined standards, allowing researchers to boost the statistical power of their work and overcome geographic and ancestry limitations of data sets

genomics↗