Search bioRxiv⌕ Search

Biology subjects

Syed, H.

Publications and source records attributed to Syed, H..

4 recordsLinked to original sources

The Use of Class Imbalanced Learning Methods on ULSAM Data to Predict the Case-Control Status in Genome-Wide Association Studies

Machine learning (ML) methods for uncovering single nucleotide polymorphisms (SNPs) in genome-wide association study (GWAS) data that can be used to predict disease outcomes are becoming increasingly used in genetic research. Two issues with the use of ML models are finding the correct method for dealing with imbalanced data and data training. This article compares three ML models to identify SNPs that predict type 2 diabetes (T2D) status using the Support vector machine SMOTE (SVM SMOTE), The Adaptive Synthetic Sampling Approach (ADASYN), Random under sampling (RUS) on GWAS data from elderly male participants (165 cases and 951 controls) from the Uppsala Longitudinal Study of Adult Men (ULSAM). It was also applied to SNPs selected by the SMOTE, SVM SMOTE, ADASYN, and RUS clumping method. The analysis was performed using three different ML models: (i) support vector machine (SVM), (ii) multilayer perceptron (MLP) and (iii) random forests (RF). The accuracy of the case-control classification was compared between these three methods. The best classification algorithm was a combination of MLP and SMOTE (97% accuracy). Both RF and SVM achieved good accuracy results of over 90%. Overall, methods used against unbalanced data, all three ML algorithms were found to improve prediction accuracy.

bioinformatics↗

Epigenetic-focused CRISPR/Cas9 screen identifies ASH2L as a regulator of glioblastoma cell survival

Glioblastoma is the most common and aggressive primary brain tumor with poor prognosis, highlighting an urgent need for novel treatment strategies. In this study, we investigated epigenetic regulators of glioblastoma cell survival through CRISPR/Cas9 based genetic ablation screens using a customized sgRNA library EpiDoKOL, which targets critical functional domains of chromatin modifiers. Screens conducted in multiple cell lines revealed ASH2L, a histone lysine methyltransferase complex subunit, as a major regulator of glioblastoma cell viability. ASH2L depletion led to cell cycle arrest and apoptosis. RNA sequencing and greenCUT&RUN together identified a set of cell cycle regulatory genes, such as TRA2B, BARD1, KIF20B, ARID4A and SMARCC1 that were downregulated upon ASH2L depletion. Mass spectrometry analysis revealed the interaction partners of ASH2L in glioblastoma cell lines as SET1/MLL family members including SETD1A, SETD1B, MLL1 and MLL2. We further showed that glioblastoma cells had a differential dependency on expression of SET1/MLL family members for survival. The growth of ASH2L-depleted glioblastoma cells was markedly slower than controls in orthotopic in vivo models. TCGA analysis showed high ASH2L expression in glioblastoma compared to low grade gliomas and immunohistochemical analysis revealed significant ASH2L expression in glioblastoma tissues, attesting to its clinical relevance. Therefore, high throughput, robust and affordable screens with focused libraries, such as EpiDoKOL, holds great promise to enable rapid discovery of novel epigenetic regulators of cancer cell survival, such as ASH2L. Together, we suggest that targeting ASH2L could serve as a new therapeutic opportunity for glioblastoma.

cancer biology↗

rareSurvival: rare variant association analysis for time-to-event outcomes.

SummaryRare variants have been proposed as contributing to the "missing heritability" of complex human traits. There has been much recent development of methodology to investigate association of complex traits with multiple rare variants within pre-defined "units" from sequence and array-based studies of the exome or genome. However, software for modelling time to event outcomes for rare variant associations has been under developed in comparison with binary and quantitative traits. We introduce a new command line application, rareSurvival, used for the analysis of rare variants with time to event outcomes. The program is compatible with high performance computing (HPC) clusters for batch processing. rareSurvival implements statistical methodology, which are a combination of widely used survival and gene-based analysis techniques such as the Cox proportional hazards model and the burden test. We introduce a novel piece of software that will be at the forefront of efforts to discover rare variants associated with a variety of complex diseases with survival endpoints. Availability & ImplementationrareSurvival is implemented in C#, available on Linux, Windows and Mac OS X operating systems. It is freely available (GNU General Public License, version 3) to download from https://www.liverpool.ac.uk/translational-medicine/research/statistical-genetics/software/. Download Mono for Linux or Mac OS X to run software. Contacthamzah.syed@liverpool.ac.uk Supplementary informationLinks to additional figures and tables are available at Bioinformatics online.

bioinformatics↗

MOPower: an R-shiny application for the simulation and power calculation of multi-omics studies.

BackgroundMulti-omics studies are increasingly used to help understand the underlying mechanisms of clinical phenotypes, integrating information from the genome, transcriptome, epigenome, metabolome, proteome and microbiome. This integration of data is of particular use in rare disease studies where the sample sizes are often relatively small. Methods development for multi-omics studies is in its early stages due to the complexity of the different individual data types. There is a need for software to perform data simulation and power calculation for multi-omics studies to test these different methodologies and help calculate sample size before the initiation of a study. This software, in turn, will optimise the success of a study. ResultsThe interactive R shiny application MOPower described below simulates data based on three different omics using statistical distributions. It calculates the power to detect an association with the phenotype through analysis of n number of replicates using a variety of the latest multi-omics analysis models and packages. The simulation study confirms the efficiency of the software when handling thousands of simulations over ten different sample sizes. The average time elapsed for a power calculation run between integration models was approximately 500 seconds. Additionally, for the given study design model, power varied with the increase in the number of features affecting each method differently. For example, using MOFA had an increase in power to detect an association when the study sample size equally matched the number of features. ConclusionsMOPower addresses the need for flexible and user-friendly software that undertakes power calculations for multi-omics studies. MOPower offers users a wide variety of integration methods to test and full customisation of omics features to cover a range of study designs.

bioinformatics↗