Search bioRxiv⌕ Search

Biology subjects

Mehmood, T.

Publications and source records attributed to Mehmood, T..

2 recordsLinked to original sources

Multiblock LASSO Framework for Cancer Gene Selection from RNA-Seq PANCAN Data

The Cancer RNA-HiSeq PANCAN dataset consists of RNA-Seq gene expression data collected from multiple cancer types. It is a high-dimensional dataset, meaning it has thousands of gene expression features (predictors) and relatively fewer samples (observations). The dataset contains thousands of genes, making it difficult to identify key biomarkers. In order to reduce data and comprehend the modeled link, variable selection is essential. Least Regression using Absolute Shrinkage and Selection Operator (LASSO) is one modeling technique that deals with high throughput data. The data might be divided into different blocks representing different biological pathways or cancer types. Many genes are correlated, which can reduce interpretability. In many areas, including modern biology, variable selection is an important problem. For instance, choosing genetic characteristics for categorization (i.e., identifying harmful bacteria, diagnosing diseases, etc.) is an example of this. Multiblock Lasso (a variant of Lasso regression) is particularly useful when data is structured into blocks (e.g., different biological processes or pathways). It helps in selecting important features across multiple blocks, improving interpretability by grouping related genes, reducing over fitting in high-dimensional datasets. In this study, we apply Multiblock Lasso to extract significant gene features for cancer classification. We preprocess the dataset, define block structures using biological pathways, and optimize the regularization parameters using cross-validation. Experimental results demonstrate that Multiblock Lasso effectively reduces dimensionality while maintaining classification accuracy, making it a powerful tool for biomarker discovery in cancer genomics.

genetics↗

Integrating Genomic and Environmental Data Using Machine Learning for Vernalization Response Prediction

This research investigates the integration of genomic and environmental data using Random Forests to predict vernalization response in barley. Vernalization, the requirement of a prolonged period of cold to induce flowering, is a critical adaptive trait for temperate cereal crops. The study compiles a comprehensive dataset of barley genotypes, gene expression levels related to vernalization (e.g., VRN1, VRN2, and FT1 genes), and detailed environmental variables including temperature, photoperiod, soil moisture, and humidity. By employing a Random Forest algorithm, the research identifies key genetic and environmental factors that influence vernalization. The findings suggest that this machine learning approach effectively models the complex interactions between genotype and environment, providing insights for breeding climate-resilient barley varieties. This integrative approach not only enhances our understanding of the genetic basis of vernalization but also aids in the development of barley varieties with optimized flowering times for diverse climatic conditions.

genomics↗