Search bioRxiv⌕ Search

Biology subjects

Arias-Gaguancela, O.

Publications and source records attributed to Arias-Gaguancela, O..

2 recordsLinked to original sources

LeukGenePipeline: Modular Workflow for Genomic Datasets

Analyzing human genome data has become increasingly common, supported by the growing availability of public repositories that enable statistical modeling and predictive analysis. However, researchers from wet lab-based disciplines often face challenges due to limited training in programming and computational tools. To address this barrier, we introduce LeukGenePipeline (LGP): a user-friendly, Python-based tool designed to automate core genomic analyses. As a proof of concept, LGP was used to perform mutation classification, copy number variation (CNV) analysis, pathway enrichment analysis (PEA), and gene ontology (GO) enrichment using data from the COSMIC public database (v101) with a focus on acute myeloid leukemia (AML). Data consisted of a mutation table with 830,978 unique rows associated with protein-coding genes, and a CNV table with 12,926 gene-level entries. LGP outputs revealed frequently mutated, CNV-altered genes, and enrichment of key transcription factors associated with leukemogenesis.

bioinformatics↗

QuickProt: A bioinformatics and visualization tool for DIA and PRM mass spectrometry-based proteomics datasets

Mass spectrometry (MS)-based proteomics focuses on identifying and quantifying peptides and proteins in biological samples. Processing of MS-derived raw data, including deconvolution, alignment, and peptide-protein prediction, has been achieved through various software platforms. However, the downstream analysis, including quality control, visualizations, and interpretation of proteomics results remains challenging due to the lack of integrated tools to facilitate the analyses. To address this challenge, we developed QuickProt, a series of Python-based Google Colab notebooks for analyzing data-independent acquisition (DIA) and parallel reaction monitoring (PRM) proteomics datasets. These pipelines are designed so that users with no coding expertise can utilize the tool. Furthermore, as open-source code, QuickProt notebooks can be customized and incorporated into existing workflows. As proof of concept, we applied QuickProt to analyze in-house DIA and stable isotope dilution (SID)-PRM MS proteomics datasets from a time-course study of human erythropoiesis. The analysis resulted in annotated tables and publication-ready figures revealing a dynamic rearrangement of the proteome during erythroid differentiation, with the abundance of proteins linked to gene regulation, metabolic, and chromatin remodeling pathways increasing early in erythropoiesis. Altogether, these tools aim to automate and streamline DIA and PRM-MS proteomics data analysis, making it more efficient and less time-consuming.

bioinformatics↗