Search bioRxivSearch

Biology subjects

Parker, S.

Publications and source records attributed to Parker, S..

3 recordsLinked to original sources

YAMDA: thousandfold speedup of EM-based motif discovery using deep learning libraries and GPU

MotivationMotif discovery in large biopolymer sequence datasets can be computationally demanding, presenting significant challenges for discovery in omics research. MEME, arguably one of the most popular motif discovery software, takes quadratic time with respect to dataset size, leading to excessively long runtimes for large datasets. Therefore, there is a demand for fast programs that can generate results of the same quality as MEME.\n\nResultsHere we describe YAMDA, a highly scalable motif discovery software package. It is built on Pytorch, a tensor computation deep learning library with strong GPU acceleration that is highly optimized for tensor operations that are also useful for motifs. YAMDA takes linear time to find motifs as accurately as MEME, completing in seconds or minutes, which translates to speedups over a thousandfold.\n\nAvailabilityYAMDA is freely available on Github (https://github.com/daquang/YAMDA)\n\nContactdaquang@umich.edu

bioinformatics

Proteomic Architecture of Human Coronary and Aortic Atherosclerosis

The inability to detect premature atherosclerosis significantly hinders implementation of personalized therapy to prevent coronary heart disease. A comprehensive understanding of arterial protein networks and how they change in early atherosclerosis could identify new biomarkers for disease detection and improved therapeutic targets. Here we describe the human arterial proteome and the proteomic features strongly associated with early atherosclerosis based on mass-spectrometry analysis of coronary artery and aortic specimens from 100 autopsied young adults (200 arterial specimens). Convex analysis of mixtures, differential dependent network modeling and bioinformatic analyses defined the composition, network re-wiring and likely regulatory features of the protein networks associated with early atherosclerosis. Among other things the results reveal major differences in mitochondrial protein mass between the coronary artery and distal aorta in both normal and atherosclerotic samples - highlighting the importance of anatomic specificity and dynamic network structures in in the study of arterial proteomics. The publicly available data resource and the description of the analysis pipeline establish a new foundation for understanding the proteomic architecture of atherosclerosis and provide a template for similar investigations of other chronic diseases characterized by multi-cellular tissue phenotypes.\n\nHighlightsO_LILC MS/MS analysis performed on 200 human aortic or coronary artery samples\nC_LIO_LINumerous proteins, networks, and regulatory pathways associated with early atherosclerosis\nC_LIO_LIMitochondrial proteins mass and selected metabolic regulatory pathways vary dramatically by disease status and anatomic location\nC_LIO_LIPublically available data resource and analytic pipeline are provided or described in detail\nC_LI

molecular biology

Cis-Compound Mutations are Prevalent in Triple Negative Breast Cancer and Can Drive Tumor Progression

About 16% of breast cancers fall into a clinically aggressive category designated triple negative (TNBC) due to a lack of ERBB2, estrogen receptor and progesterone receptor expression1-3. The mutational spectrum of TNBC has been characterized as part of The Cancer Genome Atlas (TCGA)4; however, snapshots of primary tumors cannot reveal the mechanisms by which TNBCs progress and spread. To address this limitation we initiated the Intensive Trial of OMics in Cancer (ITOMIC)-001, in which patients with metastatic TNBC undergo multiple biopsies over space and time5. Whole exome sequencing (WES) of 67 samples from 11 patients identified 426 genes containing multiple distinct single nucleotide variants (SNVs) within the same sample, instances we term Multiple SNVs affecting the Same Gene and Sample (MSSGS). We find that >90% of MSSGS result from cis-compound mutations (in which both SNVs affect the same allele), that MSSGS comprised of SNVs affecting adjacent nucleotides arise from single mutational events, and that most other MSSGS result from the sequential acquisition of SNVs. Some MSSGS drive cancer progression, as exemplified by a TNBC driven by FGFR2(S252W;Y375C). MSSGS are more prevalent in TNBC than other breast cancer subtypes and occur at higher-than-expected frequencies across TNBC samples within TCGA. MSSGS may denote genes that play as yet unrecognized roles in cancer progression.

clinical trials