Search bioRxiv⌕ Search

Biology subjects

Hamamsy, T.

Publications and source records attributed to Hamamsy, T..

2 recordsLinked to original sources

Single-cell gene regulatory network inference atscale: The Inferelator 3.0

MotivationGene regulatory networks define regulatory relationships between transcription factors and target genes within a biological system, and reconstructing them is essential for understanding cellular growth and function. Methods for inferring and reconstructing networks from genomics data have evolved rapidly over the last decade in response to advances in sequencing technology and machine learning. The scale of data collection has increased dramatically; the largest genome-wide gene expression datasets have grown from thousands of measurements to millions of single cells, and new technologies are on the horizon to increase to tens of millions of cells and above. ResultsIn this work, we present the Inferelator 3.0, which has been significantly updated to integrate data from distinct cell types to learn context-specific regulatory networks and aggregate them into a shared regulatory network, while retaining the functionality of the previous versions. The Inferelator is able to integrate the largest single-cell datasets and learn cell-type specific gene regulatory networks. Compared to other network inference methods, the Inferelator learns new and informative Saccharomyces cerevisiae networks from single-cell gene expression data, measured by recovery of a known gold standard. We demonstrate its scaling capabilities by learning networks for multiple distinct neuronal and glial cell types in the developing Mus musculus brain at E18 from a large (1.3 million) single-cell gene expression dataset with paired single-cell chromatin accessibility data. AvailabilityThe inferelator software is available on GitHub (https://github.com/flatironinstitute/inferelator) under the MIT license and has been released as python packages with associated documentation (https://inferelator.readthedocs.io/).

systems biology↗

Quantifying the severity of adverse drug reactions using social media

Adverse drug reactions (ADRs) impact the health of 100,000s of individuals annually in the United States with associated costs in the hundreds of billions. The monitoring and analysis of the severity of adverse drug reactions is limited by the current qualitative and categorical system of severity classifications. Previous efforts have generated quantitative estimates for a subset of ADRs, but were limited in scope due to the time and costs associated with the efforts. We present a semi-supervised approach that estimates ADR severity by using a lexical network of ADR word embeddings and label propagation. We use this method to estimate the severity of 28,113 ADRs, representing 12,198 unique ADR concepts from MedDRA. Our Severity of Adverse Events Derived from Reddit (SO_SCPLOWAEDRC_SCPLOW) scores have good correlations with real-world outcomes. SO_SCPLOWAEDRC_SCPLOW scores had Spearman correlations with ADR case outcomes in FAERS of 0.595, 0.633, and -0.748 for death, serious outcome, and no outcome, respectively. We investigate different methods for defining initial seed term sets and evaluate their impact on severity estimates. We analyzed severity distributions for ADRs based on their appearance in Boxed Warning drug label sections, as well as ADRs with sex-specific associations. We find that ADRs discovered postmarket have significantly greater severity compared to those discovered in the clinical trial. We create quantitative Drug RIsk Profile (DO_SCPLOWRIPC_SCPLOW) scores for 968 drugs that have a Spearman correlation of 0.377 with drugs ranked by FAERS cases resulting in death, where the given drug was the primary suspect. We make the SO_SCPLOWAEDRC_SCPLOW and DO_SCPLOWRIPC_SCPLOW scores publicly available in order to enable more quantitative analysis of pharmacovigilance data.

bioinformatics↗