Search bioRxivSearch

Biology subjects

Elemento, O.

Publications and source records attributed to Elemento, O..

14 recordsLinked to original sources

Robust Automated Assessment of Human Blastocyst Quality using Deep Learning

Morphology assessment has become the standard method for evaluation of embryo quality and selecting human blastocysts for transfer in in vitro fertilization (IVF). This process is highly subjective for some embryos and thus prone to human bias. As a result, morphological assessment results may vary extensively between embryologists and in some cases may fail to accurately predict embryo implantation and live birth potential. Here we postulated that an artificial intelligence (AI) approach trained on thousands of embryos can reliably predict embryo quality without human intervention.\n\nTo test this hypothesis, we implemented an AI approach based on deep neural networks (DNNs). Our approach called STORK accurately predicts the morphological quality of blastocysts based on raw digital images of embryos with 98% accuracy. These results indicate that a DNN can automatically and accurately grade embryos based on raw images. Using clinical data for 2,182 embryos, we then created a decision tree that integrates clinical parameters such as embryo quality and patient age to identify scenarios associated with increased or decreased pregnancy chance. This IVF data-driven analysis shows that the chance of pregnancy varies from 13.8% to 66.3%.\n\nIn conclusion, our AI-driven approach provides a novel way to assess embryo quality and uncovers new, potentially personalized strategies to select embryos with an improved likelihood of pregnancy outcome.

bioinformatics

A harmonized meta-knowledgebase of clinical interpretations of cancer genomic variants

Precision oncology relies on the accurate discovery and interpretation of genomic variants to enable individualized diagnosis, prognosis, and therapy selection. We found that knowledgebases containing clinical interpretations of somatic cancer variants are highly disparate in interpretation content, structure, and supporting primary literature, impeding consensus when evaluating variants and their relevance in a clinical setting. With the cooperation of experts of the Global Alliance for Genomics and Health (GA4GH) and six prominent cancer variant knowledgebases, we developed a framework for aggregating and harmonizing variant interpretations to produce a meta-knowledgebase of 12,856 aggregate interpretations covering 3,437 unique variants in 415 genes, 357 diseases, and 791 drugs. We demonstrated large gains in overlap between resources across variants, diseases, and drugs as a result of this harmonization. We subsequently demonstrated improved matching between a patient cohort and harmonized interpretations of potential clinical significance, observing an increase from an average of 33% per individual knowledgebase to 56% in aggregate. Our analyses illuminate the need for open, interoperable sharing of variant interpretation data. We also provide an open and freely available web interface (search.cancervariants.org) for exploring the harmonized interpretations from these six knowledgebases.

bioinformatics

Epigenetic analysis identifies factors driving racial disparity in prostate cancer

Prostate cancer (PCa) is the second most leading cause of death in men worldwide. African American men (AA) represent more aggressive form of PCa as compared to Caucasian (CA) counterparts. Evidence suggests that genetic and other biological factors could account for the observed racial disparity. We analyzed the cancer genome atlas (TCGA) dataset (2015) for existing epigenetic variation in AA and CA prostate cancer patients, and carried out Reduced Representation Bisulphite Sequencing (RRBS) analysis to identify global methylation changes in AA and CA prostate cancer patients. The TCGA dataset analysis revealed that the epigenetic heterogeneity could be categorized into 4 classes, where AA associated primarily to methylation cluster 1 (p value 0.048), and CA associated to methylation cluster 3 (p value 0.000146). We identified enrichment of Wnt signaling genes in both AA and CA, however they were differentially activated in terms of canonical and non-canonical Wnt signaling pathway activation. This was further validated using the GenomeDx expression data. Our RRBS data also suggested distinct methylation patterns in AA compared to CA, and in part validated our TCGA findings. Survival analysis using the RRBS data suggested hypomethylated genes to be significantly associated with recurrence of prostate cancer in CA (p=6.07x10-6) as well as in AA (p=0.0077). Overall, the observed racial disparity in the molecular mechanism involved in the pathogenesis of prostate cancer suggests diverse heterogeneity that potentially could affect survival and should be considered during prognosis and treatment.

cancer biology

Predicting peptide presentation by major histocompatibility complex class I using one million peptides

Improved computational tools are needed to prioritize putative neoantigens within immunotherapy pipelines for cancer treatment. Herein, we assemble a database of over one million human peptides presented by major histocompatibility complex class I (MHC-I), the largest known database of its type. We use these data to train a random forest classifier (ForestMHC) to predict likelihood of MHC-I presentation. The information content of features mirrors the canonical importance of positions two and nine in determining likelihood of binding. Our random forest-based method outperforms NetMHC and NetMHCpan on test sets, and it outperforms both these methods and MixMHCpred on new mass spectrometry data from an ovarian carcinoma sample. Furthermore, the random forest scores correlate monotonically with peptide binding affinities, when known. Finally, we examine the effect size of gene expression on peptide presentation and find a moderately strong relationship. The ForestMHC method is a promising modality to prioritize neoantigens for experimental testing in immunotherapy.

bioinformatics

A Machine Learning Approach Predicts Tissue-Specific Drug Adverse Events

One of the main causes for failure in the drug development pipeline or withdrawal post approval is the unexpected occurrence of severe drug adverse events. Even though such events should be detected by in vitro, in vivo, and human trials, they continue to unexpectedly arise at different stages of drug development causing costly clinical trial failures and market withdrawal. Inspired by the \"moneyball\" approach used in baseball to integrate diverse features to predict player success, we hypothesized that a similar approach could leverage existing adverse event and tissue-specific toxicity data to learn how to predict adverse events. We introduce MAESTER, a data-driven machine learning approach that integrates information on a compounds structure, targets, and phenotypic effects with tissue-wide genomic profiling and our toxic target database to predict the probability of a compound presenting with different types of tissue-specific adverse events. When tested on 6 different types of adverse events MAESTER maintains a high accuracy, sensitivity, and specificity across both the training data and new test sets. Additionally, MAESTER scores could flag a number of drugs that were approved, but later withdrawn due to unknown adverse events - highlighting its potential to identify events missed by traditional methods. MAESTER can also be used to identify toxic targets for each tissue type. Overall MAESTER provides a broadly applicable framework to identify toxic targets and predict specific adverse events and can accelerate the drug development pipeline and drive the design of new safer compounds.

pharmacology and toxicology

Generation of pulmonary neuro-endocrine cells and tumors resembling small cell lung cancers from human embryonic stem cells

SUMMARYBy blocking an important signaling pathway (called NOTCH) and interfering with expression of two tumor suppressor genes in cells derived from human embryonic stem cells, the authors have developed a model for studying highly lethal small cell lung cancers.\n\nABSTRACTCell culture models based on directed differentiation of human embryonic stem cells (hESCs) may reveal why certain constellations of genetic changes drive carcinogenesis in specialized human cell lineages. Here we demonstrate that up to 10 percent of lung progenitor cells derived from hESCs can be induced to form pulmonary neuroendocrine cells (PNECs), the putative normal precursors to small cell lung cancers (SCLCs), by inhibition of NOTCH signaling. By using small inhibitory RNAs in these cultures to reduce levels of retinoblastoma (RB) protein, the product of a gene commonly mutated in SCLCs, we can significantly expand the number of PNECs. Similarly reducing levels of TP53 protein, the product of another tumor suppressor gene commonly mutated in SCLCs, or expressing mutant KRAS or EGFR genes, did not induce or expand PNECs, consistent with lineage-specific sensitivity to loss of RB function. Tumors resembling early stage SCLC grew in immunodeficient mice after subcutaneous injection of PNEC-containing cultures in which expression of both RB and TP53 was blocked. Single-cell RNA profiles of PNECs are heterogeneous; when RB levels are reduced, the profiles show similarities to RNA profiles from early stage SCLC; when both RB and TP53 levels are reduced, the transcriptome is enriched with cell cycle-specific RNAs. Taken together, these findings suggest that genetic manipulation of hESC-derived pulmonary cells will enable studies of the initiation, progression, and treatment of this recalcitrant cancer.

cancer biology

Accelerated lipid catabolism and autophagy are cancer survival mechanisms under inhibited glutaminolysis

Suppressing glutaminolysis does not always induce cancer cell death in glutamine-dependent tumors because cells may switch to alternative energy sources. To reveal compensatory metabolic pathways, we investigated the metabolome-wide cellular response to inhibited glutaminolysis. We conducted metabolic profiling in the triple-negative breast cancer cell line MB-MDA-231, treated with different dosages of glutaminase inhibitor C.968 at multiple time points. We found that multiple molecules involved in lipid catabolism responded directly to glutamate deficiency as a presumed compensation for energy deficit. Accelerated lipid catabolism, together with oxidative stress induced by glutaminolysis inhibition, triggered autophagy. We therefore simultaneously inhibited glutaminolysis and autophagy, which induced cancer cell death. Our study emphasizes the potential of non-targeted metabolomics to characterize and identify metabolic escape mechanisms contributing to cancer cell survival under treatment. Our findings add to the increasing evidence that combined inhibition of glutaminolysis and autophagy may be effective in glutamine-addicted cancers.

cancer biology

Breast Cancer Histopathological Image Classification: A Deep Learning Approach

Breast cancer remains the most common type of cancer and the leading cause of cancer-induced mortality among women with 2.4 million new cases diagnosed and 523,000 deaths per year. Historically, a diagnosis has been initially performed using clinical screening followed by histopathological analysis. Automated classification of cancers using histopathological images is a chciteallenging task of accurate detection of tumor sub-types. This process could be facilitated by machine learning approaches, which may be more reliable and economical compared to conventional methods.\n\nTo prove this principle, we applied fine-tuned pre-trained deep neural networks. To test the approach we first classify different cancer types using 6, 402 tissue micro-arrays (TMAs) training samples. Our framework accurately detected on average 99.8% of the four cancer types including breast, bladder, lung and lymphoma using the ResNet V1 50 pre-trained model. Then, for classification of breast cancer sub-types this approach was applied to 7,909 images from the BreakHis database. In the next step, ResNet V1 152 classified benign and malignant breast cancers with an accuracy of 98.7%. In addition, ResNet V1 50 and ResNet V1 152 categorized either benign- (adenosis, fibroadenoma, phyllodes tumor, and tubular adenoma) or malignant- (ductal carcinoma, lobular carcinoma, mucinous carcinoma, and papillary carcinoma) sub-types with 94.8% and 96.4% accuracy, respectively. The confusion matrices revealed high sensitivity values of 1, 0.995 and 0.993 for cancer types, as well as malignant- and benign sub-types respectively. The areas under the curve (AUC) scores were 0.996,0.973 and 0.996 for cancer types, malignant and benign sub-types, respectively. Overall, our results show negligible false negative (on average 3.7 samples) and false positive (on average 2 samples) results among different models. Availability: Source codes, guidelines and data sets are temporarily available on google drive upon request before moving to a permanent GitHub repository.

bioinformatics

Challenges in Using ctDNA to Achieve Early Detection of Cancer

Early detection of cancer is a significant unmet clinical need. Improved technical ability to detect circulating tumor-derived DNA (ctDNA) in the cell-free DNA (cfDNA) component of blood plasma via next-generation sequencing and established correlations between ctDNA load and tumor burden in cancer patients have spurred excitement about the possibilities of detecting cancer early by performing ctDNA mutation detection.\n\nWe reanalyze published data on the expected ctDNA allele fraction in early-stage cancer and the population statistics of cfDNA concentration to show that under conservative technical assumptions, high-sensitivity cancer detection by ctDNA mutation detection will require either more blood volume (150-300mL) than practical for a routine screen or variant filtering that may be impossible given our knowledge of cancer evolution, and will likely remain out of economic reach for routine population screening without multiple-order-of-magnitude decreases in sequencing cost. Instead, new approaches that integrate ctDNA mutations with multiple other blood-based analytes (such as exosomes, circulating tumor cells, ctDNA epigenetics, metabolites) as well as integration of these signals over time for each individual may be needed.

cancer biology

Deep Convolutional Neural Networks Enable Discrimination of Heterogeneous Digital Pathology Images

Pathological evaluation of tumor tissue is pivotal for diagnosis in cancer patients and automated image analysis approaches have great potential to increase precision of diagnosis and help reduce human error.\n\nIn this study, we utilize various computational methods based on convolutional neural networks (CNN) and build a stand-alone pipeline to effectively classify different histopathology images across different types of cancer. In particular, we demonstrate the utility of our pipeline to discriminate between two subtypes of lung cancer, four biomarkers of bladder cancer, and five biomarkers of breast cancer. In addition, we apply our pipeline to discriminate among four immunohistochemistry (IHC) staining scores of bladder and breast cancers.\n\nOur classification pipeline utilizes a basic architecture of CNN, Googles Inceptions within three training strategies, and an ensemble of two state-of-the-art algorithms, Inception and ResNet. These strategies include training the last layer of Googles Inceptions, training the network from scratch, and fine-tunning the parameters for our data using two pre-trained version of Googles Inception architectures, Inception-V1 and Inception-V3.\n\nWe demonstrate the power of deep learning approaches for identifying cancer subtypes, and the robustness of Googles Inceptions even in presence of extensive tumor heterogeneity. Our pipeline on average achieved accuracies of 100%, 92%, 95%, and 69% for discrimination of various cancer types, subtypes, biomarkers, and scores, respectively. Our pipeline and related documentation is freely available at https://github.com/ih-lab/CNN_Smoothie.

bioinformatics

A New Big-Data Paradigm For Target Identification And Drug Discovery

Drug target identification is one of the most important aspects of pre-clinical development yet it is also among the most complex, labor-intensive, and costly. This represents a major issue, as lack of proper target identification can be detrimental in determining the clinical application of a bioactive small molecule. To improve target identification, we developed BANDIT, a novel paradigm that integrates multiple data types within a Bayesian machine-learning framework to predict the targets and mechanisms for small molecules with unprecedented accuracy and versatility. Using only public data BANDIT achieved an accuracy of approximately 90% over 2000 different small molecules - substantially better than any other published target identification platform. We applied BANDIT to a library of small molecules with no known targets and generated [~]4,000 novel molecule-target predictions. From this set we identified and experimentally validated a set of novel microtubule inhibitors, including three with activity on cancer cells resistant to clinically used anti-microtubule therapies. We next applied BANDIT to ONC201 - an active anti- cancer small molecule in clinical development - whose target has remained elusive since its discovery in 2009. BANDIT identified dopamine receptor 2 as the unexpected target of ONC201, a prediction that we experimentally validated. Not only does this open the door for clinical trials focused on target-based selection of patient populations, but it also represents a novel way to target GPCRs in cancer. Additionally, BANDIT identified previously undocumented connections between approved drugs with disparate indications, shedding light onto previously unexplained clinical observations and suggesting new uses of marketed drugs. Overall, BANDIT represents an efficient and highly accurate platform that can be used as a resource to accelerate drug discovery and direct the clinical application of small molecule therapeutics with improved precision.

pharmacology and toxicology

Discovery and reporting of clinically-relevant germline variants in advanced cancer patients assessed using whole-exome sequencing

PurposeIn precision cancer care, WES-based analysis of tumor-normal samples helps reveal somatic alterations but can also identify cancer-associated germline variants important for disease surveillance, treatment choice and cancer prevention. WES can also identify germline secondary findings impacting risk of cardiac, neurodegenerative or metabolic diseases. In patients with advanced cancer, the frequency of reportable secondary findings encountered with WES is not well defined.\n\nMethodsTo address this question, we analyzed a cohort of 343 patients with advanced, metastatic cancer for whom we have performed tumor and germline WES interrogating more than 21,000 genes using a CLIA/CLEP approved assay.\n\nResults17% of patients in our cohort have one or more reportable germline variants, including patients with pathogenic variants in the BRCA1 and BRCA2 genes. The frequency of non-cancer clinically relevant germline variants (8.8%) was within the range of two control non-cancer cohorts (11.0% and 6.5%). The frequency of variants in cancer-associated genes was significantly higher (p<0.0005) in our advanced cancer cohort (8.2%) compared to control cohorts (2.7% and 3.8%). More than 50% of patients with reportable germline cancer variants had a family history of cancer.\n\nConclusionthese results stress the importance of returning germline results found during somatic genomic tumor testing.

genomics

EthSEQ: ethnicity annotation from whole exome sequencing data

Whole exome sequencing (WES) is widely utilized both in translational cancer genomics studies and in the setting of precision medicine. Stratification of individual's ethnicity is fundamental for the correct interpretation of personal genomic variation impact. We implemented EthSEQ to provide reliable and rapid ethnicity annotation from whole exome sequencing individual's data and validated it on 1,000 Genome Project and TCGA data demonstrating high precision (>99%). EthSEQ can be integrated into any WES based processing pipeline and exploits multi-core capabilities. Source code, manual and other data is available at http://demichelislab.unitn.it/EthSEQ.

bioinformatics

AGFusion: annotate and visualize gene fusions

SummaryThe discovery of novel gene fusions in tumor samples has rapidly accelerated with the rise of next-generation sequencing. A growing number of tools enable discovery of gene fusions from RNA-seq data. However it is likely that not all gene fusions are driving tumors. Assessing the potential functional consequences of a fusion is critical to understand their driver role. It is also challenging as gene fusions are described by chromosomal breakpoint coordinates that need to be translated into an actual amino acid fusion sequence and predicted domain architecture of the fusion proteins. Currently there are no easy-to-use tools that can automatically reconstruct and visualize fusion proteins from genomic breakpoints. To facilitate the functional interpretation of gene fusions, we developed AGFusion, available as an online web tool that can be readily used by non-computational researchers as well as a python package that can be built into computational pipelines. With minimal input from the user, AGFusion predicts the cDNA, CDS, and protein sequences of all gene fusion products based on all combinations of gene isoforms. For protein coding fusions, AGFusion can annotate and visualize the protein domain architecture. AGFusion currently supports Homo sapiens (genome builds GRCh37 and GRCh38) and Mus musculus (genome build GRCm38) and new genomes can easily be added.\n\nAvailabilityAGFusion python package is freely available at https://github.com/murphycj/AGFusion under the MIT license. The AGFusion web app is available at http://agfusion.info

bioinformatics