Search bioRxivSearch

Biology subjects

Hajirasouliha, I.

Publications and source records attributed to Hajirasouliha, I..

8 recordsLinked to original sources

Robust Automated Assessment of Human Blastocyst Quality using Deep Learning

Morphology assessment has become the standard method for evaluation of embryo quality and selecting human blastocysts for transfer in in vitro fertilization (IVF). This process is highly subjective for some embryos and thus prone to human bias. As a result, morphological assessment results may vary extensively between embryologists and in some cases may fail to accurately predict embryo implantation and live birth potential. Here we postulated that an artificial intelligence (AI) approach trained on thousands of embryos can reliably predict embryo quality without human intervention.\n\nTo test this hypothesis, we implemented an AI approach based on deep neural networks (DNNs). Our approach called STORK accurately predicts the morphological quality of blastocysts based on raw digital images of embryos with 98% accuracy. These results indicate that a DNN can automatically and accurately grade embryos based on raw images. Using clinical data for 2,182 embryos, we then created a decision tree that integrates clinical parameters such as embryo quality and patient age to identify scenarios associated with increased or decreased pregnancy chance. This IVF data-driven analysis shows that the chance of pregnancy varies from 13.8% to 66.3%.\n\nIn conclusion, our AI-driven approach provides a novel way to assess embryo quality and uncovers new, potentially personalized strategies to select embryos with an improved likelihood of pregnancy outcome.

bioinformatics

Characterization of segmental duplications and large inversions using Linked-Reads

Many algorithms aimed at characterizing genomic structural variation (SV) have been developed since the inception of high-throughput sequencing. However, the full spectrum of SVs in the human genome is not yet assessed. Most of the existing methods focus on discovery and genotyping of deletions, insertions, and mobile elements. Detection of balanced SVs with no gain or loss of genomic segments (e.g., inversions) is particularly a challenging task. Long read sequencing has been leveraged to find short inversions but there is still a need to develop methods to detect large genomic inversions. Furthermore, currently there are no algorithms to predict the insertion locus of large interspersed segmental duplications.\n\nHere we propose novel algorithms to characterize large (>40Kbp) interspersed segmental duplications and (>80Kbp) inversions using Linked-Read sequencing data. Linked-Read sequencing provides long range information, where Illumina reads are tagged with barcodes that can be used to assign short reads to pools of larger (30-50 Kbp) molecules. Our methods rely on split molecule sequence signature that we have previously described [11]. Similar to the split read, split molecules refer to large segments of DNA that span an SV breakpoint. Therefore, when mapped to the reference genome, the mapping of these segments would be discontinuous. We redesign our earlier algorithm, VALOR, to specifically leverage Linked-Read sequencing data to discover large inversions and characterize interspersed segmental duplications. We implement our new algorithms in a new software package, called VALOR2.\n\nAvailabilityVALOR2 is available at https://github.com/BilkentCompGen/valor.

bioinformatics

PhISCS - A Combinatorial Approach for Sub-perfect Tumor Phylogeny Reconstruction via Integrative use of Single Cell and Bulk Sequencing Data

Recent technological advances in single cell sequencing (SCS) provide high resolution data for studying intra-tumor heterogeneity and tumor evolution. Available computational methods for tumor phylogeny inference via SCS typically aim to identify the most likely perfect phylogeny tree satisfying infinite sites assumption (ISA). However limitations of SCS technologies such as frequent allele dropout or highly variable sequence coverage, commonly result in mutational call errors and prohibit a perfect phylogeny. In addition, ISA violations are commonly observed in tumor phylogenies due to the loss of heterozygosity, deletions and convergent evolution. In order to address such limitations, we, for the first time, introduce a new combinatorial formulation that integrates single cell sequencing data with matching bulk sequencing data, with the objective of minimizing a linear combination of (i) potential false negatives (due to e.g. allele dropout or variance in sequence coverage) and (ii) potential false positives (due to e.g. read errors) among mutation calls, as well as (iii) the number of mutations that violate ISA - to define the optimal sub-perfect phylogeny. Our formulation ensures that several lineage constraints imposed by the use of variant allele frequencies (VAFs, derived from bulk sequence data) are satisfied. We express our formulation both in the form of an integer linear program (ILP) and - for the first time in the context of tumor phylogeny reconstruction - a boolean constraint satisfaction problem (CSP) and solve them by leveraging state-of-the-art ILP/CSP solvers. The resulting method, which we name PhISCS, is the first to integrate SCS and bulk sequencing data under the finite sites model. Using several simulated and real SCS data sets, we demonstrate that PhISCS is not only more general but also more accurate than the alternative tumor phylogeny inference tools. PhISCS is very fast especially when its CSP based variant is used returns the optimal solution, except in rare instances for which it provides an optimality gap. PhISCS is available at https://github.com/haghshenas/PhISCS.

bioinformatics

gpps: An ILP-based approach for inferringcancer progression with mutation losses fromsingle cell data

MotivationIn recent years, the well-known Infinite Sites Assumption (ISA) has been a fundamental feature of computational methods devised for reconstructing tumor phylogenies and inferring cancer progression where mutations are accumulated through histories. However, some recent studies leveraging Single Cell Sequencing (SCS) techniques have shown evidence of mutation losses in several tumor samples [19], making the inference problem harder.\n\nResultsWe present a new tool, gpps, that reconstructs a tumor phylogeny from single cell data, allowing each mutation to be lost at most a fixed number of times.\n\nAvailabilityThe General Parsimony Phylogeny from Single cell (gpps) tool is open source and available at https://github.com/AlgoLab/gppf.

cancer biology

Inferring Cancer Progression from Single Cell Sequencing while allowing loss of mutations

MotivationIn recent years, the well-known Infinite Sites Assumption (ISA) has been a fundamental feature of computational methods devised for reconstructing tumor phylogenies and inferring cancer progressions seen as an accumulation of mutations. However, recent studies (Kuipers et al., 2017) leveraging Single-cell Sequencing (SCS) techniques have shown evidence of the widespread recurrence and, especially, loss of mutations in several tumor samples. Still, established methods that can infer phylogenies with mutation losses are however lacking.\n\nResultsWe present the SASC (Simulated Annealing Single-Cell inference) tool which is a new and robust approach based on simulated annealing for the inference of cancer progression from SCS data. More precisely, we introduce a simple extension of the model of evolution where mutations are only accumulated, by allowing also a limited amount of back mutations in the evolutionary history of the tumor: the Dollo-k model. We demonstrate that SASC achieves high levels of accuracy when tested on both simulated and real data sets and in comparison with some other available methods.\n\nAvailabilityThe Simulated Annealing Single-cell inference (SASC) tool is open source and available at https://github.com/sciccolella/sasc.\n\nContacts.ciccolella@campus.unimib.it

bioinformatics

Breast Cancer Histopathological Image Classification: A Deep Learning Approach

Breast cancer remains the most common type of cancer and the leading cause of cancer-induced mortality among women with 2.4 million new cases diagnosed and 523,000 deaths per year. Historically, a diagnosis has been initially performed using clinical screening followed by histopathological analysis. Automated classification of cancers using histopathological images is a chciteallenging task of accurate detection of tumor sub-types. This process could be facilitated by machine learning approaches, which may be more reliable and economical compared to conventional methods.\n\nTo prove this principle, we applied fine-tuned pre-trained deep neural networks. To test the approach we first classify different cancer types using 6, 402 tissue micro-arrays (TMAs) training samples. Our framework accurately detected on average 99.8% of the four cancer types including breast, bladder, lung and lymphoma using the ResNet V1 50 pre-trained model. Then, for classification of breast cancer sub-types this approach was applied to 7,909 images from the BreakHis database. In the next step, ResNet V1 152 classified benign and malignant breast cancers with an accuracy of 98.7%. In addition, ResNet V1 50 and ResNet V1 152 categorized either benign- (adenosis, fibroadenoma, phyllodes tumor, and tubular adenoma) or malignant- (ductal carcinoma, lobular carcinoma, mucinous carcinoma, and papillary carcinoma) sub-types with 94.8% and 96.4% accuracy, respectively. The confusion matrices revealed high sensitivity values of 1, 0.995 and 0.993 for cancer types, as well as malignant- and benign sub-types respectively. The areas under the curve (AUC) scores were 0.996,0.973 and 0.996 for cancer types, malignant and benign sub-types, respectively. Overall, our results show negligible false negative (on average 3.7 samples) and false positive (on average 2 samples) results among different models. Availability: Source codes, guidelines and data sets are temporarily available on google drive upon request before moving to a permanent GitHub repository.

bioinformatics

Minerva: An Alignment and Reference Free Approach to Deconvolve Linked-Reads for Metagenomics

Emerging Linked-Read technologies (aka Read-Cloud or barcoded short-reads) have revived interest in standard short-read technology as a viable way to understand large-scale structure in genomes and metagenomes. Linked-Read technologies, such as the 10X Chromium system, use a microfluidic system and a set of specially designed 3 barcodes (aka UIDs) to tag short DNA reads which were originally sourced from the same long fragment of DNA; subsequently, these specially barcoded reads are sequenced on standard short read platforms. This approach results in interesting compromises. Each long fragment of DNA is covered only sparsely by short reads, no information about the relative ordering of reads from the same fragment is preserved, and typically each 3 barcode matches reads from 2-20 long fragments of DNA. However, compared to long read platforms like those produced by Pacific Biosciences and Oxford Nanopore the cost per base to sequence is far lower, far less input DNA is required, and the per base error rate is that of Illumina short-reads.\n\nThe use of Linked-Reads presents a new set of algorithmic challenges. In this paper, we formally describe one particular issue common to all applications of Linked-Read technology: the deconvolution of reads with a single 3 barcode into clusters that correspond to a single long fragment of DNA. We introduce Minerva, A graph-based algorithm that approximately solves the barcode deconvolution problem for metagenomic data (where reference genomes may be incomplete or unavailable). Additionally, we develop two demonstrations where the deconvolution of barcoded reads improves downstream results: improving the specificity of taxonomic assignments, and by improving clustering of related sequences. To the best of our knowledge, we are the first to address the problem of barcode deconvolution in metagenomics.

bioinformatics

Deep Convolutional Neural Networks Enable Discrimination of Heterogeneous Digital Pathology Images

Pathological evaluation of tumor tissue is pivotal for diagnosis in cancer patients and automated image analysis approaches have great potential to increase precision of diagnosis and help reduce human error.\n\nIn this study, we utilize various computational methods based on convolutional neural networks (CNN) and build a stand-alone pipeline to effectively classify different histopathology images across different types of cancer. In particular, we demonstrate the utility of our pipeline to discriminate between two subtypes of lung cancer, four biomarkers of bladder cancer, and five biomarkers of breast cancer. In addition, we apply our pipeline to discriminate among four immunohistochemistry (IHC) staining scores of bladder and breast cancers.\n\nOur classification pipeline utilizes a basic architecture of CNN, Googles Inceptions within three training strategies, and an ensemble of two state-of-the-art algorithms, Inception and ResNet. These strategies include training the last layer of Googles Inceptions, training the network from scratch, and fine-tunning the parameters for our data using two pre-trained version of Googles Inception architectures, Inception-V1 and Inception-V3.\n\nWe demonstrate the power of deep learning approaches for identifying cancer subtypes, and the robustness of Googles Inceptions even in presence of extensive tumor heterogeneity. Our pipeline on average achieved accuracies of 100%, 92%, 95%, and 69% for discrimination of various cancer types, subtypes, biomarkers, and scores, respectively. Our pipeline and related documentation is freely available at https://github.com/ih-lab/CNN_Smoothie.

bioinformatics