Search bioRxivSearch

Biology subjects

Xie, X.

Publications and source records attributed to Xie, X..

At least 19 recordsLinked to original sources

AnatomyNet: Deep 3D Squeeze-and-excitation U-Nets for fast and fully automated whole-volume anatomical segmentation

PurposeRadiation therapy (RT) is a common treatment for head and neck (HaN) cancer where therapists are often required to manually delineate boundaries of the organs-at-risks (OARs). The radiation therapy planning is time-consuming as each computed tomography (CT) volumetric data set typically consists of hundreds to thousands of slices and needs to be individually inspected. Automated head and neck anatomical segmentation provides a way to speed up and improve the reproducibility of radiation therapy planning. Previous work on anatomical segmentation is primarily based on atlas registrations, which takes up to hours for one patient and requires sophisticated atlas creation. In this work, we propose the AnatomyNet, an end-to-end and atlas-free three dimensional squeeze-and-excitation U-Net (3D SE U-Net), for fast and fully automated whole-volume HaN anatomical segmentation.\n\nMethodsThere are two main challenges for fully automated HaN OARs segmentation: 1) challenge in segmenting small anatomies (i.e., optic chiasm and optic nerves) occupying only a few slices, and 2) training model with inconsistent data annotations with missing ground truth for some anatomical structures because of different RT planning. We propose the AnatomyNet that has one down-sampling layer with the trade-off between GPU memory and feature representation capacity, and 3D SE residual blocks for effective feature learning to alleviate these challenges. Moreover, we design a hybrid loss function with the Dice loss and the focal loss. The Dice loss is a class level distribution loss that depends less on the number of voxels in the anatomy, and the focal loss is designed to deal with highly unbalanced segmentation. For missing annotations, we propose masked loss and weighted loss for accurate and balanced weights updating in the learning of the AnatomyNet.\n\nResultsWe collect 261 HaN CT images to train the AnatomyNet, and use MICCAI Head and Neck Auto Segmentation Challenge 2015 as the benchmark dataset to evaluate the performance of the AnatomyNet. The objective is to segment nine anatomies: brain stem, chiasm, mandible, optic nerve left, optic nerve right, parotid gland left, parotid gland right, submandibular gland left, and submandibular gland right. Compared to previous state-of-the-art methods for each anatomy from the MICCAI 2015 competition, the AnatomyNet increases Dice similarity coefficient (DSC) by 3.3% on average. The proposed AnatomyNet takes only 0.12 seconds on average to segment a whole-volume HaN CT image of an average dimension of 178 x 302 x 225. All the data and code will be availablea.\n\nConclusion1We propose an end-to-end, fast and fully automated deep convolutional network, AnatomyNet, for accurate and whole-volume HaN anatomical segmentation. The proposed Anato-myNet outperforms previous state-of-the-art methods on the benchmark dataset. Extensive experiments demonstrate the effectiveness and good generalization ability of the components in the AnatomyNet.

bioengineering

Characterizing Activity and Thermostability of GH5 Cellulase Chimeras from Mesophilic and Thermophilic Parents

Cellulases from glycoside hydrolase (GH) family 5 are key enzymes in the degradation of diverse polysaccharide substrates and are used in industrial enzyme cocktails to break down biomass. The GH5 family shares a canonical ({beta})8-barrel structure, where each ({beta}) module is essential for the enzyme stability and activity. Despite their shared topology, the thermostability of GH5 enzymes can vary significantly, and highly thermostable variants are often sought for industrial applications. Based on a previously characterized thermophilic GH5 cellulase from Talaromyces emersonii (TeEgl5A, with an optimal temperature of 90{degrees}C), we created ten hybrid enzymes with the mesophilic cellulase from Prosthecium opalus (PoCel5) to determine which elements are responsible for enhanced thermostability. Five of the expressed hybrid enzymes exhibit enzyme activity. Two of these hybrids exhibited pronounced increases in the temperature optima (10 and 20{degrees}C), T50 (15 and 19{degrees}C), Tm (16.5 and 22.9{degrees}C), and extended half life, t1/2 (~240- and 650-fold at 55{degrees}C) relative to the mesophilic parent enzyme, and demonstrated improved catalytic efficiency on selected substrates. The successful hybridization strategies were validated experimentally in another GH5 cellulase from Aspergillus nidulans (AnCel5), which demonstrated a similar increase in thermostability. Based on molecular dynamics simulations (MD) of both PoCel5 and TeEgl5A parent enzymes as well as their hybrids, we hypothesize that improved hydrophobic packing of the interface between 2 and 3 is the primary mechanism by which the hybrid enzymes increase their thermostability relative to the mesophilic parent PoCel5.\n\nIMPORTANCEThermal stability is an essential property of enzymes in many industrial biotechnological applications, as high temperatures improve bioreactor throughput. Many protein engineering approaches, such as rational design and directed evolution, have been employed to improve the thermal properties of mesophilic enzymes. Structure-based recombination has also been used to fuse TIM-barrel fragments and even fragments from unrelated folds, to generate new structures. However, there are not many research on GH5 cellulases. In this study, two GH5 cellulases, which showed TIM-barrel structure, PoCel5 and TeEgl5A with different thermal properties were hybridized to study the roles of different ({beta}) motifs. This work illustrates the role that structure guided recombination can play in helping to identify sequence function relationships within GH5 enzymes by supplementing natural diversity with synthetic diversity.

bioengineering

BodyMap transcriptomes reveal unique circular RNA features across tissue types and developmental stages

Circular RNAs (circRNAs) are a novel class of regulatory RNAs. Here, we present a comprehensive investigation of circRNA expression profiles across 11 tissues and 4 developmental stages in rats, along with cross-species analyses in humans and mice. Although positively correlated, circRNAs exhibit higher tissue specificity than cognate mRNAs. Also, genes with higher expression levels exhibit a larger fraction of spliced circular transcripts than their linear counterparts. Intriguingly, while we observed a monotonic increase of circRNA abundance with age in the rat brain, we further discovered a dynamic, age-dependent pattern of circRNA expression in the testes that is characterized by a dramatic increase with advancing stages of sexual maturity and a decrease with aging. The age-sensitive testicular circRNAs are highly associated with spermatogenesis, independent of cognate mRNA expression. The tissue/age implications of circRNAs suggest that they present unique physiological functions rather than simply occurring as occasional by-products of gene transcription.

bioinformatics

A computational strategy for finding novel targets and therapeutic compounds for opioid dependence

Opioids are widely used for treating different types of pains, but overuse and abuse of prescription opioids have led to opioid epidemic in the United States. Besides analgesic effects, chronic use of opioid can also cause tolerance, dependence, and even addiction. Effective treatment of opioid addiction remains a big challenge today. Studies on addictive effects of opioids focus on striatum, a main component in the brain responsible for drug dependence and addiction. Some transcription regulators have been associated with opioid addiction, but relationship between analgesic effects of opioids and dependence behaviors mediated by them at the molecular level has not been thoroughly investigated. In this paper, we developed a new computational strategy that identifies novel targets and potential therapeutic molecular compounds for opioid dependence and addiction. We employed several statistical and machine learning techniques and identified differentially expressed genes over time which were associated with dependence-related behaviors after exposure to either morphine or heroin, as well as potential transcription regulators that regulate these genes, using time course gene expression data from mouse striatum. Moreover, our findings revealed that some of these dependence-associated genes and transcription regulators are known to play key roles in opioid-mediated analgesia and tolerance, suggesting that an intricate relationship between opioid-induce pain-related pathways and dependence may develop at an early stage during opioid exposure. Finally, we determined small compounds that can potentially target the dependence-associated genes and transcription regulators. These compounds may facilitate development of effective therapy for opioid dependence and addiction. We also built a database (http://daportals.org) for all opioid-induced dependence-associated genes and transcription regulators that we discovered, as well as the small compounds that target those genes and transcription regulators.

systems biology

DeepEM: Deep 3D ConvNets With EM For Weakly Supervised Pulmonary Nodule Detection

Recently deep learning has been witnessing widespread adoption in various medical image applications. However, training complex deep neural nets requires large-scale datasets labeled with ground truth, which are often unavailable in many medical image domains. For instance, to train a deep neural net to detect pulmonary nodules in lung computed tomography (CT) images, current practice is to manually label nodule locations and sizes in many CT images to construct a sufficiently large training dataset, which is costly and difficult to scale. On the other hand, electronic medical records (EMR) contain plenty of partial information on the content of each medical image. In this work, we explore how to tap this vast, but currently unexplored data source to improve pulmonary nodule detection. We propose DeepEM, a novel deep 3D ConvNet framework augmented with expectation-maximization (EM), to mine weakly supervised labels in EMRs for pulmonary nodule detection. Experimental results show that DeepEM can lead to 1.5% and 3.9% average improvement in free-response receiver operating characteristic (FROC) scores on LUNA16 and Tianchi datasets, respectively, demonstrating the utility of incomplete information in EMRs for improving deep learning algorithms.1

bioinformatics

Anxiety induced by extra-hypothalamic BDNF deficiency instigates resistance to diet-induced obesity

Anxiety disorders are associated with body weight changes in humans. However, mechanisms underlying anxiety-related weight changes remain poorly understood. Using Emx1Cre/+ mice, we deleted the gene for brain-derived neurotrophic factor (BDNF) in the cortex, hippocampus, and some parts of the amygdala. The resulting mutant mice displayed elevated anxiety levels and were markedly lean when fed either chow diet or high-fat diet (HFD). The mice showed higher levels of sympathetic activity, thermogenesis and lipolysis in both brown and white adipose tissues, and higher oxygen consumption and body temperature, compared with control mice. They were still lean at thermoneurality when fed HFD, indicating elevated basal metabolism in addition to activated thermogenesis. Anxiety induced by site-specific Bdnf deletion similarly increased energy expenditure and minimized HFD-induced weight gain. These results reveal that anxiety can stimulate adaptive thermogenesis and basal metabolism by activating sympathetic nervous system, which enhances lipolysis and limits weight gain.

physiology

Fusion expression and anti-Aspergillus flavus activity of a novel inhibitory protein DN-AflR

The regulatory gene (aflR) of aflatoxin encodes AflR, a positive regulator that activates transcriptional pathway of genes in aflatoxin biosynthesis. New L-Asp-L-Asn (DN) extracted from Bacillus megaterium inhibited the growth of A. flavus had been elucidated in our laboratory. The genes encoding DN and binuclear zinc finger cluster protein of AflR were fused, then fusion protein could compete with the AflS-AflR complex for the AflR binding site and significantly improve anti-A. flavus activity of DN. The fusion gene dn-aflR was cloned into pET32a and recombinant plasmid was introduced into Escherichia coli BL21. The highest expression was observed after 10 h induction and purified by affinity chromatography column. Compared with DN, the novel fusion protein DN-AflR significantly inhibited the growth of A. flavus and biosynthesis of aflatoxin B1. This study promoted the use of competitive inhibition of fusion proteins to reduce the expression of regulatory genes in the biosynthetic pathway of aflatoxin. Moreover, it provided more supports for deep research and industrialization of such novel, anti-A. flavus bio-inhibitors.\n\nIMPORTANCEAflatoxin contamination has seriously influence on export of agricultural products, income of farmers and economic development. Biological methods, especially using antagonistic microorganisms to inhibit aflatoxin biosynthesis gradually become the hot spot in recent years. DN (L-Asp-L-Asn) from Bacillus megaterium, which could inhibit growth of Aspergillus flavus and synthesis of aflatoxin, has been identified. In this report, we fused the genes encoding inhibitory peptides (DN) and specific zinc finger cluster protein, and expressed the novel anti-A. flavus protein in Escherichia coli. Compared with DN, the inhibitory ability of novel protein has been improved significantly. This research showed fusion expression of anti-fungal proteins, such as DN-AflR, is a promising method to economically improve the inhibitory activity of bio-inhibitors for A. flavus.

microbiology

Parallelized Inference for Single Cell Transcriptomic Clustering with Split Merge Sampling on DPMM Model

Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been applied to the data, they face challenges in the following aspects: (1) the clustering quality still needs to be improved; (2) most models need prior knowledge on number of clusters, which is not always available; (3) there is a demand for faster computational speed.\n\nResults: We propose to tackle these challenges with Parallelized Split Merge Sampling on Dirichlet Process Mixture Model (the Para-DPMM model). Unlike classic DPMM methods that perform sampling on each single data point, the split merge mechanism samples on the cluster level, which significantly improves convergence and optimality of the result. The model is highly parallelized and can utilize the computing power of high performance computing (HPC) clusters, enabling massive inference on huge datasets. Experiment results show the model achieves about 7% improvement in clustering accuracy for small datasets and more than 20% improvement for large challenging datasets compared with current widely used models. In the mean time, the models computing speed is significantly faster.\n\nAvailability: Source code is publicly available on https://github.com/tiehangd/Para_DPMM/tree/master/Para_DPMM_package

bioinformatics

Mechanistic insight into the interactions of NAP1 with NDP52 and TAX1BP1 for the recruitment of TBK1

NDP52 and TAX1BP1, two SKICH domain-containing autophagy recetpors, play crucial roles in selective autophagy. The autophagic functions of NDP52 and TAX1BP1 are regulated by TBK1, which can indirectly associate with them through the adaptor protein NAP1. However, the molecular mechanism governing the interactions of NAP1 with NDP52 and TAX1BP1 as well as the effects induced by TBK1-mediated phosphorylation of NDP52 and TAX1BP1 remain elusive. Here, we reported the first atomic structures of the SKICH regions of NDP52 and TAX1BP1 in complex with NAP1, which not only uncover the mechanismtic basis underpinning the specific interactions of NAP1 with NDP52 and TAX1BP1, but also reveal the first binding mode of a SKICH domain. Moreover, we demonstrated that the phosphorylation of TAX1BP1 SKICH mediated by TBK1 may regulate the interaction between TAX1BP1 and NAP1. In all, our findings provide mechanistic insights into the NAP1-mediated recruitments of TBK1 to NDP52 and TAX1BP1, and are valuable for further understanding the functions of these proteins in selective autophagy.

biophysics

A Near Complete Zonal Map of Mouse Olfactory Receptors

In the mouse olfactory system, spatially regulated expression of > 1,000 olfactory receptors (ORs) - a phenomenon termed "zones" - forms a topological map in the main olfactory epithelium (MOE). However, the zones of most ORs are currently unknown. By sequencing mRNA of 12 isolated MOE pieces, we mapped out zonal information for 1,033 OR genes with an estimated accuracy of 0.3 zones, covering 81% of all intact OR genes and 99.4% of total OR mRNA abundance. Zones tend to vary gradually along chromosomes. We further identified putative non-OR genes that may exhibit zonal expression.

neuroscience

Which are major players, canonical or non-canonical strigolactones?

Strigolactones (SLs) can be classified into two structurally distinct groups: canonical and non-canonical SLs. Canonical SLs contain the ABCD ring system, and non-canonical SLs lack the A, B, or C ring but have the enol ether-D ring moiety which is essential for biological activities. The simplest non-canonical SL is the SL biosynthetic intermediate carlactone (CL). In plants, CL and its oxidized metabolites such as carlactonoic acid and methyl carlactonoate, are present in root and shoot tissues. In some plant species including black oat (Avena strigosa), sunflower (Helianthus annuus), and maize (Zea mays), non-canonical SLs are major germination stimulants in the root exudates. Various plant species such as tomato (Solanum lycopersicum) release carlactonoic acid, and poplar (Populus spp.) was found to exude methyl carlactonoate into the rhizosphere. These results suggest that both canonical and non-canonical SLs are active as host recognition signals in the rhizosphere. In contrast, limited distribution of canonical SLs in the plant kingdom and structure- and stereo-specific transportation of canonical SLs from roots to shoots suggest that plant hormones inhibiting shoot branching are not canonical SLs but are rather non-canonical SLs.\n\nSupplemental files: Synthesis of 7-hydroxy-5-deoxystrigol stereoisomers, spectroscopic data of synthetic compounds and the natural stimulant in dokudami root exudates.\n\nScheme S1. Synthesis of racemic mixture of 7- and 7{beta}-hydroxy-5-deoxystrigol.\n\nFig. S1. 1H NMR spectrum of natural stimulant in dokudami root exudates.\n\nFig. S2. LC-MS/MS chromatograms of synthetic standards and natural stimulant.\n\nFig. S3. LC-MS/MS chromatograms of synthetic 7{beta}-hydroxy-5-deoxystrigol and natural stimulant.\n\nFig. S4. LC-MS/MS chromatograms of synthetic 7{beta}-hydroxy-5-deoxystrigol and natural stimulant.\n\nHighlightThe chemistry of canonical and non-canonical strigolactones and their distribution in the plant kingdom are summarized in relation to their biological activities in the rhizosphere and in plants.

plant biology

DeepLung: 3D Deep Convolutional Nets for Automated Pulmonary Nodule Detection and Classification

In this work, we present a fully automated lung CT cancer diagnosis system, DeepLung. DeepLung contains two parts, nodule detection and classification. Considering the 3D nature of lung CT data, two 3D networks are designed for the nodule detection and classification respectively. Specifically, a 3D Faster R-CNN is designed for nodule detection with a U-net-like encoder-decoder structure to effectively learn nodule features. For nodule classification, gradient boosting machine (GBM) with 3D dual path network (DPN) features is proposed. The nodule classification subnetwork is validated on a public dataset from LIDC-IDRI, on which it achieves better performance than state-of-the-art approaches, and surpasses the average performance of four experienced doctors. For the DeepLung system, candidate nodules are detected first by the nodule detection subnetwork, and nodule diagnosis is conducted by the classification subnetwork. Extensive experimental results demonstrate the DeepLung is comparable to the experienced doctors both for the nodule-level and patient-level diagnosis on the LIDC-IDRI dataset.

bioinformatics

FactorNet: a deep learning framework for predicting cell type specific transcription factor binding from nucleotide-resolution sequential data

Due to the large numbers of transcription factors (TFs) and cell types, querying binding profiles of all TF/cell type pairs is not experimentally feasible, owing to constraints in time and resources. To address this issue, we developed a convolutional-recurrent neural network model, called FactorNet, to computationally impute the missing binding data. FactorNet trains on binding data from reference cell types to make accurate predictions on testing cell types by leveraging a variety of features, including genomic sequences, genome annotations, gene expression, and single-nucleotide resolution sequential signals, such as DNase I cleavage. To the best of our knowledge, this is the first deep learning method to study the rules governing TF binding at such a fine resolution. With FactorNet, a researcher can perform a single sequencing assay, such as DNase-seq, on a cell type and computationally impute dozens of TF binding profiles. This is an integral step for reconstructing the complex networks underlying gene regulation. While neural networks can be computationally expensive to train, we introduce several novel strategies to significantly reduce the overhead. By visualizing the neural network models, we can interpret how the model predicts binding which in turn reveals additional insights into regulatory grammar. We also investigate the variables that affect cross-cell type predictive performance to explain why the model performs better on some TF/cell types than others, and offer insights to improve upon this field. Our method ranked among the top four teams in the ENCODE-DREAM in vivo Transcription Factor Binding Site Prediction Challenge.

genomics

The Effect Of Nipped-B-Like (Nipbl) Haploinsufficiency On Genome-Wide Cohesin Binding And Target Gene Expression: Modeling Cornelia de Lange Syndrome

Cornelia de Lange Syndrome (CdLS) is a multisystem developmental disorder frequently associated with heterozygous loss-of-function mutations of Nipped-B-like (NIPBL), the human homolog of Drosophila Nipped-B. NIPBL loads cohesin onto chromatin. Cohesin mediates sister chromatid cohesion important for mitosis, but is also increasingly recognized as a regulator of gene expression. In CdLS patient cells and animal models, the presence of multiple gene expression changes with little or no sister chromatid cohesion defect suggests that disruption of gene regulation underlies this disorder. However, the effect of NIPBL haploinsufficiency on cohesin binding, and how this relates to the clinical presentation of CdLS, has not been fully investigated. Nipbl haploinsufficiency causes CdLS-like phenotype in mice. We examined genome-wide cohesin binding and its relationship to gene expression using mouse embryonic fibroblasts (MEFs) from Nipbl +/- mice that recapitulate the CdLS phenotype. We found a global decrease in cohesin binding, including at CCCTC-binding factor (CTCF) binding sites and repeat regions. Cohesin-bound genes were found to be enriched for histone H3 lysine 4 trimethylation (H3K4me3) at their promoters; were disproportionately downregulated in Nipbl mutant MEFs; and displayed evidence of reduced promoter-enhancer interaction. The results suggest that gene activation is the primary cohesin function sensitive to Nipbl reduction. Over 50% of significantly dysregulated transcripts in mutant MEFs come from cohesin target genes, including genes involved in adipogenesis that have been implicated in contributing to the CdLS phenotype. Thus, decreased cohesin binding at the gene regions directly contributes to disease-specific expression changes. Taken together, our Nipbl haploinsufficiency model allows us to analyze the dosage effect of cohesin loading on CdLS development.

developmental biology

miR-200c Suppresses Stemness And Increases Cellular Sensitivity To Trastuzumab In HER2+ Breast Cancer

Resistance to trastuzumab remains a major obstacle in HER2-overexpressing breast cancer treatment. miR-200c is important for many functions in cancer stem cells (CSCs), including tumor recurrence, metastasis and resistance. We hypothesized that miR-200c contributes to trastuzumab resistance and stemness maintenance in HER2-overexpressing breast cancer. In this study, we used HER2-positive SKBR3, HER2-negative MCF-7, and their CD44+CD24- phenotype mammospheres SKBR3-S and MCF-7-S to verify. Our results demonstrated that miR-200c was weakly expressed in breast cancer cell lines and cell line stem cells. Overexpression of miR-200c resulted in a significant reduction in the number of tumor spheres formed and the population of CD44+CD24- phenotype mammospheres in SKBR3-S. Combining miR-200c with trastuzumab can significantly reduce proliferation and increase apoptosis of SKBR3 and SKBR3-S. Overexpression of miR-200c also eliminated its downstream target genes. These genes were highly expressed and positively related in breast cancer patients. Overexpression of miR-200c also improved the malignant progression of SKBR3-S and SKBR3 in vivo. miR-200c plays an important role in the maintenance of the CSC-like phenotype and increases drug sensitivity to trastuzumab in HER2+ cells and stem cells.\n\nSummary statementmiRNAs are critical in stemness maintenance and drug resistance. These data link maintenance of the stemness-related phenotype and the sensitivity of HER2+ breast cancer to miR-200c in response to trastuzumab.

cancer biology

Understanding sequence conservation with deep learning

MotivationComparing the human genome to the genomes of closely related mammalian species has been a powerful tool for discovering functional elements in the human genome. Millions of conserved elements have been discovered. However, understanding the functional roles of these elements still remain a challenge, especially in noncoding regions. In particular, it is still unclear why these elements are evolutionarily conserved and what kind of functional elements are encoded within these sequences.\n\nResultsWe present a deep learning framework, called DeepCons, to uncover potential functional elements within conserved sequences. DeepCons is a convolutional neural net (CNN) that receives a short segment of DNA sequence as input and outputs the probability of the sequence of being evolutionary conserved. DeepCons utilizes hundreds of convolution kernels to detect features within DNA sequences, and automatically learns these kernels after training the CNN model using 887,577 conserved elements and a similar number of nonconserved elements in the human genome. On a balanced test dataset, DeepCons can achieve an accuracy of 75% in determining whether a sequence element is conserved or not, and the area under the ROC curve of 0.83, based on information from the human genome alone. We further investigate the properties of the learned kernels. Some kernels are directly related to well-known regulatory motifs corresponding to transcription factors. Many kernels show positional biases relative to transcriptional start sites or transcription end sites. But most of discovered kernels do not correspond to any known functional element, suggesting that they might represent unknown categories of functional elements. We also utilize DeepCons to annotate how changes at each individual nucleotide might impact the conservation properties of the surrounding sequences.\n\nAvailabilityThe source code of DeepCons and all the learned convolution kernels in motif format is publicly available online at https://github.com/uci-cbcl/DeepCons.\n\nContactxhx@ics.uci.edu

bioinformatics

HLA class I binding prediction via convolutional neural networks

Many biological processes are governed by protein-ligand interactions. One such example is the recognition of self and non-self cells by the immune system. This immune response process is regulated by the major histocompatibility complex (MHC) protein which is encoded by the human leukocyte antigen (HLA) complex. Understanding the binding potential between MHC and peptides can lead to the design of more potent, peptide-based vaccines and immunotherapies for infectious autoimmune diseases.\n\nWe apply machine learning techniques from the natural language processing (NLP) domain to address the task of MHC-peptide binding prediction. More specifically, we introduce a new distributed representation of amino acids, name HLA-Vec, that can be used for a variety of downstream proteomic machine learning tasks. We then propose a deep convolutional neural network architecture, name HLA-CNN, for the task of HLA class I-peptide binding prediction. Experimental results show combining the new distributed representation with our HLA-CNN architecture acheives state-of-the-art results in the majority of the latest two Immune Epitope Database (IEDB) weekly automated benchmark datasets. We further apply our model to predict binding on the human genome and identify 15 genes with potential for self binding. Codes are available at https://github.com/uci-cbcl/HLA-bind.

bioinformatics

Adversarial Deep Structural Networks for Mammographic Mass Segmentation

Mass segmentation is an important task in mammogram analysis, providing effective morphological features and regions of interest (ROI) for mass detection and classification. Inspired by the success of using deep convolutional features for natural image analysis and conditional random fields (CRF) for structural learning, we propose an end-to-end network for mammographic mass segmentation. The network employs a fully convolutional network (FCN) to model potential function, followed by a CRF to perform structural learning. Because the mass distribution varies greatly with pixel position, the FCN is combined with position priori for the task. Due to the small size of mammogram datasets, we use adversarial training to control over-fitting. Four models with different convolutional kernels are further fused to improve the segmentation results. Experimental results on two public datasets, INbreast and DDSM-BCRP, show that our end-to-end network combined with adversarial training achieves the-state-of-the-art results.

bioinformatics