Search bioRxivSearch

Biology subjects

Michal Linial

Publications and source records attributed to Michal Linial.

3 recordsLinked to original sources

Single Cell Expression Data Reveal Human Genes that Escape X-Chromosome Inactivation

Sex chromosomes pose an inherent genetic imbalance between genders. In mammals, one of the females X-chromosomes undergoes inactivation (Xi). Indirect measurements estimate that about 20% of Xi genes completely or partially escape inactivation. The identity of these escapee genes and their propensity to escape inactivation remain unsolved. A direct method for identifying escapees was applied by quantifying differential allelic expression from single cells. RNA-Seq fragments were assigned to informative SNPs which were labeled by the appropriate parental haplotype. This method was applied for measuring allelic specific expression from Chromosome-X (ChrX) and an autosomal chromosome as a control. We applied the protocol for measuring biallelic expression from ChrX to 104 primary fibroblasts. Out of 215 genes that were considered, only 13 genes (6%) were associated with biallelic expression. The sensitivity of escapees' identification was increased by combining SNP mapping for parental diploid genomes together with RNA-Seq from clonal single cells (25 lymphoblasts). Using complementary protocols, referred to as strict and relaxed, we confidently identified 25 and 31escapee genes, respectively. When pooled versions of 30 and 100 cells were used, <50% of these genes were revealed. We assessed the generality of our protocols in view of an escapee catalog compiled from indirect methods. The overlap between the escapee catalog and the genes list from this study is statistically significant (P-value of E-07). We conclude that single cells expression data are instrumental for studying X-inactivation with an improved sensitivity. Finally, our results support the emerging notion of the non-deterministic nature of genes that escape X-chromosome inactivation.

Genetics

ASAP: A Machine-Learning Framework for Local Protein Properties

Determining residue level protein properties, such as the sites for post-translational modifications (PTMs) are vital to understanding proteins at all levels of function. Experimental methods are costly and time-consuming, thus high confidence predictions become essential for functional knowledge at a genomic scale. Traditional computational methods based on strict rules (e.g. regular expressions) fail to annotate sites that lack substantial similarity. Thus, Machine Learning (ML) methods become fundamental in annotating proteins with unknown function. We present ASAP (Amino-acid Sequence Annotation Prediction), a universal ML framework for residue-level predictions. ASAP extracts efficiently and fast large set of window-based features from raw sequences. The platform also supports easy integration of external features such as secondary structure or PSSM profiles. The features are then combined to train underlying ML classifiers. We present a detailed case study for ASAP that was used to train CleavePred, a state-of-the-art protein precursor cleavage sites predictor. Protein cleavage is a fundamental PTM shared by a wide variety of protein groups with minimal sequence similarity. Current computational methods have high false positive rates, making them suboptimal for this task. CleavePred has a simple Python API, and is freely accessible via a web-based application. The high performance of ASAP toward the task of precursor cleavage is suited for analyzing new proteomes at a genomic scale. The tool is attractive to protein design, mass spectrometry search engines and the discovery of new peptide hormones. In summary, we illustrate ASAP as an entry point for predicting PTMs. The approach and flexibility of the platform can easily be extended for additional residue specific tasks. ASAP and CleavePred source code available at https://github.com/ddofer/asap.

Bioinformatics

Overlapping Genes and Size Constraints in Viruses - An Evolutionary Perspective

Viruses are the simplest replicating units, characterized by a limited number of coding genes and an exceptionally high rate of overlapping genes. We sought a unified explanation for the evolutionary constraints that govern genome sizes, gene overlapping and capsid properties. We performed an unbiased statistical analysis over the [~]100 known viral families, and came to refute widespread assumptions regarding viral evolution. We found that the volume utilization of viral capsids is often low, and greatly varies among families. Most notably, we show that the total amount of gene overlapping is tightly bounded. Although viruses expand three orders of magnitude in genome length, their absolute amount of gene overlapping almost never exceeds 1500 nucleotides, and mostly confined to <4 significant overlapping instances. Our results argue against the common theory by which gene overlapping is driven by a necessity of viruses to compress their genome. Instead, we support the notion that overlapping has a role in gene novelty and evolution exploration.

Evolutionary Biology