Search bioRxivSearch

Biology subjects

Mahajan, S.

Publications and source records attributed to Mahajan, S..

3 recordsLinked to original sources

Massive gene amplification on a recently formed Drosophila Y chromosome

Widespread loss of genes on the Y is considered a hallmark of sex chromosome differentiation. Here we show that the initial stages of Y evolution are driven by massive amplification of distinct classes of genes. The neo-Y chromosome of Drosophila miranda initially contained about 3000 protein-coding genes, but has gained over 3200 genes since its formation about 1.5 MY ago, primarily by tandem amplification of protein-coding genes ancestrally present on this chromosome. We show that distinct evolutionary processes may account for this drastic increase in gene number on the Y. Testis-specific and dosage sensitive genes appear to have amplified on the Y to increase male fitness. A distinct class of meiosis-related multi-copy Y genes independently co-amplified on the X, and their expansion is likely driven by conflicts over segregation. Co-amplified X/Y genes are highly expressed in testis, enriched for meiosis and RNAi functions, and are frequently targeted by small RNAs in testis. This suggests that their amplification is driven by X vs. Y antagonism for increased transmission, where sex chromosome drive suppression is likely mediated by sequence homology between the suppressor and distorter, through RNAi mechanism. Thus, our analysis suggests that newly emerged sex chromosomes are a battleground for sexual and meiotic conflict.

evolutionary biology

De novo assembly of a young Drosophila Y chromosome using Single-Molecule sequencing and Chromatin Conformation capture

While short-read sequencing technology has resulted in a sharp increase in the number of species with genome assemblies, these assemblies are typically highly fragmented. Repeats pose the largest challenge for reference genome assembly, and pericentromeric regions and the repeat-rich Y chromosome are typically ignored from sequencing projects. Here, we assemble the genome of Drosophila miranda using long reads for contig formation, chromatin interaction maps for scaffolding and short reads, optical mapping and BAC clone sequencing for consensus validation. Our assembly recovers entire chromosomes and contains large fractions of repetitive DNA, including ~41.5 Mb of pericentromeric and telomeric regions, and >100Mb of the recently formed highly repetitive neo-Y chromosome. While Y chromosome evolution is typically characterized by global sequence loss and shrinkage, the neo-Y increased in size by almost 3-fold, due to the accumulation of repetitive sequences. Our high-quality assembly allows us to reconstruct the chromosomal events that have led to the unusual sex chromosome karyotype in D. miranda, including the independent de novo formation of a pair of sex chromosomes at two distinct time points, or the reversion of a former Y chromosome to an autosome.

genomics

PB-kPRED: Knowledge-Based Prediction Of Protein Backbone Conformation Using A Structural Alphabet

Libraries of structural prototypes that abstract protein local structures are known as structural alphabets and have proven to be very useful in various aspects of protein structure analyses and predictions. One such library, Protein Blocks (PBs), is composed of 16 standard 5-residues long structural prototypes. This form of analyzing proteins involves drafting its structure as a string of PBs. Thus, predicting the local structure of a protein in terms of protein blocks is a step towards the objective of predicting its 3-D structure. Here a new approach, kPred, is proposed towards this aim that is independent of the evolutionary information available. It involves (i) organizing the structural knowledge in the form of a database of pentapeptide fragments extracted from all protein structures in the PDB and (ii) apply a purely knowledge-based algorithm, not relying on secondary structure predictions or sequence alignment profiles, to scan this database and predict most probable backbone conformations for the protein local structures.\n\nBased on the strategy used for scanning the database, the method was able to achieve efficient mean Q16 accuracies between 40.8% and 66.3% for a non-redundant subset of the PDB filtered at 30% sequence identity cut-off. The impact of these scanning strategies on the prediction was evaluated and is discussed. A scoring function that gives a good estimate of the accuracy of prediction was further developed. This score estimates very well the accuracy of the algorithm (R2 of 0.82). An online version of the tool is provided freely for non-commercial usage at http://www.bo-protscience.fr/kpred/.

bioinformatics