Search bioRxivSearch

Biology subjects

Weng, Z.

Publications and source records attributed to Weng, Z..

15 recordsLinked to original sources

The RNA-binding ATPase, Armitage, Couples piRNA Amplification in Nuage to Phased piRNA Production on Mitochondria

PlWI-interacting RNAs (piRNAs) silence transposons in Drosophila ovaries, ensuring female fertility. Two coupled pathways generate germline piRNAs: the ping-pong cycle, in which the PIWI proteins Aubergine and Ago3 increase the abundance of pre-existing piRNAs, and the phased piRNA pathway, which generates strings of tail-to-head piRNAs, one after another. Proteins acting in the ping-pong cycle localize to nuage, whereas phased piRNA production requires Zucchini, an endonuclease on the mitochondrial surface. Here, we report that Armitage (Armi), an RNA-binding ATPase localized to both nuage and mitochondria, links the ping-pong cycle to the phased piRNA pathway. Mutations that block phased piRNA production deplete Armi from nuage. Armi ATPase mutants cannot support phased piRNA production and inappropriately bind mRNA instead of piRNA precursors. We propose that Armi shuttles between nuage and mitochondria, feeding precursor piRNAs generated by Ago3 cleavage into the Zucchini-dependent production of Aubergine- and Piwi-bound piRNAs on the mitochondrial surface.

genetics

Comprehensive Genomic Characterization of Breast Tumors with BRCA1 and BRCA2 Mutations

BackgroundGermline mutations in the BRCA1 and BRCA2 genes predispose carriers to breast and ovarian cancer, and there remains a need to identify the specific genomic mechanisms by which cancer evolves in these patients. Here we present a systematic genomic analysis of breast tumors with BRCA1 and BRCA2 mutations.\n\nMethodsWe analyzed genomic data from breast tumors, with a focus on comparing tumors with BRCA1/BRCA2 gene mutations with common classes of sporadic breast tumors.\n\nResultsWe identify differences between BRCA-mutated and sporadic breast tumors in patterns of point mutation, DNA methylation and structural variation. We show that structural variation disproportionately affects tumor suppressor genes and identify specific driver gene candidates that are enriched for structural variation.\n\nConclusionsCompared to sporadic tumors, BRCA-mutated breast tumors show signals of reduced DNA methylation, more ancestral cell divisions, and elevated rates of structural variation that tend to disrupt highly expressed protein-coding genes and known tumor suppressors. Our analysis suggests that BRCA-mutated tumors are more aggressive than sporadic breast cancers because loss of the BRCA pathway causes multiple processes of mutagenesis and gene dysregulation.

cancer biology

An Evolutionarily Conserved piRNA-producing Locus Required for Male Mouse Fertility

Pachytene piRNAs, which comprise >80% of small RNAs in the adult mouse testis, have been proposed to bind and regulate target RNAs like miRNAs, cleave targets like siRNAs, or lack biological function altogether. Although piRNA pathway protein mutants are male sterile, no biological function has been identified for any mammalian piRNA-producing locus. Here, we report that males lacking piRNAs from a conserved mouse pachytene piRNA locus on chromosome 6 (pi6) produce sperm with defects in capacitation and egg fertilization. Moreover, heterozygous embryos sired by pi6-/- fathers show reduced viability in utero. Molecular analyses suggest that pi6 piRNAs repress gene expression by cleaving mRNAs encoding proteins required for sperm function. pi6 also participates in a network of piRNA-piRNA precursor interactions that initiate piRNA production from a second piRNA locus on chromosome 10 as well as pi6 itself. Our data establish a direct role for pachytene piRNAs in spermiogenesis and embryo viability.\n\nHighlightsO_LINormal male mouse fertility and spermiogenesis require piRNAs from the pi6 locus\nC_LIO_LISperm capacitation and binding to the zona pellucida of the egg require pi6 piRNAs\nC_LIO_LIHeterozygous embryos sired by pi6-/- fathers show reduced viability in utero\nC_LIO_LIDefects in pi6 mutant sperm reflect changes in the abundance of specific mRNAs.\nC_LI

genetics

Maelstrom Represses Canonical Polymerase II Transcription within Bi-Directional piRNA Clusters in Drosophila melanogaster

In Drosophila, 23-30 nt long PIWI-interacting RNAs (piRNAs) direct the protein Piwi to silence germline transposon transcription. Most germline piRNAs derive from dual-strand piRNA clusters, heterochromatic transposon graveyards that are transcribed from both genomic strands. These piRNA sources are marked by the Heterochromatin Protein 1 homolog, Rhino (Rhi), which facilitates their promoter-independent transcription, suppresses splicing, and inhibits transcriptional termination. Here, we report that the protein Maelstrom (Mael) represses canonical, promoter-dependent transcription in dual-strand clusters, allowing Rhi to initiate piRNA precursor transcription. In addition to Mael, the piRNA biogenesis factors Armitage and Piwi, but not Rhi, are required to repress canonical transcription in dual-strand clusters. We propose that Armitage, Piwi, and Mael collaborate to repress potentially dangerous transcription of individual transposon mRNAs within clusters, while Rhi allows non-canonical transcription of the clusters into piRNA precursors without generating transposase-encoding mRNAs.

genetics

Integrating cross-linking experiments with ab initio protein-protein docking

Ab initio protein-protein docking algorithms often rely on experimental data to identify the most likely complex structure. We integrated protein-protein docking with the experimental data of chemical cross-linking followed by mass spectrometry. We tested our approach using 12 cases that resulted from an exhaustive search of the Protein Data Bank for protein complexes with cross-links identified in our experiments. We implemented cross-links as constraints based on Euclidean distance or void-volume distance. For most test cases the rank of the top-scoring near-native prediction was improved by at least two fold compared with docking without the cross-link information, and the success rates for the top 5 and top 10 predictions doubled. Our results demonstrate the delicate balance between retaining correct predictions and eliminating false positives. Several test cases had multiple components with distinct interfaces, and we present an approach for assigning cross-links to the interfaces. Employing the symmetry information for these cases further improved the performance of complex structure prediction.\n\nHighlightsO_LIIncorporating low-resolution cross-linking experimental data in protein-protein docking algorithms improves performance more than two fold.\nC_LIO_LIIntegration of protein-protein docking with chemical cross-linking reveals information on the configuration of higher order complexes.\nC_LIO_LISymmetry analysis of protein-protein docking results improves the predictions of multimeric complex structures\nC_LI

bioinformatics

Culture-free generation of microbial genomes from human and marine microbiomes

Our understanding of natural microbial communities is shaped by the careful investigation of a relatively small number of isolated and cultured organisms, and by analysis of genomic sequences obtained by culture-free metagenomic sequencing approaches. Metagenomic shotgun sequencing has facilitated partial reconstruction of strain-level community structure and functional repertoire. Unfortunately, it remains difficult to cost-effectively produce high quality genome drafts for individual microbes without isolation and culture. Recent molecular techniques that partition long DNA fragments and then barcode short fragments derived from them produce \"read clouds\", which are short-read sequences containing long-range information. Here, we present a novel application of a read cloud technique to microbiome samples, as well as Athena, a de novo assembler that uses these barcodes to produce improved metagenomic assemblies. We apply our approach to sequence human stool samples from two healthy individuals, and compare it to existing short read and synthetic long read metagenomic sequencing approaches. We find that read cloud metagenomic sequencing and Athena assembly produce the most complete individual genome drafts. These genome drafts are also highly contiguous (>200kb N50, <10 contigs), even for bacteria that have relatively low (20x) raw short read sequence coverage. We also apply this approach to a significantly more complex marine sediment sample and obtain 23 genome drafts with valuable 16S ribosomal RNA taxonomic marker sequences, nine of which are complete genome drafts. Read cloud metagenomic sequencing allows culture-free generation of high quality microbial genome drafts using only a single shotgun experiment.

genomics

Identification of piRNA binding sites reveals the Argonaute regulatory landscape of the C. elegans germline

piRNAs (Piwi-interacting small RNAs) engage Piwi Argonautes to silence transposons and promote fertility in animal germlines. Genetic and computational studies have suggested that C. elegans piRNAs tolerate mismatched pairing and in principle could target every transcript. Here we employ in vivo cross-linking to identify transcriptome-wide interactions between piRNAs and target RNAs. We show that piRNAs engage all germline mRNAs and that piRNA binding follows microRNA-like pairing rules. Targeting correlates better with binding energy than with piRNA abundance, suggesting that piRNA concentration does not limit targeting. In mRNAs silenced by piRNAs, secondary small RNAs accumulate at the center and ends of piRNA binding sites. In germline-expressed mRNAs, however, targeting by the CSR-1 Argonaute correlates with reduced piRNA binding density and suppression of piRNA-associated secondary small RNAs. Our findings reveal physiologically important and nuanced regulation of individual piRNA targets and provide evidence for a comprehensive post transcriptional regulatory step in germline gene expression.

genetics

The TRIM-NHL protein NHL-2 is a Novel Co-Factor of the CSR-1 and HRDE-1 22G-RNA Pathways

Proper regulation of germline gene expression is essential for fertility and maintaining species integrity. In the C. elegans germline, a diverse repertoire of regulatory pathways promote the expression of endogenous germline genes and limit the expression of deleterious transcripts to maintain genome homeostasis. Here we show that the conserved TRIM-NHL protein, NHL-2, plays an essential role in the C. elegans germline, modulating germline chromatin and meiotic chromosome organization. We uncover a role for NHL-2 as a co-factor in both positively (CSR-1) and negatively (HRDE-1) acting germline 22G-small RNA pathways and the somatic nuclear RNAi pathway. Furthermore, we demonstrate that NHL-2 is a bona fide RNA binding protein and, along with RNA-seq data point to a small RNA independent role for NHL-2 in regulating transcripts at the level of RNA stability. Collectively, our data implicate NHL-2 as an essential hub of gene regulatory activity in both the germline and soma.

genetics

Elimination of PCR duplicates in RNA-seq and small RNA-seq using unique molecular identifiers

RNA-seq and small RNA-seq are powerful, quantitative tools to study gene regulation and function. Common high-throughput sequencing methods rely on polymerase chain reaction (PCR) to expand the starting material, but not every molecule amplifies equally, causing some to be overrepresented. Unique molecular identifiers (UMIs) can be used to distinguish undesirable PCR duplicates derived from a single molecule and identical but biologically meaningful reads from different molecules. We have incorporated UMIs into RNA-seq and small RNA-seq protocols and developed tools to analyze the resulting data. Our UMIs contain stretches of random nucleotides whose lengths sufficiently capture diverse molecule species in both RNA-seq and small RNA-seq libraries generated from mouse testis. Our approach yields high-quality data while allowing unique tagging of all molecules in high-depth libraries. Using simulated and real datasets, we demonstrate that our methods increase the reproducibility of RNA-seq and small RNA-seq data. Notably, we find that the amount of starting material and sequencing depth, but not the number of PCR cycles, determine PCR duplicate frequency. Finally, we show that computational removal of PCR duplicates based only on their mapping coordinates introduces substantial bias into data analysis.

genomics

Transcriptome-wide analysis of the functional intronome using spliceosome profiling

Full understanding of eukaryotic transcriptomes and how they respond to different conditions requires deep knowledge of all sites of intron excision. Although RNA-Seq provides much of this information, the low abundance of many spliced transcripts (often due to their rapid cytoplasmic decay) limits the ability of RNA-Seq alone to reveal the full repertoire of spliced species. Here we present \"spliceosome profiling\", a strategy based on deep sequencing of RNAs co-purifying with late stage spliceosomes. Spliceosome profiling allows for unambiguous mapping of intron ends to single nucleotide resolution and branchpoint identification at unprecedented depths. Our data reveal hundreds of new introns in S. pombe and numerous others that were previously misannotated. By providing a means to directly interrogate sites of spliceosome assembly and catalysis genome-wide, spliceosome profiling promises to transform our understanding of RNA processing in the nucleus much like ribosome profiling has transformed our understanding mRNA translation in the cytoplasm.

genomics

Genome-Wide Identification of Early-Firing Human Replication Origins by Optical Replication Mapping

The timing of DNA replication is largely regulated by the location and timing of replication origin firing. Therefore, much effort has been invested in identifying and analyzing human replication origins. However, the heterogeneous nature of eukaryotic replication kinetics and the low efficiency of individual origins in metazoans has made mapping the location and timing of replication initiation in human cells difficult. We have mapped early-firing origins in HeLa cells using Optical Replication Mapping, a high-throughput single-molecule approach based on Bionano Genomics genomic mapping technology. The single-molecule nature and 290-fold coverage of our dataset allowed us to identify origins that fire with as little as 1% efficiency. We find sites of human replication initiation in early S phase are not confined to well-defined efficient replication origins, but are instead distributed across broad initiation zones consisting of many inefficient origins. These early-firing initiation zones co-localize with initiation zones inferred from Okazaki-fragment-mapping analysis and are enriched in ORC1 binding sites. Although most early-firing origins fire in early-replication regions of the genome, a significant number fire in late-replicating regions, suggesting that the major difference between origins in early and late replicating regions is their probability of firing in early S-phase, as opposed to qualitative differences in their firing-time distributions. This observation is consistent with stochastic models of origin timing regulation, which explain the regulation of replication timing in yeast.

genomics

The genome of Trichoplusia ni, an agricultural pest and novel model for small RNA biology

The cabbage looper, Trichoplusia ni (Lepidoptera: Noctuidae), is a destructive insect pest that feeds on a wide range of plants. The High Five cell line (Hi5), originally derived from T. ni ovaries, is often used for efficient expression of recombinant proteins. Here, we report a draft assembly of the 368.2 Mb T. ni genome, with 90.6% of all bases assigned to one of its 28 chromosomes and predicted 14,037 predicted protein-coding genes. Manual curation of gene families involved in chemoreception and detoxification reveals T. ni-specific gene expansions that may explain its widespread distribution and rapid adaptation to insecticides. Using male and female genome sequences, we define Z-linked and repeat-rich W-linked sequences. Transcriptome and small RNA data from T. ni thorax, ovary, testis, and Hi5 cells reveal distinct expression profiles for 295 microRNA- and >393 piRNA-producing loci, as well as 39 genes encoding core small RNA pathway proteins. siRNAs target both endogenous transposons and the exogenous TNCL virus. Surprisingly, T. ni siRNAs are not 2{acute}-O-methylated. Five piRNA-producing loci account for 34.9% piRNAs in the ovary, 49.3% piRNAs in the testis, and 44.0% piRNAs in Hi5 cells. Nearly all of the W chromosome is devoted to piRNA production: >76.0% of bases in the assembled W produce piRNAs in ovary. To enable use of the T. ni germline-derived Hi5 cell line as a model system, we have established efficient genome editing and single-cell cloning protocols. Taken together, the T. ni genome provides insights into pest control and allows Hi5 cells to become a new tool for studying small RNAs ex vivo.

genomics

Human PGBD5 DNA transposase promotes site-specific oncogenic mutations in rhabdoid tumors

Genomic rearrangements are a hallmark of childhood solid tumors, but their mutational causes remain poorly understood. Here, we identify the piggyBac transposable element derived 5 (PGBD5) gene as an enzymatically active human DNA transposase expressed in the majority of rhabdoid tumors, a lethal childhood cancer. Using assembly-based whole-genome DNA sequencing, we observed previously unknown somatic genomic rearrangements in primary human rhabdoid tumors. These rearrangements were characterized by deletions and inversions involving PGBD5-specific signal (PSS) sequences at their breakpoints, with some recurrently targeting tumor suppressor genes, leading to their inactivation. PGBD5 was found to be physically associated with human genomic PSS sequences that were also sufficient to mediate PGBD5-induced DNA rearrangements in rhabdoid tumor cells. We found that ectopic expression of PGBD5 in primary immortalized human cells was sufficient to promote penetrant cell transformation in vitro and in immunodeficient mice in vivo. This activity required specific catalytic residues in the PGBD5 transposase domain, as well as end-joining DNA repair, and induced distinct structural rearrangements, involving PSS-associated breakpoints, similar to those found in primary human rhabdoid tumors. This defines PGBD5 as an oncogenic mutator and provides a plausible mechanism for site-specific DNA rearrangements in childhood and adult solid tumors.

cancer biology

A unified encyclopedia of human functional DNA elements through fully automated annotation of 164 human cell types

Semi-automated genome annotation methods such as Segway enable understanding of chromatin activity. Here we present chromatin state annotations of 164 human cell types using 1,615 genomics data sets. To produce these annotations, we developed a fully-automated annotation strategy in which we train separate unsupervised annotation models on each cell type and use a machine learning classifier to automate the state interpretation step. Using these annotations, we developed a measure of the importance of each genomic position called the \"conservation-associated activity score,\" which we use to aggregate information across cell types into a multi-cell type view. The aggregated conservation-associated activity score provides a measure of importance directly attributable to a specific activity in a specific set of cell types. In contrast to evolutionary conservation, this measure is not biased to detect only elements shared with related species. Using the conservation-associated activity score, we combined all our annotations into a single, cell type-agnostic encyclopedia that catalogs all human transcriptional and regulatory elements, enabling easy and intuitive interpretation of the effect of genome variants on phenotype, such as in disease-associated, evolutionarily conserved or positively selected loci. These resources, including cell type-specific annotations, encyclopedia, and a visualization server, are available at http://noble.gs.washington.edu/proj/encyclopedia.\n\nAuthor SummaryGenome annotation algorithms are an effective class of tools for understanding the function of the genome. These algorithms take as input a set of genome-wide measurements about the activity at each base pair in a given tissue, such as where a given protein is binding or how accessible the DNA is to being read by a protein. The genome is then partitioned and each segment is assigned a label such that positions with the same label exhibit similar patterns in the input data. Such annotations are widely used for many applications, such as to understand the mechanism of impact of a given genetic variant. Here we present, to our knowledge, the most comprehensive set of genome annotations created so far, encompassing 164 human cell types and including 1,615 genomics data sets. These comprehensive annotations are made possible by a strategy that automates the previous interpretation step. Furthermore, we present several methodological innovations that make these genome annotations more useful.

genomics

LR-DNase: Predicting TF binding from DNase-seq data

Transcription factors play a key role in the regulation of gene expression. Hypersensitivity to DNase I cleavage has long been used to gauge the accessibility of genomic DNA for transcription factor binding and as an indicator of regulatory genomic locations. An increasing amount of ChIP-seq data on a large number of TFs is being generated, mostly in a small number of cell types. DNase-seq data are being produced for hundreds of cell types. We aimed to develop a computational method that could combine ChIP-seq and DNase-seq data to predict TF binding sites in a wide variety of cell types. We trained and tested a logistic regression model, called LR-DNase, to predict binding sites for a specific TF using seven features derived from DNase-seq and genomic sequence. We calculated the area under the precision-recall curve at a false discovery rate cutoff of 0.5 for the LR-DNase model, a number of logistic regression models with fewer features, and several existing state-of-the-art TF binding prediction methods. The LR-DNase model outperformed existing unsupervised and supervised methods. Additionally, for many TFs, a model that uses only two features, DNase-seq reads and motif score, was sufficient to match the performance of the best existing methods.

bioinformatics