Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Identification of five patterns of nucleosome positioning that globally describe transcription factor function

Following the binding of transcription factors (TF) to specific regions, chromatin remodeling including alterations in nucleosome positioning (NP) occurs. These changes in NP cause selective gene expression to determine cell function. However whether specific NP patterns upon TF binding determine the transcriptional regulation such as gene activation or suppression is unclear. Here we identified five patterns of NP around TF binding sites (TFBSs) using fixed MNase-Seq analysis. The most frequently observed NP pattern described the transcription state. The five patterns explained approximately 80% of the whole NP pattern on the genome in mouse C2C12 cells. We further performed ChIP-Seq using the input obtained from the fixed MNase-Seq. The result showed that a single trial of ChIP-Seq could visualize the NP patterns around the TFBS and predict the function of the transcriptional regulation at the same time. These findings indicate that NP can directly predict the function of TFs.

Molecular Biology

Neural lineage induction reveals multi-scale dynamics of 3D chromatin organization

Regulation of gene expression underlies cell identity. Chromatin structure and gene activity are linked at multiple levels, via positioning of genomic loci to transcriptionally permissive or repressive environments and by connecting cis-regulatory elements such as promoters and enhancers. However, the genome-wide dynamics of these processes during cell differentiation has not been characterized. Using tethered chromatin conformation capture (TCC) sequencing we determined global three-dimensional chromatin structures in mouse embryonic stem (ES) and neural stem (NS) cell derivatives. We found that changes in the propensity of genomic regions to form inter-chromosomal contacts are pervasive in neural induction and are associated with the regulation of gene expression. Moreover, we found a pronounced contribution of euchromatic domains to the intra-chromosomal interaction network of pluripotent cells, indicating the existence of an ES cell-specific mode of chromatin organization. Mapping of promoter-enhancer interactions in pluripotent and differentiated cells revealed that spatial proximity without enhancer element activity is a common architectural feature in cells undergoing early developmental changes. Activity-independent formation of higher-order contacts between cis-regulatory elements, predominant at complex loci, may thus provide an additional layer of transcriptional control.

Molecular Biology

Comprehensive mutational scanning of a kinase in vivo reveals substrate-dependent fitness landscapes

Deep mutational scanning has emerged as a promising tool for mapping sequence-activity relationships in proteins1-4, RNA5 and DNA6-8. In this approach, diverse variants of a sequence of interest are first ranked according to their activities in a relevant pooled assay, and this ranking is then used to infer the shape of the fitness landscape around the wild-type sequence. Little is currently know, however, about the degree to which such fitness landscapes are dependent on the specific assay conditions from which they are inferred. To explore this issue, we performed deep mutational scanning of APH(3)II, a Tn5 transposon-derived kinase that confers resistance to aminoglycoside antibiotics9, in E. coli under selection with each of six structurally diverse antibiotics at a range of inhibitory concentrations. We found that the resulting fitness landscapes showed significant dependence on both antibiotic structure and concentration. This shows that the notion of essential amino acid residues is context-dependent, but also that this dependence can be exploited to guide protein engineering. Specifically, we found that differential analysis of fitness landscapes allowed us to generate synthetic APH(3)II variants with orthogonal substrate specificities.

Molecular Biology

A genome-wide analysis of Cas9 binding specificity using ChIP-seq and targeted sequence capture

Clustered regularly interspaced short palindromic repeat (CRISPR) RNA-guided nucleases have gathered considerable excitement as a tool for genome engineering. However, questions remain about the specificity of their target site recognition. Most previous studies have examined predicted off-target binding sites that differ from the perfect target site by one to four mismatches, which represent only a subset of genomic regions. Here, we use ChIP-seq to examine genome-wide CRISPR binding specificity at gRNA-specific and gRNA-independent sites. For two guide RNAs targeting the murine Snurf gene promoter, we observed very high binding specificity at the intended target site while off-target binding was observed at 2- to 6-fold lower intensities. We also identified significant gRNA-independent off-target binding. Interestingly, we found that these regions are highly enriched in the PAM site, a sequence required for target site recognition by CRISPR. To determine the relationship between Cas9 binding and endonuclease activity, we used targeted sequence capture as a high-throughput approach to survey a large number of the potential off-target sites identified by ChIP-seq or computational prediction. A high frequency of indels was observed at both target sites and one off-target site, while no cleavage activity could be detected at other ChIP-bound regions. Our data is consistent with recent finding that most interactions between the CRISPR nuclease complex and genomic PAM sites are transient and do not lead to DNA cleavage. The interactions are stabilized by gRNAs with good matches to the target sequence adjacent to the PAM site, resulting in target cleavage activity.

Molecular Biology

ADP-ribose derived Nuclear ATP is Required for Chromatin Remodeling and Hormonal Gene Regulation

Highlights- Hormonal gene regulation requires synthesis of PAR and its degradation to ADP-ribose by PARG\n- ADP-ribose is converted to ATP in the cell nuclei by hormone-activated NUDIX5/NUDT5\n- Blocking nuclear ATP formation precludes hormone-induced chromatin remodeling, gene regulation and cell proliferation\n\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=165 HEIGHT=200 SRC=\"FIGDIR/small/006593v1_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (46K):\norg.highwire.dtl.DTLVardef@e2c9fborg.highwire.dtl.DTLVardef@13ab596org.highwire.dtl.DTLVardef@16799d7org.highwire.dtl.DTLVardef@a53086_HPS_FORMAT_FIGEXP M_FIG C_FIG SummaryKey nuclear processes in eukaryotes including DNA replication or repair and gene regulation require extensive chromatin remodeling catalyzed by energy consuming enzymes. How the energetic demands of such processes are ensured in response to rapid stimuli remains unclear. We have analyzed this question in the context of the massive gene regulation changes induced by progestins in breast cancer cells and found that ATP is generated in the cell nucleus via the hydrolysis of poly-ADP-ribose to ADP-ribose. Nuclear ATP synthesis requires the combined enzymatic activities of PARP1, PARG and NUDIX5/NUDT5. Although initiated via mitochondrial derived ATP, the nuclear source of ATP is essential for hormone induced chromatin remodeling, gene regulation and cell proliferation and may also participate in DNA repair. This novel pathway reveals exciting avenues of research for drug development.

Molecular Biology

Reagent contamination can critically impact sequence-based microbiome analyses

The study of microbial communities has been revolutionised in recent years by the widespread adoption of culture independent analytical techniques such as 16S rRNA gene sequencing and metagenomics. One potential confounder of these sequence-based approaches is the presence of contamination in DNA extraction kits and other laboratory reagents. In this study we demonstrate that contaminating DNA is ubiquitous in commonly used DNA extraction kits, varies greatly in composition between different kits and kit batches, and that this contamination critically impacts results obtained from samples containing a low microbial biomass. Contamination impacts both PCR based 16S rRNA gene surveys and shotgun metagenomics. These results suggest that caution should be advised when applying sequence-based techniques to the study of microbiota present in low biomass environments. We provide an extensive list of potential contaminating genera, and guidelines on how to mitigate the effects of contamination. Concurrent sequencing of negative control samples is strongly advised.

Molecular Biology

Conditional U1 Gene Silencing in Toxoplasma gondii

In absence of powerful siRNA approaches, the functional characterisation of essential genes in apicomplexan parasites, such as Toxoplasma gondii or Plasmodium falciparum, relies on conditional mutagenesis systems. Here we present a novel strategy based on U1 snRNP-mediated gene silencing. U1 snRNP is critical in pre-mRNA splicing by defining the exonintron boundaries. When a U1 recognition site is placed into the 3-terminal exon or adjacent to the termination codon, pre-mRNA is cleaved at the 3-end and degraded, leading to an efficient knockdown of the gene of interest (GOI). Here we describe a simple one-step approach that combines endogenous tagging with DiCre-mediated positioning of U1 recognition sites adjacent to the termination codon of the GOI which leads to a conditional knockdown of the GOI in Ku80 knockout and RH T. gondii tachyzoites. Specific knockdown mutants of the reporter gene GFP and several endogenous genes of T. gondii including the clathrin heavy chain gene 1 (chc1), the vacuolar protein sorting gene 26 (vps26), and the dynamin-related protein C gene (drpC) were silenced using this new approach. This new gene silencing tool kit allows protein tracking and functional studies simultaneously.

Molecular Biology

Microbial community composition and diversity via 16S rRNA gene amplicons: evaluating the Illumina platform

As new sequencing technologies become cheaper and older ones disappear, laboratories switch vendors and platforms. Validating the new setups is a crucial part of conducting rigorous scientific research. Here we report on the reliability and biases of performing bacterial 16S rRNA gene amplicon paired-end sequencing on the MiSeq Illumina platform. We designed a protocol using 50 barcode pairs to run samples in parallel and coded a pipeline to process the data. Sequencing the same sediment sample in 248 replicates as well as 70 samples from alkaline soda lakes, we evaluated the performance of the method with regards to estimates of alpha and beta diversity.\n\nUsing different purification and DNA quantification procedures we always found up to 5-fold differences in the yield of sequences between individually barcodes samples. Using either a one-step or a two-step PCR preparation resulted in significantly different estimates in both alpha and beta diversity. Comparing with a previous method based on 454 pyrosequencing, we found that our Illumina protocol performed in a similar manner - with the exception for evenness estimates where correspondence between the methods was low.\n\nWe further quantified the data loss at every processing step eventually accumulating to 50% of the raw reads. When evaluating different OTU clustering methods, we observed a stark contrast between the results of QIIME with default settings and the more recent UPARSE algorithm when it comes to the number of OTUs generated. Still, overall trends in alpha and beta diversity corresponded highly using both clustering methods.\n\nOur procedure performed well considering the precisions of alpha and beta diversity estimates, with insignificant effects of individual barcodes. Comparative analyses suggest that 454 and Illumina sequence data can be combined if the same PCR protocol and bioinformatic workflows are used for describing patterns in richness, beta-diversity and taxonomic composition.

Molecular Biology

Resolving microbial microdiversity with high accuracy full length 16S rRNA Illumina sequencing

We describe a method for sequencing full-length 16S rRNA gene amplicons using the high throughput Illumina MiSeq platform. The resulting sequences have about 100-fold higher accuracy than standard Illumina reads and are chimera filtered using information from a single molecule dual tagging scheme that boosts the signal available for chimera detection. We demonstrate that the data provides fine scale phylogenetic resolution not available from Illumina amplicon methods targeting smaller variable regions of the 16S rRNA gene.

Molecular Biology

Polymerase ζ activity is linked to replication timing in humans: evidence from mutational signatures

Replication timing is an important determinant of germline mutation patterns, with a higher rate of point mutations in late replicating regions. Mechanisms underlying this association remain elusive. One of the suggested explanations is the activity of error-prone DNA polymerases in late-replicating regions. Polymerase {zeta} (pol {zeta}), an essential error-prone polymerase biased towards transversions, also has a tendency to produce dinucleotide mutations (DNMs), complex mutational events that simultaneously affect two adjacent nucleotides. Experimental studies have shown that pol {zeta} is strongly biased towards GC->AA/TT DNMs. Using primate divergence data, we show that the GC->AA/TT pol {zeta} mutational signature is the most frequent among DNMs, and its rate exceeds the mean rate of other DNM types by a factor of ~10. Unlike the overall rate of DNMs, the pol {zeta} signature drastically increases with the replication time in the human genome. Finally, the pol {zeta} signature is enriched in transcribed regions, and there is a strong prevalence of GC->TT over GC->AA DNMs on the non-template strand, indicating association with transcription. A recurrently occurring GC->TT DNM in HRAS gene causes the Costello syndrome; we find a 2-fold increase in the mutation rate, and a 2-fold decrease in the transition/transversion ratio, at distances of up to 1 kb from the DNM, suggesting a link between the Costello syndrome and pol {zeta} activity. This study uncovers the genomic preferences of pol {zeta}, shedding light on a novel cause of mutational heterogeneity along the genome.

Molecular Biology

Pollen feeding proteomics: salivary proteins of the passion flower butterfly, Heliconius melpomene

While most adult Lepidoptera use flower nectar as their primary food source, butterflies in the genus Heliconius have evolved the novel ability to acquire amino acids from consuming pollen. Heliconius butterflies collect pollen on their proboscis, moisten the pollen with saliva, and use a combination of mechanical disruption and chemical degradation to release free amino acids that are subsequently re-ingested in the saliva. Little is known about the molecular mechanisms of this complex pollen feeding adaptation. Here we report an initial shotgun proteomic analysis of saliva from Heliconius melpomene. Results from liquid-chromatography tandem mass-spectrometry confidently identified 31 salivary proteins, most of which contained predicted signal peptides, consistent with extracellular secretion. Further bioinformatic annotation of these salivary proteins indicated the presence of four distinct functional classes: proteolysis (10 proteins), carbohydrate hydrolysis (5), immunity (6), and \"housekeeping\"(4). Additionally, six proteins could not be functionally annotated beyond containing a predicted signal sequence. The presence of several salivary proteases is consistent with previous demonstrations that Heliconius saliva has proteolytic capacity. It is likely these proteins play a key role in generating free amino acids during pollen digestion. The identification of proteins functioning in carbohydrate hydrolysis is consistent with Heliconius butterflies consuming nectar, like other lepidopterans, as well as pollen. Immune-related proteins in saliva are also expected, given that ingestion of pathogens is a very likely route to infection. The few \"housekeeping\" proteins are likely not true salivary proteins and reflect a modest level of contamination that occurred during saliva collection. Among the unannotated proteins were two sets of paralogs, each seemingly the result of a relatively recent tandem duplication. These results offer a first glimpse into the molecular foundation of Heliconius pollen feeding and provide a substantial advance towards comprehensively understanding this striking evolutionary novelty.

Molecular Biology

Strong spurious transcription likely a cause of DNA insert bias in typical metagenomic clone libraries

BackgroundClone libraries provide researchers with a powerful resource with which to study nucleic acid from diverse sources. Metagenomic clone libraries in particular have aided in studies of microbial biodiversity and function, as well as allowed the mining of novel enzymes for specific functions of interest. These libraries are often constructed by cloning large-inserts ([~]30 kb) into a cosmid or fosmid vector. Recently, there have been reports of GC bias in fosmid metagenomic clone libraries, and it was speculated that the bias may be a result of fragmentation and loss of AT-rich sequences during the cloning process. However, evidence in the literature suggests that transcriptional activity or gene product toxicity may play a role in library bias.\n\nResultsTo explore the possible mechanisms responsible for sequence bias in clone libraries, and in particular whether fragmentation is involved, we constructed a cosmid clone library from a human microbiome sample, and sequenced DNA from three different steps of the library construction process: crude extract DNA, size-selected DNA, and cosmid library DNA. We confirmed a GC bias in the final constructed cosmid library, and we provide strong evidence that the sequence bias is not due to fragmentation and loss of AT-rich sequences but is likely occurring after the DNA is introduced into E. coli. To investigate the influence of strong constitutive transcription, we searched the sequence data for consensus promoters and found that rpoD/{sigma}70 promoter sequences were underrepresented in the cosmid library. Furthermore, when we examined the reference genomes of taxa that were differentially abundant in the cosmid library relative to the original sample, we found that the bias appears to be more closely correlated with the number of rpoD/{sigma}70 consensus sequences in the genome than with simple GC content.\n\nConclusionsThe GC bias of metagenomic clone libraries does not appear to be due to DNA fragmentation. Rather, analysis of promoter consensus sequences provides support for the hypothesis that strong constitutive transcription from sequences recognized as rpoD/{sigma}70 consensus-like in E. coli may lead to plasmid instability or loss of insert DNA. Our results suggest that despite widespread use of E. coli to propagate foreign DNA, the effects of in vivo transcriptional activity may be under-appreciated. Further work is required to tease apart the effects of transcription from those of gene product toxicity.

Molecular Biology

Folding of Aquaporin 1: Multiple evidence that helix 3 can shift out of the membrane core

The folding of most integral membrane proteins follows a two-step process: Initially, individual transmembrane helices are inserted into the membrane by the Sec translocon. Thereafter, these helices fold to shape the final conformation of the protein. However, for some proteins, including Aquaporin 1 (AQP1), the folding appears to follow a more complicated path. AQP1 has been reported to first insert as a four-helical intermediate, where helix 2 and 4 are not inserted into the membrane. In a second step this intermediate is folded into a six-helical topology. During this process, the orientation of the third helix is inverted. Here, we propose a mechanism for how this reorientation could be initiated: First, helix 3 slides out from the membrane core resulting in that the preceding loop enters the membrane. The final conformation could then be formed as helix 2, 3 and 4 are inserted into the membrane and the reentrant regions come together. We find support for the first step in this process by showing that the loop preceding helix 3 can insert into the membrane. Further, hydrophobicity curves, experimentally measured insertion efficiencies and MD-simulations suggest that the barrier between these two hydrophobic regions is relatively low, supporting the idea that helix 3 can slide out of the membrane core, initiating the rearrangement process

Molecular Biology

The positive inside rule is stronger when followed by a transmembrane helix.

The translocon recognizes transmembrane helices with sufficient level of hydrophobicity and inserts them into the membrane. However, sometimes less hydrophobic helices are also recognized. Positive inside rule, orientational preferences of and specific interactions with neighboring helices have been shown to aid in the recognition of these helices, at least in artificial systems. To better understand how the translocon inserts marginally hydrophobic helices, we studied three naturally occurring marginally hydrophobic helices, which were previously shown to require the subsequent helix for efficient translocon recognition. We find no evidence for specific interactions when we scan all residues in the subsequent helices. Instead, we identify arginines located at the N-terminal part of the subsequent helices that are crucial for the recognition of the marginally hydrophobic transmembrane helices, indicating that the positive inside rule is important. However, in two of the constructs these arginines do not aid in the recognition without the rest of the subsequent helix, i.e. the positive inside rule alone is not sufficient. Instead, the improved recognition of marginally hydrophobic helices can here be explained as follows; the positive inside rule provides an orientational preference of the subsequent helix, which in turn allows the marginally hydrophobic helix to be inserted, i.e. the effect of the positive inside rule is stronger if positively charged residues are followed by a transmembrane helix. Such a mechanism can obviously not aid C-terminal helices and consequently we find that the terminal helices in multi-spanning membrane proteins are more hydrophobic than internal helices.

Molecular Biology

A unique HMG-box domain of mouse Maelstrom binds structured RNA but not double stranded DNA

Piwi-interacting piRNAs are a major and essential class of small RNAs in the animal germ cells with a prominent role in transposon control. Efficient piRNA biogenesis and function require a cohort of proteins conserved throughout the animal kingdom. Here we studied Maelstrom (MAEL), which is essential for piRNA biogenesis and germ cell differentiation in flies and mice. MAEL contains a high mobility group (HMG)-box domain and a Maelstrom-specific domain with a presumptive RNase H-fold. We employed a combination of sequence analyses, structural and biochemical approaches to evaluate and compare nucleic acid binding of mouse MAEL HMG-box to that of canonical HMG-box domain proteins (SRY and HMGB1a). MAEL HMG-box failed to bind double-stranded (ds)DNA but bound to structured RNA. We also identified important roles of a novel cluster of arginine residues in MAEL HMG-box in these interactions. Cumulatively, our results suggest that the MAEL HMG-box domain may contribute to MAEL function in selective processing of retrotransposon RNA into piRNAs. In this regard, a cellular role of MAEL HMG-box domain is reminiscent of that of HMGB1 as a sentinel of immunogenic nucleic acids in the innate immune response.

Molecular Biology

Proteomic analysis of Dhh1 complexes reveals a role for Hsp40 chaperone Ydj1 in yeast P-body assembly

P-bodies (PB) are ribonucleoprotein (RNP) complexes that aggregate into cytoplasmic foci when cells are exposed to stress. While the conserved mRNA decay and translational repression machineries are known components of PB, how and why cells assemble RNP complexes into large foci remain unclear. Using mass spectrometry to analyze proteins immunoisolated with the core PB protein Dhh1, we show that a considerable number of proteins contain low-complexity (LC) sequences, similar to proteins highly represented in mammalian RNP granules. We also show that the Hsp40 chaperone Ydj1, which contains an LC domain and controls prion protein aggregation, is required for the formation of Dhh1-GFP foci upon glucose depletion. New classes of proteins that reproducibly co-enrich with Dhh1-GFP during PB induction include proteins involved in nucleotide or amino acid metabolism, glycolysis, tRNA aminoacylation, and protein folding. Many of these proteins have been shown to form foci in response to other stresses. Finally, analysis of RNA associated with Dhh1-GFP shows enrichment of mRNA encoding the PB protein Pat1 and catalytic RNAs along with their associated mitochondrial RNA-binding proteins, suggesting an active role for RNA in PB function. Thus, global characterization of PB composition has uncovered proteins and RNA that are important for PB assembly.

Molecular Biology

Accurate Isothermal Amplification Reaction Rate Determination Using High-Frequency Sampling

BackgroundNucleic acids quantification by amplification is currently done primarily by real-time amplification for relative quantification, or by statistical inference from replicated endpoint assays for absolute quantification. The polymerase chain reaction (PCR) has been the dominant amplification technology, although alternative isothermal technologies have been described. Theoretical analysis of amplification kinetics and of amplification data interpretation have almost exclusively considered the PCR.\n\nResultsReal-time measurements of isothermal amplification reactions can be made continuously, in contrast to the discrete per-cycle measurements of real-time PCR. Isothermal ramified rolling circle amplification (RAM) reactions were measured at frequent intervals, and amplification data subsets were fitted to an exponential amplification model. Signal-change-over-time slopes and time-zero signal intercepts were derived from the chosen subset data. Slope measurements were sufficient to determine a reaction rate (the isothermal equivalent of PCR efficiency) for each reaction. Analysis of slope and intercept together suggest that amplification reactions that were initiated from a single target molecule can be distinguished from reactions that that were initiated from greater than one target molecule.\n\nConclusionsThe constant reaction environment of isothermal nucleic acid amplification allows continuous monitoring of reaction rate. Functional regions of interest in real-time data can be determined directly from the data. Accurate per-reaction efficiency can be readily measured. Improved estimation of low target copy number should improve quantification efficiency.

Molecular Biology

Thiosulfate-hydrogen peroxide redox oscillator as pH driver for ribozyme activity in the RNA world

The RNA world of more than 3.7 billion years ago may have drawn on thermal and pH oscillations set up by the oxidation of thiosulfate by hydrogen peroxide (the THP oscillator) as a power source to drive replication. Since this primordial RNA also must have developed enzyme functionalities, in this work we examine the responses of two simple ribozymes to a THP periodic drive, using experimental rate and thermochemical data in a dynamical model for the coupled, self-consistent evolution of all reactants and intermediates. The resulting time traces show that ribozyme performance can be enhanced under pH cycling, and that thermal cycling may have been necessary to achieve large performance gains. We discuss three important ways in which the dynamic hydrogen peroxide medium may have acted as an agent for development of the RNA world towards a cellular world: proton gradients, resolution of the ribozyme versus replication paradox, and vesicle formation.

Molecular Biology