Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Trans-branching of polyubiquitin chains orchestrates the DNA replication stress response

Polyubiquitin chain geometry dictates functional consequences of ubiquitylation. Although branched polyubiquitin chains are abundant in cells, little is known about their functions. Here we show that branching on the DNA replication factor PCNA, mediated by the ubiquitin-conjugating enzyme UBE2K and involving lysines 63 and 48 of ubiquitin, orchestrates the sequence of events in response to replication stress. By inducing VCP-dependent extraction of PCNA from chromatin, branching promotes re-priming of stalled forks and necessitates a BRCA1-dependent pathway of daughter-strand gap repair. Our study identifies hyper-accumulation of daughter-strand gaps as the mechanistic basis underlying the toxicity of inhibitors of the PCNA-specific isopeptidase, USP1, in BRCA1-deficient cells. Moreover, an unexpected preference of UBE2K to operate in trans suggests a general timing mechanism to organize hierarchies amongst ubiquitin signals.

molecular biology

Impaired proteostasis is an early feature of the diabetic heart in humans and mice

Diabetes and obesity increase cardiac lipid levels leading to cardiomyopathy and heart failure. We hypothesized that intermittent fasting would reduce cardiac lipid levels. Surprisingly, intermittent fasting increased myocardial triglyceride content, but rescued mortality and attenuated cardiomyopathy in mice overexpressing cardiomyocyte acyl-CoA synthetase 1 (MHC-ACSL1). Lipid overload caused cardiomyocyte accumulation of polyubiquitinated protein aggregates containing desmin, a scaffolding intermediate filament protein, which intermittent fasting prevented. Furthermore, intermittent fasting reversed elevated myocardial C16:0 ceramide content, and knockdown of ceramide synthase CerS5 and CerS6 reduced palmitate-induced protein aggregation, highlighting a role for C16:0 ceramides in this pathology. Conversely, impairing aggrephagy with cardiomyocyte-specific p62 ablation induced heart failure in mice fed a high-fat diet, with paradoxically reduced cardiac lipid content. Crucially, non-failing diabetic human hearts also exhibited protein aggregate pathology. Taken together, these results demonstrate that impaired proteostasis characterizes cardiomyopathy from cardiac lipid overload and identify a promising new therapeutic target for this condition.

molecular biology

Structural mechanism of nuclear membrane sealing by LEM2-ESCRT-III

In open mitosis, re-establishing nucleocytoplasmic compartmentalization requires the LEM2-ESCRT machinery to coordinate spindle clearance with sealing of the remaining nuclear envelope pores. The structural basis of this topologically unique and fundamental membrane-remodeling process is poorly understood. Here, we combine biochemical reconstitution, cryo-electron tomography, subtomogram averaging and large-scale molecular dynamics simulations to define the structural mechanism of nuclear membrane sealing. We structurally resolve that LEM2s winged-helix domain (WH) co-polymerizes with the ESCRT-II/III protein CHMP7 to form a membrane-bound scaffold whose geometry is progressively remodeled by downstream ESCRT-III proteins as it transitions from the flat membrane surrounding the pore towards the negatively curved membrane neck. In parallel, LEM2 positions its intrinsically disordered low-complexity domain within the pore, where condensation around spindle microtubules mechanically couples the membrane-ESCRT-LEM2 scaffold to the spindle and narrows the remaining diffusion path, restoring compartmentalization before membrane closure is complete. Remarkably, the LEM2-WH domain alone forms tightly constricted membrane tubes, coating the negatively curved inner surface, revealing an intrinsic membrane-remodeling activity of the receptor itself. Together, our work establishes a structural framework for how receptor-ESCRT co-polymerization, low complexity domain-mediated sealing and receptor-driven membrane remodeling guide nuclear-envelope pores from spindle-containing openings to terminal constriction and fusion.

molecular biology

SARAF represses the mild hypothermia response through the regulation of JUN

The mild hypothermia response (MHR) is a conserved mammalian cytoprotective program activated upon exposure to mild hypothermia (32 degrees C) that contributes to the neuroprotective effects of therapeutic hypothermia following hypoxic injury. Although rapid changes in intracellular calcium occur upon cooling, the mechanisms linking calcium dynamics to the activation of core MHR factors such as SP1 and RBM3 remain incompletely defined. In this study, we used siRNA-mediated knockdown (KD) of candidate regulators in conjunction with novel mild hypothermia indicator (MHI) reporters to identify upstream modulators of MHR-associated transcription. We identify SARAF, a negative regulator of store-operated calcium entry (SOCE), as a repressor of both SP1 and RBM3 under normothermic conditions. SARAF depletion is associated with increased intracellular calcium release and enhanced SP1- and RBM3-linked transcriptional outputs. We identify JUN as an important downstream factor mediating SARAF depletion-dependent de-repression of the MHR and demonstrate that it undergoes activation rapidly upon cooling. Finally, SARAF depletion conferred significant cytoprotection against hypoxia-induced early apoptosis. Collectively, these findings establish SARAF as an upstream regulator of MHR-associated transcription and provide a functional link between cold-induced intracellular calcium dynamics and the induction of core MHR effectors.

molecular biology

A novel target associated with senescence and inflammatory signaling in human intervertebral disc degeneration

Background Intervertebral disc degeneration (IDD) is a leading cause of chronic low back pain and disability worldwide, affecting most individuals over 50 years of age. Despite its prevalence, no disease-modifying therapies exist, and current interventions are limited to reducing pain. Cellular senescence and the associated secretory phenotype (SASP) have been increasingly recognized as major drivers of disc matrix degradation and inflammation. However, the upstream molecular mechanisms that lead to IDD degeneration are still unknown. Connexin 43 (Cx43), a gap junction protein implicated in progression of age-related diseases, has emerged as a key regulator of cellular senescence and inflammatory signalling in musculoskeletal tissues. Methods Human primary cells were isolated from intervertebral disc samples obtained from patients classified into clinically meaningful groups: healthy controls, chronic/mechanical degeneration (DDD, ADJ, ASD), and acute/inflammatory event (herniated nucleus pulposus, HNP). Cx43 expression was assessed by qPCR and Western blotting. Cellular senescence was evaluated through SA-{beta}-gal staining and analysis of p53/p21 expression. SASP factors and EMT-related markers were measured by qPCR. Protein expression was quantified by immunoblotting across different age groups and degeneration grades. Results In this current study Cx43, was identified as the most abundant connexin isoform in human intervertebral discs, showing a progressive increase in expression with age and disc degeneration. Also, high Cx43 expression correlated with increased expression of the senescent markers p53 and p21 and increased SA-{beta}-gal activity. Besides, increased expression of EMT-related and differentiation markers has been correlated with high Cx43 levels in human IDD samples, consistent with fibrotic remodeling processes. Conclusions These findings identify aberrant upregulation of Cx43 signaling as a potential mechanistic link between intervertebral disc cellular senescence and extracellular matrix degradation, with the ensuing inflammatory response, representing a novel potential therapeutic target to modulate senescence-driven pathogenesis and modulate IDD progression.

molecular biology

SOX2 can associate with chromatin directly by binding to DNA or indirectly via association with other chromatin-bound proteins

It is widely assumed that SOX2 regulates gene expression and facilitates the opening of chromatin by binding directly at SOX motifs. To test this assumption, we created a SOX2 DNA binding mutant to determine whether other regions of SOX2 contribute to gene target specificity. When exogenously expressed in cells, this SOX2 mutant [SOX2(G76P)], like elevated unmodified SOX2, dramatically alters the transcriptome, but it does so by regulating vastly different gene sets and gene networks than SOX2. Consistent with their differential effects on the transcriptome, ChIP-seq analysis demonstrates that SOX2 and SOX2(G76P) associate primarily with different genomic loci, and motif analysis indicates that SOX2 binds primarily at SOX motifs, whereas SOX2(G76P) associates with chromatin at non-SOX motifs, including AP-1 motifs. Additionally, ATAC-seq analysis indicates that SOX2 substantially increases chromatin accessibility, but SOX2(G76P) does not. The findings presented lead to the conclusion that SOX2(G76P) associates with chromatin indirectly by a "piggyback" mechanism through its association with other chromatin-associated proteins, including AP-1 complexes. Remarkably, we also show that SOX2 and SOX2(G76P) each associate with a subset of the same gene loci that contain several different DNA motifs, including AP-1 motifs, but no high confidence SOX motifs. Overall, our findings provide new perspectives on SOX2 and lead to two important conclusions: 1) selection of gene targets by SOX2 is not solely determined by its DNA binding domain, and 2) SOX2 not only associates with chromatin directly by binding to SOX motifs but can also associate with a subset of gene loci indirectly through its association with other chromatin-associated proteins.

molecular biology

Multiscale spatial analysis implicates chromosomal metaloops in gene patterning across the Drosophila brain

Scores of chromosome-scale loops, or metaloops, arise in the Drosophila brain, but their spatial organization and relationship to neural gene expression patterns remain unclear. Here, we used multiplexed Optical Reconstruction of Chromatin Architecture (ORCA) to examine the multiscale spatial organization of metaloops in cross-sections of 100s of larval and adult Drosophila brains. We find metaloops form preferentially in the central regions of the brain, where they nucleate the formation of metadomains, characterized by the intermingling of distal topologically associating domains (TADs). At the sub-cellular scale, metaloops tend to arise towards the nuclear center, and multiple metaloops in the same cell have a preference to form hubs (3 or more contacts). Each brain nucleus generally harbors only a few loops or hubs. An in-depth analysis of the hub centered on DIP-epsilon, a synaptic wiring gene, identified a three-way metadomain that brings together the DIP-epsilon TAD; a distal TAD carrying a paralog of DIP-epsilon, DIP-zeta; and a putative regulatory TAD, across 3 Mb. This metadomain adopts distinct conformations depending on gene expression; cells expressing DIP-epsilon or DIP-zeta show preferential interactions between the TAD carrying the corresponding gene and the putative regulatory TAD. We posit that the neuron-specific formation of different subsets of metadomains might coordinate the expression of diverse combinations of synaptic wiring genes underlying complex brain architecture.

molecular biology

A simple cloning-free method to efficiently induce gene expression using CRISPR/Cas9

Gain-of-function studies often require the tedious cloning of transgene cDNA into vectors for overexpression beyond the physiological expression levels. The rapid development of CRISPR/Cas technology presents promising opportunities to address these issues. Here we report a simple, cloning-free method to induce gene expression at endogenous locus using CRISPR/Cas9 activators. Our strategy utilises synthesized sgRNA expression cassettes to direct a nuclease-null Cas9 complex fused with transcriptional activators (VP64, p65 and Rta) for site-specific induction of endogenous genes. This strategy allows rapid initiation of gain-of-function studies in the same day. Using this cloning-free approach, we tested two CRISPR activation systems, dSpCas9VPR and dSaCas9VPR, for induction of multiple genes in human and rat cells. Our results showed that both CRISPR activators allow efficient induction of six different neural development genes (CRX, RORB, RAX, OTX2, ASCL1 and NEUROD1) in human cells, whereas the rat cells exhibit a more variable and less efficient levels of gene induction, as observed in three different genes (Ascl1, Neurod1, Nrl). Altogether, this study provides a simple method to efficiently activate endogenous gene expression using CRISPR/Cas9 activators, which can be applies as a rapid workflow to initiate gain-of-function studies for a range of molecular and cell biology disciplines.

molecular biology

Bayesian Energy Landscape Tilting: Towards Concordant Models of Molecular Ensembles

Predicting biological structure has remained challenging for systems such as disordered proteins that take on myriad conformations. Hybrid simulation/experiment strategies have been undermined by difficulties in evaluating errors from computa- tional model inaccuracies and data uncertainties. Building on recent proposals from maximum entropy theory and nonequilibrium thermodynamics, we address these issues through a Bayesian Energy Landscape Tilting (BELT) scheme for computing Bayesian \"hyperensembles\" over conformational ensembles. BELT uses Markov chain Monte Carlo to directly sample maximum-entropy conformational ensembles consistent with a set of input experimental observables. To test this framework, we apply BELT to model trialanine, starting from disagreeing simulations with the force fields ff96, ff99, ff99sbnmr-ildn, CHARMM27, and OPLS-AA. BELT incorporation of limited chemical shift and 3J measurements gives convergent values of the peptides , {beta}, and PPII conformational populations in all cases. As a test of predictive power, all five BELT hyperensembles recover set-aside measurements not used in the fitting and report accu- rate errors, even when starting from highly inaccurate simulations. BELTs principled fxramework thus enables practical predictions for complex biomolecular systems from discordant simulations and sparse data.

Biophysics

A robust method for RNA extraction and purification from a single adult mouse tendon

BackgroundMechanistic understanding of tendon molecular and cellular biology is crucial towards furthering our abilities to design new therapies for tendon and ligament injuries and disease. Recent transcriptomic and epigenomic studies in the field have harnessed the power of mouse genetics to reveal new insights into tendon biology. However, many mouse studies pool tendon tissues or use amplification methods to perform RNA analysis, which can significantly increase the experimental costs and limit the ability to detect changes in expression of low copy transcripts.\n\nMethodsSingle Achilles tendons were harvested from uninjured, contralateral injured, and wild type mice between 3-5 months of age, and RNA was extracted. RNA Integrity Number (RIN) and concentration were determined, and RT-qPCR gene expression analysis was performed.\n\nResultsAfter testing several RNA extraction approaches on single adult mouse Achilles tendons, we developed a protocol that was successful at obtaining high RIN and sufficient concentrations suitable for RNA analysis. We found that the RNA quality was sensitive to the time between tendon harvest and homogenization, and the RNA quality and concentration was dependent on the duration of homogenization. Using this method, we demonstrate that analysis of Scx gene expression in single mouse tendons reduces the biological variation caused by pooling tendons from multiple mice. We also show successful use of this approach to analyze Sox9 and Col1a2 gene expression changes in injured compared with uninjured control tendons.\n\nDiscussionOur work presents a robust, cost-effective, and straightforward method to extract high quality RNA from a single adult mouse Achilles tendon at sufficient amounts for RNA-seq and RT-qPCR. We show this can reduce biological variation and decrease the overall costs associated with experiments. This approach can also be applied to other skeletal tissues as well as precious human samples.

molecular biology

Promoter boundaries for the luxCDABE and betIBA-proXWV operons in Vibrio harveyi defined by the method RAIL: Rapid Arbitrary PCR Insertion Libraries

Experimental studies of transcriptional regulation in bacteria require the ability to precisely measure changes in gene expression, often accomplished through the use of reporter genes. However, the boundaries of promoter sequences required for transcription are often unknown, thus complicating construction of reporters and genetic analysis of transcriptional regulation. Here, we analyze reporter libraries to define the promoter boundaries of the luxCDABE bioluminescence operon and the betIBA-proXWV osmotic stress operon in Vibrio harveyi. We describe a new method called RAIL (Rapid Arbitrary PCR Insertion Libraries) that combines the power of arbitrary PCR and isothermal DNA assembly to rapidly clone promoter fragments of various lengths upstream of reporter genes to generate large libraries. To demonstrate the versatility and efficiency of RAIL, we analyzed the promoters driving expression of the luxCDABE and betIBA-proXWV operons and created libraries of DNA fragments from these loci fused to fluorescent reporters. Using flow cytometry sorting and deep sequencing, we identified the DNA regions necessary and sufficient for maximum gene expression for each promoter. These analyses uncovered previously unknown regulatory sequences and validated known transcription factor binding sites. We applied this high-throughput method to gfp, mCherry, and lacZ reporters and multiple promoters in V. harveyi. We anticipate that the RAIL method will be easily applicable to other model systems for genetic, molecular, and cell biological applications.\n\nImportanceGene reporter constructs have long been essential tools for studying gene regulation in bacteria, particularly following the recent advent of fluorescent gene reporters. We developed a new method that enables efficient construction of promoter fusions to reporter genes to study gene regulation. We demonstrate the versatility of this technique in the model bacterium Vibrio harveyi by constructing promoter libraries for three bacterial promoters using three reporter genes. These libraries can be used to determine the DNA sequences required for gene expression, revealing regulatory elements in promoters. This method is applicable to various model systems and reporter genes for assaying gene expression.

molecular biology

Conservation of conformational dynamics across prokaryotic actins

The actin family of cytoskeletal proteins is essential to the physiology of virtually all archaea, bacteria, and eukaryotes. While X-ray crystallography and electron microscopy have revealed structural homologies among actin-family proteins, these techniques cannot probe molecular-scale conformational dynamics. Here, we use all-atom molecular dynamic simulations to reveal conserved dynamical behaviors in four prokaryotic actin homologs: MreB, FtsA, ParM, and crenactin. We demonstrate that the majority of the conformational dynamics of prokaryotic actins can be explained by treating the four subdomains as rigid bodies. MreB, ParM, and FtsA monomers exhibited nucleotide-dependent dihedral and opening angles, while crenactin monomer dynamics were nucleotide-independent. We further determine that the opening angle of ParM is sensitive to a specific interaction between subdomains. Steered molecular dynamics simulations of MreB, FtsA, and crenactin dimers revealed that changes in subunit dihedral angle lead to intersubunit bending or twist, suggesting a conserved mechanism for regulating filament structure. Taken together, our results provide molecular-scale insights into the nucleotide and polymerization dependencies of the structure of prokaryotic actins, suggesting mechanisms for how these structural features are linked to their diverse functions.\n\nSignificance StatementSimulations are a critical tool for uncovering the molecular mechanisms underlying biological form and function. Here, we use molecular-dynamics simulations to identify common and specific dynamical behaviors in four prokaryotic homologs of actin, a cytoskeletal protein that plays important roles in cellular structure and division in eukaryotes. Dihedral angles and opening angles in monomers of bacterial MreB, FtsA, and ParM were all sensitive to whether the subunit was bound to ATP or ADP, unlike in the archaeal homolog crenactin. In simulations of MreB, FtsA, and crenactin dimers, changes in subunit dihedral angle led to bending or twisting in filaments of these proteins, suggesting a mechanism for regulating the properties of large filaments. Taken together, our simulations set the stage for understanding and exploiting structure- function relationships of bacterial cytoskeletons.

biophysics

Calculating biological module enrichment or depletion and visualizing data on large-scale molecular maps with ACSNMineR and RNaviCell R packages

Biological pathways or modules represent sets of interactions or functional relationships occurring at the molecular level in living cells. A large body of knowledge on pathways is organized in public databases such as the KEGG, Reactome, or in more specialized repositories, such as the Atlas of Cancer Signaling Network (ACSN). All these open biological databases facilitate analyses, improving our understanding of cellular systems. We hereby describe the R package ACSNMineR for calculation of enrichment or depletion of lists of genes of interest in biological pathways. ACSNMineR integrates ACSN molecular pathways, but can use any molecular pathway encoded as a GMT file, for instance sets of genes available in the Molecular Signatures Database (MSigDB). We also present the R package RNaviCell, that can be used in conjunction with ACSNMineR to visualize different data types on web-based, interactive ACSN maps. We illustrate the functionalities of the two packages with biological data taken from large-scale cancer datasets.

Bioinformatics

Energetic costs, precision, and efficiency of a biological motor in cargo transport

Molecular motors play key roles in organizing the interior of cells. An efficient motor in cargo transport would travel with a high speed and a minimal error in transport time (or distance) while consuming minimal amount of energy. The travel distance and its variance of motor are, however, physically constrained by energy consumption, the principle of which has recently been formulated into the thermodynamic uncertainty relation. Here, we reinterpret the uncertainty measure ([Q]) defined in the thermodynamic uncertainty relation such that a motor efficient in cargo transport is characterized with a small [Q]. Analyses on the motility data from several types of molecular motors show that [Q] is a nonmonotic function of ATP concentration and load (f). For kinesin-1, [Q] is locally minimized at [ATP] {approx} 200 M and f {approx} 4 pN. Remarkably, for the mutant with a longer neck-linker this local minimum vanishes, and the energetic cost to achieve the same precision as the wild-type increases significantly, which underscores the importance of molecular structure in transport properties. For the biological motors studied here, their value of [Q] is semi-optimized under the cellular condition ([ATP] {approx} 1 mM, f = 0 - 1 pN). We find that among the motors, kinesin-1 at single molecule level is the most efficient in cargo transport.

biophysics

Quantitative study of the somitogenetic wavefront in zebrafish

A quantitative description of the molecular networks that sustain morphogenesis is one of the challenges of developmental biology. Specifically, a molecular understanding of the segmentation of the antero-posterior axis in vertebrates has yet to be achieved. This process known as somitogenesis is believed to result from the interactions between a genetic oscillator and a posterior-moving determination wavefront. Here we quantitatively study and perturb the network in zebrafish that sustains this wavefront and compare our observations to a model whereby the wavefront is due to a switch between stable states resulting from reciprocal negative feedbacks of Retinoic Acid (RA) on the activation of ERK and of ERK on RA synthesis. This model quantitatively accounts for the near linear shortening of the post-somitic mesoderm (PSM) in response to the observed exponential decrease during somitogenesis of the mRNA concentration of a morphogen (Fgf8). It also accounts for the observed dynamics of the PSM when the molecular components of the network are perturbed. The generality of our model and its robustness allows for its test in other model organisms.

developmental biology

CHEMOMETRIC APPROACHES FOR DEVELOPING INFRARED NANOSENSORS TO IMAGE ANTHRACYCLINES

Generation, identification, and validation of optical probes to image molecular targets in a biological milieu remains a challenge. Synthetic molecular recognition approaches leveraging the intrinsic near-infrared fluorescence of single-walled carbon nanotubes is a promising approach for chronic biochemical imaging in tissues. However, generation of nanosensors for selective imaging of molecular targets requires a heuristic approach. Here, we present a chemometric platform for rapidly screening libraries of candidate single-walled carbon nanotube nanosensors against biochemical analytes to quantify fluorescence response to small molecules including vitamins, neurotransmitters, and chemotherapeutics. We further show this approach can be leveraged to identify biochemical analytes that selectively modulate the intrinsic near-infrared fluorescence of candidate nanosensors. Chemometric analysis thus enables identification of nanosensor-analyte hits and also nanosensor fluorescence signaling modalities such as wavelength-shifts that are optimal for translation to biological imaging. Through this approach, we identify and characterize a nanosensor for the chemotherapeutic anthracycline doxorubicin, which provides an up to 17 nm fluorescence red-shift and exhibits an 8 {micro}M limit of detection, compatible with peak circulatory concentrations of doxorubicin common in therapeutic administration. We demonstrate selectivity of this nanosensor over dacarbazine, a chemotherapeutic commonly co-injected with DOX. Lastly, we demonstrate nanosensor tissue compatibility for imaging of doxorubicin in muscle tissue by incorporating nanosensors into the mouse hindlimb and measuring nanosensor response to exogenous DOX administration. Our results motivate chemometric approaches to nanosensor discovery for chronic imaging of drug partitioning into tissues and towards real-time monitoring of drug accumulation.

biochemistry

Comprehensive single cell transcriptional profiling of a multicellular organism by combinatorial indexing

Conventional methods for profiling the molecular content of biological samples fail to resolve heterogeneity that is present at the level of single cells. In the past few years, single cell RNA sequencing has emerged as a powerful strategy for overcoming this challenge. However, its adoption has been limited by a paucity of methods that are at once simple to implement and cost effective to scale massively. Here, we describe a combinatorial indexing strategy to profile the transcriptomes of large numbers of single cells or single nuclei without requiring the physical isolation of each cell (Single cell Combinatorial Indexing RNA-seq or sci-RNA-seq). We show that sci-RNA-seq can be used to efficiently profile the transcriptomes of tens-of-thousands of single cells per experiment, and demonstrate that we can stratify cell types from these data. Key advantages of sci-RNA-seq over contemporary alternatives such as droplet-based single cell RNA-seq include sublinear cost scaling, a reliance on widely available reagents and equipment, the ability to concurrently process many samples within a single workflow, compatibility with methanol fixation of cells, cell capture based on DNA content rather than cell size, and the flexibility to profile either cells or nuclei. As a demonstration of sci-RNA-seq, we profile the transcriptomes of 42,035 single cells from C. elegans at the L2 stage, effectively 50-fold \"shotgun cellular coverage\" of the somatic cell composition of this organism at this stage. We identify 27 distinct cell types, including rare cell types such as the two distal tip cells of the developing gonad, estimate consensus expression profiles and define cell-type specific and selective genes. Given that C. elegans is the only organism with a fully mapped cellular lineage, these data represent a rich resource for future methods aimed at defining cell types and states. They will advance our understanding of developmental biology, and constitute a major step towards a comprehensive, single-cell molecular atlas of a whole animal.

genomics

UMI-Reducer: Collapsing duplicate sequencing reads via Unique Molecular Identifiers

Short Structured AbstractO_ST_ABSSummaryC_ST_ABSEvery sequencing library contains duplicate reads. While many duplicates arise during polymerase chain reaction (PCR), some duplicates derive from multiple identical fragments of mRNA present in the original lysate (termed \"biological duplicates\"). Unique Molecular Identifiers (UMIs) are random oligonucleotide sequences that allow differentiation between technical and biological duplicates. Here we report the development of UMI-Reducer, a new computational tool for processing and differentiating PCR duplicates from biological duplicates. UMI-Reducer uses UMIs and the mapping position of the read to identify and collapse reads that are technical duplicates. Remaining true biological reads are further used for bias-free estimate of mRNA abundance in the original lysate. This strategy is of particular use for libraries made from low amounts of starting material, which typically require additional cycles of PCR and therefore are most prone to PCR duplicate bias.\n\nAvailability and ImplementationThe UMI-Reducer is an open source Python software and is freely available for non-commercial use (GPL-3.0) at https://sergheimangul.wordpress.com/umi-reducer/. Documentation and tutorials are available at https://github.com/smangul1/UMI-Reducer/wiki/.\n\nContactsmangul@ucla.edu, SVanDriesche@mednet.ucla.edu\n\nSupplementary informationFlowchart of Library Construction

bioinformatics