Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

A genetic screen suggests an alternative mechanism for inhibition of SecA by azide

Sodium azide prevents bacterial growth by inhibiting the activity of SecA, which is required for translocation of proteins across the cytoplasmic membrane. Azide inhibits ATP turnover in vitro, but its mechanism of action in vivo is unclear. To investigate how azide inhibits SecA in cells, we used transposon directed insertion-site sequencing (TraDIS) to screen a library of transposon insertion mutants for mutations that affect the susceptibility of E. coli to azide. Insertions disrupting components of the Sec machinery generally increased susceptibility to azide, but insertions truncating the C-terminal tail (CTT) of SecA decreased susceptibility of E. coli to azide. Treatment of cells with azide caused increased aggregation of the CTT, suggesting that azide disrupts its structure. Analysis of the metal-ion content of the CTT indicated that SecA binds to iron and the azide disrupts the interaction of the CTT with iron. Azide also disrupted binding of SecA to membrane phospholipids, as did alanine substitutions in the metal-coordinating amino acids. Furthermore, treating purified phospholipid-bound SecA with azide in the absence of added nucleotide disrupted binding of SecA to phospholipids. Our results suggest that azide does not inhibit SecA by inhibiting the rate of ATP turnover in vivo. Rather, azide inhibits SecA by causing it to \"backtrack\" from the ADP-bound to the ATP-bound conformation, which disrupts the interaction of SecA with the cytoplasmic membrane.\n\nSignificance statementSecA is a bacterial ATPase that is required for the translocation of a subset of secreted proteins across the cytoplasmic membrane. Sodium azide is a well-known inhibitor of SecA, but its mechanism of action in vivo is poorly understood. To investigate this mechanism, we examined the effect of azide on the growth of a library of [~]1 million transposon insertion mutations. Our results suggest that azide causes SecA to backtrack in its ATPase cycle, which disrupts binding of SecA to the membrane and to its metal cofactor, which is iron. Our results provide insight into the molecular mechanism by which SecA drives protein translocation and how this essential biological process can be disrupted.

microbiology

Enter the matrix: Interpreting unsupervised feature learning with matrix decomposition to discover hidden knowledge in high-throughput omics data

Omics data contains signal from the molecular, physical, and kinetic inter- and intra-cellular interactions that control biological systems. Matrix factorization techniques can reveal low-dimensional structure from high-dimensional data that reflect these interactions. These techniques can uncover new biological knowledge from diverse high-throughput omics data in topics ranging from pathway discovery to time course analysis. We review exemplary applications of matrix factorization for systems-level analyses. We discuss appropriate application of these methods, their limitations, and focus on analysis of results to facilitate optimal biological interpretation. The inference of biologically relevant features with matrix factorization enables discovery from high-throughput data beyond the limits of current biological knowledge--answering questions from high-dimensional data that we have not yet thought to ask.

systems biology

Synergism between a simple sugar and a small intrinsically disordered protein mitigate the lethal stresses of severe water loss

Anhydrobiotes are rare microbes, plants and animals that tolerate severe water loss. Understanding the molecular basis for their desiccation tolerance may provide novel insights into stress biology and critical tools for engineering drought-tolerant crops. Using the anhydrobiote, budding yeast, we show that trehalose and Hsp12, a small intrinsically disordered protein (sIDP) of the hydrophilin family, synergize to mitigate completely the inviability caused by the lethal stresses of desiccation. We show that these two molecules help to stabilize the activity and prevent aggregation of model proteins both in vivo and in vitro. We also identify a novel role for Hsp12 as a membrane remodeler, a protective feature not shared by another yeast hydrophilin, suggesting that sIDPs have distinct biological functions.

cell biology

Parallelized Inference for Single Cell Transcriptomic Clustering with Split Merge Sampling on DPMM Model

Motivation: With the development of droplet based systems, massive single cell transcriptome data has become available, which enables analysis of cellular and molecular processes at single cell resolution and is instrumental to understanding many biological processes. While state-of-the-art clustering methods have been applied to the data, they face challenges in the following aspects: (1) the clustering quality still needs to be improved; (2) most models need prior knowledge on number of clusters, which is not always available; (3) there is a demand for faster computational speed.\n\nResults: We propose to tackle these challenges with Parallelized Split Merge Sampling on Dirichlet Process Mixture Model (the Para-DPMM model). Unlike classic DPMM methods that perform sampling on each single data point, the split merge mechanism samples on the cluster level, which significantly improves convergence and optimality of the result. The model is highly parallelized and can utilize the computing power of high performance computing (HPC) clusters, enabling massive inference on huge datasets. Experiment results show the model achieves about 7% improvement in clustering accuracy for small datasets and more than 20% improvement for large challenging datasets compared with current widely used models. In the mean time, the models computing speed is significantly faster.\n\nAvailability: Source code is publicly available on https://github.com/tiehangd/Para_DPMM/tree/master/Para_DPMM_package

bioinformatics

Label-free imaging of cholesterol and lipid distributions in model membranes

Over recent decades, lipid membranes have become standard models for examining the biophysics and biochemistry of cell membranes. Interrogation of lipid domains within biomembranes is generally done with fluorescence microscopy via exogenous chemical probes. However, most fluorophores have limited partitioning tunability, with the majority segregating in the least biologically relevant domains (i.e., low-density liquid domains). Therefore, a molecular-level picture of the majority of non-labeled lipids forming the membrane is still elusive. Here, we present simple, label-free imaging of domain formation in lipid monolayers, with chemical selectivity in unraveling lipid and cholesterol composition in all domain types. Exploiting conventional vibrational contrast in spontaneous Raman imaging, combined with chemometrics analysis, allows for examination of ternary systems containing saturated lipids, unsaturated lipids, and cholesterol. We confirm features commonly observed by fluorescence microscopy, and provide an unprecedented analysis of cholesterol distribution at the single-membrane level.\n\n\n\nO_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=197 SRC=\"FIGDIR/small/279794_ufig1.gif\" ALT=\"Figure 1\">\nView larger version (76K):\norg.highwire.dtl.DTLVardef@1395718org.highwire.dtl.DTLVardef@1851c93org.highwire.dtl.DTLVardef@1702f5org.highwire.dtl.DTLVardef@c33555_HPS_FORMAT_FIGEXP M_FIG C_FIG

biophysics

Switching toxic protein function in life cells

Toxic proteins are prime targets for molecular farming and efficient tools for targeted cell ablation in genetics, developmental biology, and biotechnology. Achieving conditional activity of cytotoxins and their maintenance in form of stably transformed transgenes is challenging. We demonstrate here a switchable version of the highly cytotoxic bacterial ribonuclease barnase by using efficient temperature-dependent control of protein accumulation in living multicellular organisms. By tuning the levels of the protein, we were able to control the fate of a plant organ in vivo. The on-demand-formation of specialized epidermal cells (trichomes) through manipulating stabilization versus destabilization of barnase is a proof-of-concept for a robust and powerful tool for conditional switchable cell arrest. We present this tool both as a potential novel strategy for the manufacture and accumulation of cytotoxic proteins and toxic high-value products in plants or for conditional genetic cell ablation.

plant biology

DNA microscopy: Optics-free spatio-genetic imaging by a stand-alone chemical reaction

Analyzing the spatial organization of molecules in cells and tissues is a cornerstone of biological research and clinical practice. However, despite enormous progress in profiling the molecular constituents of cells, spatially mapping these constituents remains a disjointed and machinery-intensive process, relying on either light microscopy or direct physical registration and capture. Here, we demonstrate DNA microscopy, a new imaging modality for scalable, optics-free mapping of relative biomolecule positions. In DNA microscopy of transcripts, transcript molecules are tagged in situ with randomized nucleotides, labeling each molecule uniquely. A second in situ reaction then amplifies the tagged molecules, concatenates the resulting copies, and adds new randomized nucleotides to uniquely label each concatenation event. An algorithm decodes molecular proximities from these concatenated sequences, and infers physical images of the original transcripts at cellular resolution. Because its imaging power derives entirely from diffusive molecular dynamics, DNA microscopy constitutes a chemically encoded microscopy system.

bioengineering

GNMCADS: Sampling For Protein Conformation Diversity With Gaussian Network Model Guided Condition Annealed Diffusion Sampler

Proteins are dynamic molecules existing in diverse conformational states underlying their biological functions. Although recent approaches have enabled diverse conformational sampling by emulating molecular dynamics simulations, perturbing evolutionary information, or steering internal mechanisms of structure prediction models, predicting conformations resulting from major domain motions or motions that occur over long timescales still remains a challenge. To this end, we introduce GNMCADS, a conformational sampling strategy that enhances the diversity of protein diffusion models by selectively annealing the conditioning signal guided by the intrinsic dynamical organization of the sampled protein. Further, we implement GNMCADS in the diffusion module of AlphaFold3, enabling the generation of diverse protein conformations. When benchmarked across 92 proteins that include 54 class A GPCRs, 15 transporters, and 23 proteins with major domain movements, GNMCADS exhibits improved sampling diversity compared to other current conformational sampling methods.

bioinformatics

Molecular recording of mammalian embryogenesis

Understanding the emergence of complex multicellular organisms from single totipotent cells, or ontogenesis, represents a foundational question in biology. The study of mammalian development is particularly challenging due to the difficulty of monitoring embryos in utero, the variability of progenitor field sizes, and the indeterminate relationship between the generation of uncommitted progenitors and their progression to subsequent stages. Here, we present a flexible, high information, multi-channel molecular recorder with a single cell (sc) readout and apply it as an evolving lineage tracer to define a mouse cell fate map from fertilization through gastrulation. By combining lineage information with scRNA-seq profiles, we recapitulate canonical developmental relationships between different tissue types and reveal an unexpected transcriptional convergence of endodermal cells from extra-embryonic and embryonic origins, illustrating how lineage information complements scRNA-seq to define cell types. Finally, we apply our cell fate map to estimate the number of embryonic progenitor cells and the degree of asymmetric partitioning within the pluripotent epiblast during specification. Our approach enables massively parallel, high-resolution recording of lineage and other information in mammalian systems to facilitate a quantitative framework for describing developmental processes.

developmental biology

Data-driven characterization of molecular phenotypes across heterogenous sample collections

The existing large gene expression data repositories hold enormous potential to elucidate disease mechanisms, characterize changes in cellular pathways, and to stratify patients based on their molecular profile. To achieve this goal, integrative resources and tools are needed that allow comparison of results across datasets and data types. We propose an intuitive approach for data-driven stratifications of molecular profiles and benchmark our methodology using the dimensional reduction algorithm t-SNE with multi-center and multi-platform data representing hematological malignancies. Our approach enables assessing the contribution of biological versus technical variation to sample clustering, direct incorporation of additional datasets to the same low dimensional representation of molecular disease subtypes, comparison of sample groups between separate t-SNE representations, or maps, and characterization of the obtained clusters based on pathway databases and additional multi-omics data. In the example application, our approach revealed differential activity of SAM-dependent DNA methylation pathway in the acute myeloid leukemia patient cluster characterized with CEBPA mutations that accordingly was validated to have globally elevated DNA methylation levels.

bioinformatics

A comparative analysis of network mutation burdens across 21 tumor types augments discovery from cancer genomes

Heterogeneity across cancer makes it difficult to find driver genes with intermediate (2-20%) and low frequency (<2%) mutations1, and we are potentially missing entire classes of networks (or pathways) of biological and therapeutic value. Here, we quantify the extent to which cancer genes across 21 tumor types have an increased burden of mutations in their immediate gene network derived from functional genomics data. We formalize a classifier that accurately calculates the significance level of a genes network mutation burden (NMB) and show it can accurately predict known cancer genes and recently proposed driver genes in the majority of tested tumours. Our approach predicts 62 putative cancer genes, including 35 with clear connection to cancer and 27 genes, which point to new cancer biology. NMB identifies proportionally more (4x) low-frequency mutated genes as putative cancer genes than gene-based tests, and provides molecular clues in patients without established driver mutations. Our quantitative and comparative analysis of pan-cancer networks across 21 tumour types gives new insights into the biological and genetic architecture of cancers and enables additional discovery from existing cancer genomes. The framework we present here should become increasingly useful with more sequencing data in the future.

Cancer Biology

Bayesian Node Dating based on Probabilities of Fossil Sampling Supports Trans-Atlantic Dispersal of Cichlid Fishes

Divergence-time estimation based on molecular phylogenies and the fossil record has provided insights into fundamental questions of evolutionary biology. In Bayesian node dating, phylogenies are commonly time calibrated through the specification of calibration densities on nodes representing clades with known fossil occurrences. Unfortunately, the optimal shape of these calibration densities is usually unknown and they are therefore often chosen arbitrarily, which directly impacts the reliability of the resulting age estimates. As possible solutions to this problem, two non-exclusive alternative approaches have recently been developed, the \"fossilized birth-death\" model and \"total-evidence dating\". While these approaches have been shown to perform well under certain conditions, they require including all (or a random subset) of the fossils of each clade in the analysis, rather than just relying on the oldest fossils of clades. In addition, both approaches assume that fossil records of different clades in the phylogeny are all the product of the same underlying fossil sampling rate, even though this rate has been shown to differ strongly between higher-level taxa. We here develop a flexible new approach to Bayesian node dating that combines advantages of traditional node dating and the fossilized birth-death model. In our new approach, calibration densities are defined on the basis of first fossil occurrences and sampling rate estimates that can be specified separately for all clades. We verify our approach with a large number of simulated datasets, and compare its performance to that of the fossilized birth-death model. We find that our approach produces reliable age estimates that are robust to model violation, on par with the fossilized birth-death model. By applying our approach to a large dataset including sequence data from over 1000 species of teleost fishes as well as 147 carefully selected fossil constraints, we recover a timeline of teleost diversification that is incompatible with previously assumed vicariant divergences of freshwater fishes. Our results instead provide strong evidence for trans-oceanic dispersal of cichlids and other groups of teleost fishes. (Keywords: node dating; calibration density; relaxed molecular clock; fossil record; Cichlidae; marine dispersal)

Evolutionary Biology

Modeling Zika Virus Congenital Eye Disease: Differential Susceptibility of Fetal Retinal Progenitor Cells and iPSC-Derived Retinal Stem Cells to Zika Virus Infection

Zika virus (ZIKV) causes microcephaly and congenital eye disease that is characterized by macular pigment mottling, macular atrophy, and loss of foveal reflex. The cell and molecular basis of congenital ZIKV infection are not well understood. Here, we utilized a biologically relevant cell-based system on human fetal retinal pigment epithelial cells (FRPE) and iPSC-derived retinal stem cells (iRSCs) to model ZIKV-ocular cell injury processes. FRPEs were highly susceptible to ZIKV, resulting in apoptosis and decreased viability, whereas iRSCs showed reduced susceptibility. Transcriptomics and proteomics analyses of infected FRPE cells revealed the activation of innate immune and inflammatory response genes, and dysregulation of cell survival pathways, mitochondrial transmembrane potential, phagocytosis, and particle internalization. Nucleoside analogue drug treatment inhibited ZIKV replication and prevented apoptosis. In conclusion, ZIKV affects ocular cells of different developmental stages resulting in cellular injury and death, further providing molecular insight into the pathogenesis of congenital eye disease.

microbiology

Inferring biochemical reactions and metabolite structures to cope with metabolic pathway drift

Inferring genome-scale metabolic networks in emerging model organisms is challenging because of incomplete biochemical knowledge and incomplete conservation of biochemical pathways during evolution. This limits the possibility to automatically transfer knowledge from well-established model organisms. Therefore, specific bioinformatic tools are necessary to infer new biochemical reactions and new metabolic structures that can be checked experimentally. Using an integrative approach combining both genomic and metabolomic data in the red algal model Chondrus crispus, we show that, even metabolic pathways considered as conserved, like sterol or mycosporine-like amino acids (MAA) synthesis pathways, undergo substantial turnover. This phenomenon, which we formally define as \"metabolic pathway drift\", is consistent with findings from other areas of evolutionary biology, indicating that a given phenotype can be conserved even if the underlying molecular mechanisms are changing. We present a proof of concept with a new methodological approach to formalize the logical reasoning necessary to infer new reactions and new molecular structures, based on previous biochemical knowledge. We use this approach to infer previously unknown reactions in the sterol and MAA pathways.\n\nAuthor summaryGenome-scale metabolic models describe our current understanding of all metabolic pathways occuring in a given organism. For emerging model species, where few biochemical data are available about really occurring enzymatic activities, such metabolic models are mainly based on transferring knowledge from other more studied species, based on the assumption that the same genes have the same function in the compared species. However, integration of metabolomic data into genome-scale metabolic models leads to situations where gaps in pathways cannot be filled by known enzymatic reactions from existing databases. This is due to structural variation in metabolic pathways accross evolutionary time. In such cases, it is necessary to use complementary approaches to infer new reactions and new metabolic intermediates using logical reasoning, based on available partial biochemical knowledge. Here we present a proof of concept that this is feasible and leads to hypotheses that are precise enough to be a starting point for new experimental work.

systems biology

Hysteresis in the thermal relaxation dynamics of an immune complex as basis for molecular memory

Proteins search their vast conformational space in order to attain the native fold and bind productively to relevant biological partners. In particular, most proteins must be able to alternate between at least one active conformational state and back to an inactive conformer, especially for the macromolecules that perform work and need to optimize energy usage. This property may be invoked by a physical stimulus (temperature, radiation) or by a chemical ligand, and may occur through mapping of the protein external environment onto a subset of protein conformers. We have stimulated with temperature cycles two partners of an immune complex before and after assembly, and revealed that properties of the external stimulus (period, phase) are also found in the characteristics of the immune complex (i.e. periodic variations in the binding affinity). These results are important for delineating the bases of molecular memory ex vivo and could serve in the optimization of protein based sensors.

biophysics

Autoamplification and competition drive symmetry breaking: Initiation of centriole duplication by the PLK4-STIL network

Symmetry breaking, a central principle of physics, has been hailed as the driver of self-organization in biological systems in general and biogenesis of cellular organelles in particular, but the molecular mechanisms of symmetry breaking only begin to become understood. Centrioles, the structural cores of centrosomes and cilia, must duplicate every cell cycle to ensure their faithful inheritance through cellular divisions. Work in model organisms identified conserved proteins required for centriole duplication and found that altering their abundance affects centriole number. However, the biophysical principles that ensure that, under physiological conditions, only a single procentriole is produced on each mother centriole remain enigmatic. Here we propose a mechanistic biophysical model for the initiation of procentriole formation in mammalian cells. We posit that interactions between the master regulatory kinase PLK4 and its activator-substrate STIL form the basis of the procentriole initiation network. The model faithfully recapitulates the experimentally observed transition from PLK4 uniformly distributed around the mother centriole, the \"ring\", to a unique PLK4 focus, the \"spot\", that triggers the assembly of a new procentriole. This symmetry breaking requires a dual positive feedback based on autocatalytic activation of PLK4 and enhanced centriolar anchoring of PLK4-STIL complexes by phosphorylated STIL. We find that, contrary to previous proposals, in situ degradation of active PLK4 is insufficient to break symmetry. Instead, the model predicts that competition between transient PLK4 activity maxima for PLK4-STIL complexes explains both the instability of the PLK4 ring and formation of the unique PLK4 spot. In the model, strong competition at physiologically normal parameters robustly produces a single procentriole, while increasing overexpression of PLK4 and STIL weakens the competition and causes progressive addition of procentrioles in agreement with experimental observations.

cell biology

Comparing cancer cell lines and tumor samples by genomic profiles

Cancer cell lines are often used in laboratory experiments as models of tumors, although they can have substantially different genetic and epigenetic profiles compared to tumors. We have developed a general computational method - TumorComparer - to systematically quantify similarities and differences between tumor material when detailed genetic and molecular profiles are available. The comparisons can be flexibly tailored to a particular biological question by placing a higher weight on functional alterations of interest ( weighted similarity). In a first pan-cancer application, we have compared 260 cell lines from the Cancer Cell Line Encyclopaedia (CCLE) and 1914 tumors of six different cancer types from The Cancer Genome Atlas (TCGA), using weights to emphasize genomic alterations that frequently recur in tumors. We report the potential suitability of particular cell lines as tumor models and identify apparently unsuitable outlier cell lines, some of which are in wide use, for each of the six cancer types. In future, this weighted similarity method may be generalized for use in a clinical setting to compare patient profiles consisting of genomic patterns combined with clinical attributes, such as diagnosis, treatment and response to therapy.

Cancer Biology

Synthetic analysis of natural variants yields insights into the evolution and function of auxin signaling F-box proteins in Arabidopsis thaliana

The evolution of complex body plans in land plants has been paralleled by gene duplication and divergence within nuclear auxin-signaling networks. A deep mechanistic understanding of auxin signaling proteins therefore may allow rational engineering of novel plant architectures. Towards that end, we analyzed natural variation in the auxin receptor F-box family of wild accessions of the reference plant Arabidopsis thaliana and used this information to populate a structure/function map. We employed a synthetic assay to identify natural hypermorphic F-box variants, and then assayed auxin-associated phenotypes in accessions expressing these variants. To more directly measure the impact of the strongest variant in our synthetic assay on auxin sensitivity, we generated transgenic plants expressing this allele. Together, our findings link evolved sequence variation to altered molecular performance and auxin sensitivity. This approach demonstrates the potential for combining synthetic biology approaches with quantitative phenotypes to harness the wealth of available sequence information and guide future engineering efforts of diverse signaling pathways.

plant biology