Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,567 records · Page 87Linked to original sources

ChromoTrace: Reconstruction of 3D Chromosome Configurations by Super-Resolution Microscopy

MotivationThe three-dimensional structure of chromatin plays a key role in genome function, including gene expression, DNA replication, chromosome segregation, and DNA repair. Furthermore the location of genomic loci within the nucleus, especially relative to each other and nuclear structures such as the nuclear envelope and nuclear bodies strongly correlates with aspects of function such as gene expression. Therefore, determining the 3D position of the 6 billion DNA base pairs in each of the 23 chromosomes inside the nucleus of a human cell is a central challenge of biology. Recent advances of super-resolution microscopy in principle enable the mapping of specific molecular features with nanometer precision inside cells. Combined with highly specific, sensitive and multiplexed fluorescence labeling of DNA sequences this opens up the possibility of mapping the 3D path of the genome sequence in situ.\n\nResultsHere we develop computational methodologies to reconstruct the sequence configuration of all human chromosomes in the nucleus from a super-resolution image of a set of fluorescent in situ probes hybridized to the genome in a cell. To test our approach we develop a method for the simulation of chromatin packing in an idealized human nucleus. Our reconstruction method, ChromoTrace, uses suffix trees to assign a known linear ordering of in situ probes on the genome to an unknown set of 3D in situ probe positions in the nucleus from super-resolved images using the known genomic probe spacing as a set of physical distance constraints between probes. We find that ChromoTrace can assign the 3D positions of the majority of loci with high accuracy and reasonable sensitivity to specific genome sequences. By simulating spatial resolution, label multiplexing and noise scenarios we assess algorithm performance under realistic experimental constraints. Our study shows that it is feasible to achieve chromosome-wide reconstruction of the 3D DNA path in chromatin based on super-resolution microscopy images.

bioinformatics

Signatures of the evolution of parthenogenesis and cryptobiosis in the genomes of panagrolaimid nematodes

Most animal species reproduce sexually, but parthenogenesis, asexual reproduction of various forms, has arisen repeatedly. Parthenogenetic lineages are usually short lived in evolution; though in some environments parthenogenesis may be advantageous, avoiding the cost of sex. Panagrolaimus nematodes have colonised environments ranging from arid deserts to arctic and antarctic biomes. Many are parthenogenetic, and most have cryptobiotic abilities, being able to survive repeated complete desiccation and freezing. It is not clear which genomic and molecular mechanisms led to the successful establishment of parthenogenesis and the evolution of cryptobiosis in animals in general. At the same time, model systems to study these traits in the laboratory are missing.\n\nWe compared the genomes and transcriptomes of parthenogenetic and sexual Panagrolaimus able to survive crybtobiosis, as well as a non-cryptobiotic Propanogrolaimus species, to identify systems that contribute to these striking abilities. The parthenogens are most probably tripoids originating from hybridisation (allopolyploids). We identified genomic singularities like expansion of gene families, and selection on genes that could be linked to the adaptation to cryptobiosis. All Panagrolaimus have acquired genes through horizontal transfer, some of which are likely to contribute to cryptobiosis. Many genes acting in C. elegans reproduction and development were absent in distant nematode species (including the Panagrolaimids), suggesting molecular pathways cannot directly be transferred from the model system.\n\nThe easily cultured Panagrolaimus nematodes offer a system to study developmental diversity in Nematoda, the molecular evolution of parthenogens, the effects of triploidy on genomes stability, and the origin and biology of cryptobiosis.

evolutionary biology

accuMUlate: A mutation caller designed for mutation accumulation experiments

MotivationMutation accumulation (MA) is the most widely used method for directly studying the effects of mutation. Modern sequencing technologies have led to an increased interest in MA experiments. By sequencing whole genomes from MA lines, researchers can directly study the rate and molecular spectra of spontaneous mutations and use these results to understand how mutation contributes to biological processes. At present there is no software designed specifically for identifying mutations from MA lines. Studies that combine MA with whole genome sequencing use custom bioinformatic pipelines that implement heuristic rules to identify putative mutations.\n\nResultsHere we describe O_SCPLOWACCUC_SCPLOWMUO_SCPLOWLATEC_SCPLOW, a program that is designed to detect mutations from MA experiments. O_SCPLOWACCUC_SCPLOWMUO_SCPLOWLATEC_SCPLOW implements a probabilistic model that reflects the design of a typical MA experiments while being flexible enough to accommodate properties unique to any particular experiment. For each putative mutation identified from this model O_SCPLOWACCUC_SCPLOWMUO_SCPLOWLATEC_SCPLOW calculates a set of summary statistics that can be used to filter sites that may be false positives. A companion tool, O_SCPLOWDENOMINATEC_SCPLOW, can be used to apply filtering rules based on these statistics to simulated mutations and thus identify the number of callable sites per sample.\n\nAvailabilitySource code and releases available from https://github.com/dwinter/accuMUlate.

bioinformatics

Genomes from uncultivated prokaryotes: a comparison of metagenome-assembled and single-amplified genomes

BackgroundProkaryotes dominate the biosphere and regulate biogeochemical processes essential to all life. Yet, our knowledge about their biology is for the most part limited to the minority that has been successfully cultured. Molecular techniques now allow for obtaining genome sequences of uncultivated prokaryotic taxa, facilitating in-depth analyses that may ultimately improve our understanding of these key organisms.\n\nResultsWe compared results from two culture-independent strategies for recovering bacterial genomes: single-amplified genomes and metagenome-assembled genomes. Single-amplified genomes were obtained from samples collected at an offshore station in the Baltic Sea Proper and compared to previously obtained metagenome-assembled genomes from a time series at the same station. Among 16 single-amplified genomes analyzed, seven were found to match metagenome-assembled genomes, affiliated with a diverse set of taxa. Notably, genome pairs between the two approaches were nearly identical (>98.7% identity) across overlapping regions (30-80% of each genome). Within matching pairs, the single-amplified genomes were consistently smaller and less complete, whereas the genetic functional profiles were maintained. For the metagenome-assembled genomes, only on average 3.6% of the bases were estimated to be missing from the genomes due to wrongly binned contigs; the metagenome assembly was found to cause incompleteness to a higher degree than the binning procedure.\n\nConclusionsThe strong agreement between the single-amplified and metagenome-assembled genomes emphasizes that both methods generate accurate genome information from uncultivated bacteria. Importantly, this implies that the research questions and the available resources are allowed to determine the selection of genomics approach for microbiome studies.

bioinformatics

Microbiota-dependent elevation of Alcohol Dehydrogenase in Drosophila is associated with changes in alcohol-induced hyperactivity and alcohol preference

The gut microbiota impacts diverse aspects of host biology including metabolism, immunity, and behavior, but the scope of those effects and their underlying molecular mechanisms are poorly understood. To address these gaps, we used Two-dimensional Difference Gel Electrophoresis (2D-DIGE) to identify proteomic differences in male and female Drosophila heads raised with a conventional microbiota and those raised in a sterile environment (axenic). We discovered 22 microbiota-dependent protein differences, and identified a specific elevation in Alcohol Dehydrogenase (ADH) in axenic male flies. Because ADH is a key enzyme in alcohol metabolism, we asked whether physiological and behavioral responses to alcohol were altered in axenic males. Here we show that alcohol induced hyperactivity, the first response to alcohol exposure, is significantly increased in axenic males, requires ADH activity, and is modified by genetic background. While ADH activity is required, we did not detect significant microbe-dependent differences in systemic ADH activity or ethanol level. Like other animals, Drosophila exhibit a preference for ethanol consumption, and here we show significant microbiota-dependent differences in ethanol preference specifically in males. This work demonstrates that male Drosophilas association with their microbiota affects their physiological and behavioral responses to ethanol.

microbiology

Molecular dynamics ensemble refinement of the heterogeneous native state of NCBD using chemical shifts and NOEs

Many proteins display complex dynamical properties that are often intimately linked to their biological functions. As the native state of a protein is best described as an ensemble of confor-mations, it is important to be able to generate models of native state ensembles with high accuracy. Due to limitations in sampling efficiency and force field accuracy it is, however, challenging to obtain accurate ensembles of protein conformations by the use of molecular simulations alone. Here we show that dynamic ensemble refinement, which combines an accurate atomistic force field with commonly available nuclear magnetic resonance (NMR) chemical shifts and NOEs, can provide a detailed and accurate description of the conformational ensemble of the native state of a highly dynamic protein. As both NOEs and chemical shifts are averaged on timescales up to milliseconds, the resulting ensembles reflect the structural heterogeneity that goes beyond that probed e.g. by NMR relaxation order parameters. We selected the small protein domain NCBD as object of our study since this protein, which has been characterized experimentally in substantial detail, displays a rich and complex dynamical behaviour. In particular, the protein has been described as having a molten-globule like structure, but with a relatively rigid core. Our approach allowed us to describe the conformational dynamics of NCBD in solution, and to probe the structural heterogeneity resulting from both short- and long-time-scale dynamics by the calculation of order parameters on different time scales. These results illustrate the usefulness of our approach since they show that NCBD is rather rigid on the nanosecond timescale, but interconverts within a broader ensemble on longer timescales, thus enabling the derivation of a coherent set of conclusions from various NMR experiments on this protein, which could otherwise appear in contradiction with each other.

biophysics

How the Central American Seaway and an ancient northern passage affected Flatfish diversification

While the natural history of flatfish has been debated for decades, the mode of diversification of this biologically and economically important group has never been elucidated. To address this question, we assembled the largest molecular data set to date, covering > 300 species (out of ca. 800 extant), from 13 of the 14 known families over nine genes, and employed relaxed molecular clocks to uncover their patterns of diversification. As the fossil record of flatfish is contentious, we used sister species distributed on both sides of the American continent to calibrate clock models based on the closure of the Central American Seaway (CAS), and on their current species range. We show that flatfish diversified in two bouts, as species that are today distributed around the Equator diverged during the closure of CAS, while those with a northern range diverged after this, hereby suggesting the existence of a post-CAS closure dispersal for these northern species, most likely along a trans-Arctic northern route, a hypothesis fully compatible with paleogeographic reconstructions.

evolutionary biology

Identification of Key Proteins Involved in Axon Guidance Related Disorders: A Systems Biology Approach

Axon guidance is a crucial process for growth of the central and peripheral nervous systems. In this study, 3 axon guidance related disorders, namely-Duane Retraction Syndrome (DRS), Horizontal Gaze Palsy with Progressive Scoliosis (HGPPS) and Congenital fibrosis of the extraocular muscles type 3 (CFEOM3) were studied using various Systems Biology tools to identify the genes and proteins involved with them to get a better idea about the underlying molecular mechanisms including the regulatory mechanisms. Based on the analyses carried out, 7 significant modules have been identified from the PPI network. Five pathways/processes have been found to be significantly associated with DRS, HGPPS and CFEOM3 associated genes. From the PPI network, 3 have been identified as hub proteins-DRD2, UBC and CUL3.

systems biology

Are Genetic Interactions Influencing Gene Expression Evidence for Biological Epistasis or Statistical Artifacts?

The importance of epistasis - or statistical interactions between genetic variants - to the development of complex disease in humans has long been controversial. Genome-wide association studies of statistical interactions influencing human traits have recently become computationally feasible and have identified many putative interactions. However, several factors that are difficult to address confound the statistical models used to detect interactions and make it unclear whether statistical interactions are evidence for true molecular epistasis. In this study, we investigate whether there is evidence for epistasis regulating gene expression after accounting for technical, statistical, and biological confounding factors that affect interaction studies. We identified 1,119 (FDR=5%) interactions within cis-regulatory regions that regulate gene expression in human lymphoblastoid cell lines, a tightly controlled, largely genetically determined phenotype. Approximately half of these interactions replicated in an independent dataset (363 of 803 tested). We then performed an exhaustive analysis of both known and novel confounders, including ceiling/floor effects, missing genotype combinations, haplotype effects, single variants tagged through linkage disequilibrium, and population stratification. Every replicated interaction could be explained by at least one of these confounders, and replication in independent datasets did not protect against this issue. Assuming the confounding factors provide a more parsimonious explanation for each interaction, we find it unlikely that cis-regulatory interactions contribute strongly to human gene expression. As this calls into question the relevance of interactions for other human phenotypes, the analytic framework used here will be useful for protecting future studies of epistasis against confounding.

Genomics

A transcript-wide association study in physical activity intervention implicates molecular pathways in chronic disease

BackgroundPhysical activity is associated with decreased risk for several chronic and acute conditions including obesity, diabetes, cardiovascular disease, mental health and aging. However, the biological mechanisms associated with this decreased risk are elusive. One way to ascertain biological changes influenced by physical activity is by monitoring changes in how genes are expressed. In this investigation, we conducted a transcriptome-wide association study of physical activity, meta-analyzing 20 independent studies to increase power for discovery of genes expressed before and after physical activity. Further, we hypothesize that genes identified in physical activity are expressed in obesity, inflammation, major depressive disorder and healthy aging.\n\nResultsOur analysis identified thirty (30) transcripts induced by physical activity (PA signature), at an FDR < 0.05. Twenty (20) of these transcripts, including COL4A3, CAMKD1, SLC4A5, EPS15L1, RBM33, and CACNG1, are up-regulated and ten (10) transcripts including CRY1, ZNF346, SDF4, ANXA1 and YWHAZ are down-regulated. We find that several of these physical activity transcripts are associated and biologically concordant in direction with body mass index, white blood cell count, and healthy aging.\n\nConclusionspowerful approach, we found thirty genes that were putatively influenced by physical activity, eight of which are inversely associated with body mass index, thirteen inversely associated with white blood cell count, and three associated and concordant with healthy aging. One gene was significant and concordant with major depressive disorder. These results highlight the potential molecular basis for the protective benefit of physical activity for a broad set of chronic conditions.

genomics

RNA-Seq and Protein Mass Spectrometry in Microdissected Kidney Tubules Reveal Signaling Processes that Initiate Lithium-Induced Diabetes Insipidus

ABSTRACT1Lithium salts, used for treatment of bipolar disorder, frequently induce nephrogenic diabetes insipidus (NDI), limiting therapeutic success. NDI is associated with loss of expression of the molecular water channel, aquaporin-2, in the renal collecting duct (CD). Here, we use the methods of systems biology in a well-established rat model of lithium-induced NDI to identify signaling pathways activated at the onset of polyuria. Using single-tubule RNA-Seq, full transcriptomes were determined in microdissected cortical CDs of rats 72 hrs after initiation of lithium chloride (LiCl) administration (vs. time-controls without LiCl). Transcriptome-wide changes in mRNA abundances were mapped to gene sets associated with curated canonical signaling pathways, showing evidence for activation of NF-{kappa}B signaling with induction of genes coding for multiple chemokines as well as most components of the Major Histocompatibility Complex (MHC) Class I antigen-presenting complex. Administration of antiinflammatory doses of dexamethasone to LiCl-treated rats countered the loss of aquaporin-2 protein. RNA-Seq also confirmed prior evidence of a shift from quiescence into the cell cycle with arrest. Time course studies demonstrated an early (12 hrs) increase in multiple immediate early genes including several transcription factors. Protein mass spectrometry in microdissected cortical CDs provided corroborative evidence but also identified decreased abundance of several anti-oxidant proteins. Integration of new data with prior data about lithium effects at a molecular level leads to a signaling model in which lithium increases ERK activation leading to induction of NF-{kappa}B signaling and an inflammatory-like response that represses Aqp2 gene transcription.

systems biology

Contrasting patterns of divergence at the regulatory and sequence level in European Daphnia galeata natural populations

Understanding the genetic basis of local adaptation has long been a focus of evolutionary biology. Recently there has been increased interest in deciphering the evolutionary role of Daphnias plasticity and the molecular mechanisms of local adaptation. Using transcriptome data, we assessed the differences in gene expression profiles and sequences in four European Daphnia galeata populations. In total, ~33% of 32,903 transcripts were differentially expressed between populations. Among 10,280 differentially expressed transcripts, 5,209 transcripts deviated from neutral expectations and their population-specific expression pattern is likely the result of local adaptation processes. Furthermore, a SNP analysis allowed inferring population structure and distribution of genetic variation. The population divergence at the sequence-level was comparatively higher than the gene expression level by several orders of magnitude and consistent with strong founder effects and lack of gene flow between populations. Using sequence information, the candidate transcripts were annotated using a comparative genomics approach. Thus, we identified candidate transcriptomic regions for local adaptation in a key species of aquatic ecosystems in the absence of any laboratory induced stressor.

evolutionary biology

Pancreatic adenocarcinoma human organoids share structural and genetic features with primary tumors

Patient-derived pancreatic ductal adenocarcinoma (PDAC) organoid systems show great promise for understanding the biological underpinnings of disease and advancing therapeutic precision medicine. Despite the increased use of organoids, the fidelity of molecular features, genetic heterogeneity, and drug response to the tumor of origin remain important unanswered questions limiting their utility. To address this gap in knowledge, we created primary tumor- and PDX-derived organoids, and 2D cultures for in-depth genomic and histopathological comparisons to the primary tumor. Histopathological features and PDAC representative protein markers showed strong concordance. DNA and RNA sequencing of single organoids revealed patient-specific genomic and transcriptomic consistency. Single-cell RNAseq demonstrated that organoids are primarily a clonal population. In drug response assays, organoids displayed patient-specific sensitivities. Additionally, we examined the in vivo PDX response to FOLFIRINOX and Gemcitabine/Abraxane treatments, which was recapitulated in vitro by organoids. The patient-specific molecular and histopathological fidelity of organoids indicate that they can be used to understand the etiology of the patients tumor and the differential response to therapies and suggests utility for predicting drug responses.

cancer biology

Forecasting autism gene discovery with machine learning and genome-scale data

BackgroundGenes are one of the most powerful windows into the biology of autism, and it has been estimated that perhaps a thousand or more genes may confer risk. However, less than 100 genes are currently viewed as having robust enough evidence to be considered true \"autism genes\". Massive genetic studies are underway to produce data to implicate additional genes, but this approach, although necessary, is costly and slow-moving.\n\nMethodsWe approach autism gene discovery as a machine learning problem, rather than a genetic association problem, and use genome-scale data as predictors for identifying further genes that have similar properties in the feature space compared to established autism risk genes. This approach, which we call forecASD, integrates spatiotemporal gene expression, heterogeneous network data, and previous gene-level predictors of autism association into an ensemble classifier that yields a single score that indexes each genes evidence for being involved in the etiology of autism.\n\nResultsWe demonstrate that forecASD has substantially increased sensitivity and specificity compared to previous gene-level predictors of autism association, including genetic measures such as TADA. On an independent test set, consisting of newly-released pilot data from the SPARK Genomics Consortium, we show that forecASD best predicts which genes will have an excess of likely gene disrupting (LGD) de novo mutations. We further use independent data from a recent post mortem study of case/control gene expression to show that forecASD is also a significant predictor of genes implicated in ASD through differential expression. Using forecASD results, we show which molecular pathways are currently under-represented in the autism literature and likely represent under-appreciated biological mechanisms of autism. Finally, forecASD correctly predicted 12 of 16 genes implicated at FDR=0.2 by the latest ASD gene discovery study, while also identifying the most likely false positives among the candidate genes.\n\nConclusionsThese results demonstrate that forecASD bridges the gap between genetic- and expression-based ASD gene discovery, and provides a data-driven replacement to much of the manual filtering and curation that is a critical step in ensuring the robustness of gene discovery studies.

bioinformatics

Reverse-Engineering Biological Networks From Large Data Sets

Much of contemporary systems biology owes its success to the abstraction of a network, the idea that diverse kinds of molecular, cellular, and organismal species and interactions can be modeled as relational nodes and edges in a graph of dependencies. Since the advent of high-throughput data acquisition technologies in fields such as genomics, metabolomics, and neuroscience, the automated inference and reconstruction of such interaction networks directly from large sets of activation data, commonly known as reverse-engineering, has become a routine procedure. Whereas early attempts at network reverse-engineering focused predominantly on producing maps of system architectures with minimal predictive modeling, reconstructions now play instrumental roles in answering questions about the statistics and dynamics of the underlying systems they represent. Many of these predictions have clinical relevance, suggesting novel paradigms for drug discovery and disease treatment. While other reviews focus predominantly on the details and effectiveness of individual network inference algorithms, here we examine the emerging field as a whole. We first summarize several key application areas in which inferred networks have made successful predictions. We then outline the two major classes of reverse-engineering methodologies, emphasizing that the type of prediction that one aims to make dictates the algorithms one should employ. We conclude by discussing whether recent breakthroughs justify the computational costs of large-scale reverse-engineering sufficiently to admit it as a mainstay in the quantitative analysis of living systems.

bioinformatics

Molecular Recognition of Dopamine with Dual Near Infrared Excitation-Emission Two-Photon Microscopy

A key limitation for achieving deep imaging in biological structures lies in photon absorption and scattering leading to attenuation of fluorescence. In particular, neurotransmitter imaging is challenging in the biologically-relevant context of the intact brain, for which photons must traverse the cranium, skin and bone. Thus, fluorescence imaging is limited to the surface cortical layers of the brain, only achievable with craniotomy. Herein, we describe optimal excitation and emission wavelengths for through-cranium imaging, and demonstrate that near-infrared emissive nanosensors can be photoexcited using a two-photon 1560 nm excitation source. Dopamine-sensitive nanosensors can undergo two-photon excitation, and provide chirality-dependent responses selective for dopamine with fluorescent turn-on responses varying between 20% and 350%. We further calculate the two-photon absorption cross-section and quantum yield of dopamine nanosensors, and confirm a two-photon power law relationship for the nanosensor excitation process. Finally, we show improved image quality of the nanosensors embedded 2 mm deep into a brain-mimetic tissue phantom, whereby one-photon excitation yields 42% scattering, in contrast to 4% scattering when the same object is imaged under two-photon excitation. Our approach overcomes traditional limitations in deep-tissue fluorescence microscopy, and can enable neurotransmitter imaging in the biologically-relevant milieu of the intact and living brain.

neuroscience

Natural selection defines the cellular complexity

Current biology is perplexed by the lack of a theoretical framework for understanding the organization principles of the molecular system within a cell. Here we first studied growth rate, one of the seemingly most complex cellular traits, using functional data of yeast single-gene deletion mutants. We observed nearly one thousand expression informative genes (EIGs) whose expression levels are linearly correlated to the trait within an unprecedentedly large functional space. A simple model considering six EIG-formed protein modules revealed a variety of novel mechanistic insights, and also explained [~]50% of the variance of cell growth rates measured by Bar-seq technique for over 400 yeast mutants (Pearsons R = 0.69), a performance comparable to the microarray-based (R = 0.77) or colony-size-based (R = 0.66) experimental approach. We then applied the same strategy to 501 morphological traits of the yeast and achieved successes in most fitness-coupled traits each with hundreds of trait-specific EIGs. Surprisingly, there is no any EIG found for most fitness-uncoupled traits, indicating that they are controlled by super-complex epistases that allow no simple expression-trait correlation. Thus, EIGs are recruited exclusively by natural selection, which builds a rather simple functional architecture for fitness-coupled traits, and the endless complexity of a cell lies primarily in its fitness-uncoupled features.

Evolutionary Biology

Flawed evidence for convergent evolution of the circadian CLOCK gene in mole-rats

Convergently evolved mole-rats (Mammalia, Rodentia) provide a fascinating model for studying convergent molecular evolution. Three genome sequences have recently been made available for the blind mole-rat (Nannospalax galili; Spalacidae; Muroidea)1, and the convergently evolved naked mole-rat (Heterocephalus glaber; Heterocephalidae; Ctenohystrica)2 and its close relative the Damaraland mole-rat (Fukomys damarensis; Bathyergidae; Ctenohystrica)3. In their genome paper1, Fang et al. evaluated convergent molecular evolution related to the subterranean life-style between the naked mole-rat and the blind mole-rat. One particularly striking result was the strong signal for amino acid convergence detected in the circadian rhythm CLOCK gene. Here I show that this unexpected result is erroneous because it is based on the use of the wrong sequence for the naked mole-rat, which has been mistakenly replaced by a sequence from a blind mole-rat. When the correct sequence is used, the evidence for convergent molecular evolution in this gene appears very limited.

Evolutionary Biology