Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Molecular Biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,585 records · Page 88Linked to original sources

Phylogenetic expression profiling reveals widespread coordinated evolution of gene expression

Phylogenetic profiling, which infers functional relationships between genes based on patterns of gene presence/absence across species, has proven to be highly effective. Here we introduce a complementary approach, phylogenetic expression profiling (PEP), which detects gene sets with correlated expression levels across a phylogeny. Applying PEP to RNA-seq data consisting of 657 samples from 309 diverse unicellular eukaryotes, we found several hundred gene sets evolving in a coordinated fashion. These allowed us to predict a role of the Golgi apparatus in Alzheimer's disease, as well as novel genes related to diabetes pathways. We also detected adaptive evolution of tRNA ligase levels to match genome-wide codon usage. In sum, we found that PEP is an effective method for inferring functional relationships - especially among core cellular components that are never lost, to which phylogenetic profiling cannot be applied - and that many subunits of the most conserved molecular machines are coexpressed across eukaryotes.

Evolutionary Biology

Consensus Phylogenetic trees of Fifteen Prokaryotic Aminoacyl-tRNA Synthetase Polypeptides based on Euclidean Geometry of All-Pairs Distances and Concatenation

BackgroundMost molecular phylogenetic trees depict the relative closeness or the extent of similarity among a set of taxa based on comparison of sequences of homologous genes or proteins. Since the tree topology for individual monogenic traits varies among the same set of organisms and does not overlap taxonomic hierarchy, hence there is a need to generate multidimensional phylogenetic trees.\n\nResultsPhylogenetic trees were constructed for 119 prokaryotes representing 2 phyla under Archaea and 11 phyla under Bacteria after comparing multiple sequence alignments for 15 different aminoacyl-tRNA synthetase polypeptides. The topology of Neighbor Joining (NJ) trees for individual tRNA synthetase polypeptides varied substantially. We use Euclidean geometry to estimate all-pairs distances in order to construct phylogenetic trees. Further, we used a novel \"Taxonomic fidelity\" algorithm to estimate clade by clade similarity between the phylogenetic tree and the taxonomic tree. We find that, as compared to trees for individual tRNA synthetase polypeptides and rDNA sequences, the topology of our Euclidean tree and that for aligned and concatenated sequences of 15 proteins are closer to the taxonomic trees and offer the best consensus. We have also aligned sequences after concatenation, and find that by changing the order of sequence joining prior to alignment, the tree topologies vary. In contrast, changing the types of polypeptides in the grouping for Euclidean trees does not affect the tree topologies.\n\nConclusionsWe show that a consensus phylogenetic tree of 15 polypeptides from 14 aminoacyl-tRNA synthetases for 119 prokaryotes using Euclidean geometry exhibits better taxonomic fidelity than trees for individual tRNA synthetase polypeptides as well as 16S rDNA. We have also examined Euclidean N-dimensional trees for 15 tRNA synthetase polypeptides which give the same topology as that constructed after amalgamating 3-dimensional Euclidean trees for groups of 3 polypeptides. Euclidean N-dimensional trees offer a reliable future to multi-genic molecular phylogenetics.

Evolutionary Biology

Differential Strengths Of Molecular Determinants Guide Environment Specific Mutational Fates

Under the influence of selection pressures imposed by natural environments, organisms maintain competitive fitness through underlying molecular evolution of individual genes across the genome. For molecular evolution, how multiple interdependent molecular constraints play a role in determination of fitness under different environmental conditions is largely unknown. Here, using Deep Mutational Scanning (DMS), we quantitated empirical fitness of [~]2000 single site mutants of Gentamicin-resistant gene (GmR). This enabled a systematic investigation of effects of different physical and chemical environments on the fitness landscape of the gene. Molecular constraints of the fitness landscapes seem to bear differential strengths in an environment dependent manner. Among them, conformity of the identified directionalities of the environmental selection pressures with known effects of the environments on protein folding proves that along with substrate binding, protein stability is the common strong constraint of the fitness landscape. Our study thus provides mechanistic insights into the molecular constraints that allow accessibility of mutational fates in environment dependent manner.\n\nAuthor SummaryEnvironmental conditions play a central role in both organismal adaptations and underlying molecular evolution. Understanding of environmental effects on evolution of genotype is still lacking a depth of mechanistic insights needed to assist much needed ability to forecast mutational fates. Here, we address this issue by culminating high throughput mutational scanning using deep sequencing. This approach allowed comprehensive mechanistic investigation of environmental effects on molecular evolution. We monitored effects of various physical and chemical environments onto single site mutants of model antibiotic resistant gene. Alongside, to get mechanistic understanding, we identified multiple molecular constraints which contribute to various degrees in determining the resulting survivabilities of mutants. Across all tested environments, we find that along with substrate binding, protein stability stands out as the common strong constraints. Remarkable direct dependence of the environmental fitness effects on the type of environmental alteration of protein folding further proves that protein stability is the major constraint of the gene. So, our findings reveal that under the influence of environmental conditions, mutational fates are channeled by various degrees of strengths of underlying molecular constraints.

evolutionary biology

Cadmium Disrupts Vestibular Function by Interfering with Otolith Formation

Cadmium (Cd2+) is a transition metal found ubiquitously in the earths crust and is extracted in the production of other metals such as copper, lead, and zinc1,2. Human exposure to Cd2+ occurs through food consumption, cigarette smoking, and the combustion of fossil fuels. Cd2+ has been shown to be nephrotoxic, neurotoxic, and osteotoxic, and is a known carcinogen. Animal studies and epidemiological studies have linked prenatal Cd2+ exposure to hyperactivity and balance disorders although the mechanisms remain unknown. In this study we show that zebrafish developmentally exposed to Cd2+ exhibit abnormal otolith development and show an increased tendency to swim in circles, observations that are consistent with an otolith-mediated vestibular defect, in addition to being hyperactive. We also demonstrate that the addition of calcium rescues otolith malformation and reduces circling behavior but has no ameliorating effect on hyperactivity, suggesting that hyperactivity and balance disorders in human populations exposed to Cd are manifestations of separate underlying molecular pathways.

developmental biology

Discovery of Alstrom syndrome gene as a regulator of centrosome duplication in asymmetrically dividing stem cells in Drosophila.

Stereotypical inheritance of the mother vs. daughter centrosomes has been reported in several stem cells that divide asymmetrically. We report the identification of a protein that exhibits asymmetric localization between mother and daughter centrosomes in asymmetrically dividing Drosophila male germline stem cells (GSCs). We show that Alms1a, a Drosophila homolog of the causative gene for the human ciliopathy Alstrom Syndrome, is a ubiquitous mother centriole protein with a unique additional localization to the daughter centriole only in the mother centrosome of GSCs. Depletion of alms1a results in rapid loss of centrosomes due to failure in daughter centriole duplication. We reveal that alms1a is specifically required for centriole duplication in asymmetrically dividing cells but not in symmetrically dividing differentiating cells in multiple stem cell lineages. The unique requirement of alms1a in asymmetric dividing cells may shed light onto the molecular mechanisms of Alstrom syndrome pathogenesis.

developmental biology

The unexpected provenance of components in eukaryotic nucleotide-excision-repair and kinetoplast DNA-dynamics from bacterial mobile elements

BackgroundProtein weaponry deployed in biological conflicts between selfish elements and their hosts are increasingly recognized as being re-purposed for diverse molecular adaptations in the evolution of several uniquely eukaryotic systems. The anti-restriction protein ArdC, transmitted along with the DNA during invasion, is one such factor deployed by plasmids and conjugative transposons against their bacterial hosts.\n\nResultsUsing sensitive computational methods we unify the N-terminal single-stranded DNA-binding domain of ArdC (ArdC-N) with the DNA-binding domains of the nucleotide excision repair (NER) XPC/Rad4 protein and Trypanosoma Tc-38 (p38) protein implicated in kinetoplast(k) DNA replication and dynamics. We show that the ArdC-N domain was independently acquired twice by eukaryotes from bacterial mobile elements. One gave rise to the beta-hairpin domains of XPC/Rad4 and the other to the Tc-38-like proteins in the stem kinetoplastid. Eukaryotic ArdC-N domains underwent tandem duplications to form an extensive DNA-binding interface. In XPC/Rad4, the ArdC-N domain combined with the inactive transglutaminase domain of a peptide-N-glycanase originally derived from an active archaeal version, often incorporated in systems countering invasive DNA. We also show that parallel acquisitions from conjugative elements and bacteriophages gave rise to the Topoisomerase IA, DNA polymerases IB-Ds, and DNA ligases involved in kDNA dynamics.\n\nConclusionsWe resolve two outstanding questions in eukaryote-biology: 1) origin of the unique DNA lesion-recognition component of NER; 2) origin of the unusual, plasmid-like features of kDNA. These represent a more general trend in the origin of distinctive components of systems involved in DNA dynamics and their links to the ubiquitin system.

genomics

A systems biology approach uncovers the core gene regulatory network governing iridophore fate choice from the neural crest.

Multipotent neural crest (NC) progenitors generate an astonishing array of derivatives, including neuronal, skeletal components and pigment cells (chromatophores), but the molecular mechanisms allowing balanced selection of each fate remain unknown. In zebrafish, melanocytes, iridophores and xanthophores, the three chromatophore lineages, are thought to share progenitors and so lend themselves to investigating the complex gene regulatory networks (GRNs) underlying fate segregation of NC progenitors. Although the core GRN governing melanocyte specification has been previously established, those guiding iridophore and xanthophore development remain elusive. Here we focus on the iridophore GRN, where mutant phenotypes identify the transcription factors Sox10, Tfec and Mitfa and the receptor tyrosine kinase, Ltk, as key players. We present expression data, as well as loss and gain of function results, guiding the derivation of an initial iridophore specification GRN. Moreover, we use an iterative process of mathematical modelling, supplemented with a novel, Monte Carlo screening algorithm suited to the qualitative nature of the experimental data, to allow for rigorous predictive exploration of the GRN dynamics. Predictions were experimentally evaluated and testable hypotheses were derived to construct an improved version of the GRN, which we showed produced outputs consistent with experimentally observed gene expression dynamics. Our study reveals multiple important regulatory features, notably a sox10-dependent positive feedback loop between tfec and ltk driving iridophore specification; the molecular basis of sox10 maintenance throughout iridophore development; and the cooperation between sox10 and tfec in driving expression of pnp4a, a key differentiation gene. We also assess a candidate repressor of mitfa, a melanocyte-specific target of sox10. Surprisingly, our data challenge the reported role of Foxd3, an established mitfa repressor, in iridophore regulation. Our study builds upon our previous systems biology approach, by incorporating physiologically-relevant parameter values and rigorous evaluation of parameter values within a qualitative data framework, to establish for the first time the core GRN guiding specification of the iridophore lineage.\n\nAuthor SummaryMultipotent neural crest (NC) progenitors generate an astonishing array of derivatives, including neuronal, skeletal components and pigment cells, but the molecular mechanisms allowing balanced selection of each fate remain unknown. In zebrafish, melanocytes, iridophores and xanthophores, the three chromatophore lineages, are thought to share progenitors and so lend themselves to investigating the complex gene regulatory networks (GRNs) underlying fate segregation of NC progenitors. Although the core GRN governing melanocyte specification has been previously established, those guiding iridophore and xanthophore development remain elusive. Here we present expression data, as well as loss and gain of function results, guiding the derivation of a core iridophore specification GRN. Moreover, we use a process of mathematical modelling and rigorous computational exploration of the GRN to predict gene expression dynamics, assessing them by criteria suited to the qualitative nature of our current understanding of iridophore development. Predictions were experimentally evaluated and testable hypotheses were derived to construct an improved version of the GRN, which we showed produced outputs consistent with experimentally observed gene expression dynamics. The core iridophore GRN defined here is a key stepping stone towards exploring how chromatophores fate decisions are made in multipotent NC progenitors.

developmental biology

Pathway-Structured Predictive Model for Cancer Survival Prediction: A Two-Stage Approach

Heterogeneity in terms of tumor characteristics, prognosis, and survival among cancer patients has been a persistent problem for many decades. Currently, prognosis and outcome predictions are made based on clinical factors and/or by incorporating molecular profiling data. However, inaccurate prognosis and prediction may result by using only clinical or molecular information directly. One of the main shortcomings of past studies is the failure to incorporate prior biological information into the predictive model, given strong evidence of pathway-based genetic nature of cancer, i.e. the potential for oncogenes to be grouped into pathways based on biological functions such as cell survival, proliferation and metastatic dissemination.\n\nTo address this problem, we propose a two-stage procedure to incorporate pathway information into the prognostic modeling using large-scale gene expression data. In the first stage, we fit all predictors within each pathway using penalized Cox model (Lasso, Ridge and Elastic Net) and Bayesian hierarchical Cox model. In the second stage, we combine the cross-validated prognostic scores of all pathways obtained in the first stage as new predictors to build an integrated prognostic model for prediction. We apply the proposed method to analyze breast cancer data from The Cancer Genome Atlas (TCGA), predicting overall survival using clinical data and gene expression profiling. The data includes ~20000 genes mapped into 109 pathways for 505 patients. The results show that the proposed approach not only improves survival prediction compared with the alternative analysis that ignores the pathway information, but also identifies significant biological pathways.

Bioinformatics

Paradoxical signaling regulates structural plasticity in dendritic spines

Transient spine enlargement (3-5 min timescale) is an important event associated with the structural plasticity of dendritic spines. Many of the molecular mechanisms associated with transient spine en{-}largement have been identified experimentally. Here, we use a systems biology approach to construct a mathematical model of biochemical signaling and actin-mediated transient spine expansion in response to calcium-influx due to NMDA receptor activation. We have identified that a key feature of this signaling network is the paradoxical signaling loop. Paradoxical components act bifunctionally in signaling net{-}works and their role is to control both the activation and inhibition of a desired response function (protein activity or spine volume). Using ordinary differential equation (ODE)-based modeling, we show that the dynamics of different regulators of transient spine expansion including CaMKII, RhoA, and Cdc42 and the spine volume can be described using paradoxical signaling loops. Our model is able to capture the experimentally observed dynamics of transient spine volume. Furthermore, we show that actin remod{-}eling events provide a robustness to spine volume dynamics. We also generate experimentally testable predictions about the role of different components and parameters of the network on spine dynamics.

Neuroscience

Acute induction of anomalous blood clotting by molecular amplification of highly substoichiometric levels of bacterial lipopolysaccharide (LPS)

It is well known that a variety of inflammatory diseases are accompanied by hypercoagulability, and a number of more-or-less longer-term signalling pathways have been shown to be involved. In recent work, we have suggested a direct and primary role for bacterial lipopolysaccharide in this hypercoagulability, but it seems never to have been tested directly. Here we show that the addition of tiny concentrations (0.2 ng.L-1) of bacterial lipopolysaccharide (LPS) to both whole blood and platelet-poor plasma of normal, healthy donors leads to marked changes in the nature of the fibrin fibres so formed, as observed by ultrastructural and fluorescence microscopy (the latter implying that the fibrin is actually in an amyloid {beta}-sheet-rich form. They resemble those seen in a number of inflammatory (and also amyloid) diseases, consistent with an involvement of LPS in their aetiology. These changes are mirrored by changes in their viscoelastic properties as measured by thromboelastography. Since the terminal stages of coagulation involve the polymerisation of fibrinogen into fibrin fibres, we tested whether LPS would bind to fibrinogen directly. We demonstrated this using isothermal calorimetry. Finally, we show that these changes in fibre structure are mirrored when the experiment is done simply with purified fibrinogen and thrombin ({+/-} 0.2 ng.L-1 LPS). This ratio of concentrations of LPS:fibrinogen in vivo represents a molecular amplification by the LPS of more than 108-fold, a number that is probably unparalleled in biology. The observation of a direct effect of such highly substoichiometric amounts of LPS on both fibrinogen and coagulation can account for the role of very small numbers of dormant bacteria in disease progression, and opens up this process to further mechanistic analysis and possible treatment.\n\nSignificance statementMost chronic diseases (including those classified as cardiovascular, neurodegenerative, or autoimmune) are accompanied by long-term inflammation. Although typically mediated by inflammatory cytokines, the origin of this inflammation is unclear. We have suggested that one explanation is a dormant microbiome that can shed the highly inflammatory lipopolysaccharide LPS. Such inflammatory diseases are also accompanied by a hypercoagulable phenotype. We here show directly (using 6 different methods) that very low concentrations of LPS can affect the terminal stages of the coagulation properties of blood and plasma significantly, and that this may be mediated via a direct binding of LPS to a small fraction of fibrinogen monomers as assessed biophysically. Such amplification methods may be of more general significance.

Microbiology

Harnessing Big Data for Systems Pharmacology

Systems pharmacology aims to holistically understand genetic, molecular, cellular, organismal, and environmental mechanisms of drug actions through developing mechanistic or predictive models. Data-driven modeling plays a central role in systems pharmacology, and has already enabled biologists to generate novel hypotheses. However, more is needed. The drug response is associated with genetic/epigenetic variants and environmental factors, is coupled with molecular conformational dynamics, is affected by possible off-targets, is modulated by the complex interplay of biological networks, and is dependent on pharmacokinetics. Thus, in order to gain a comprehensive understanding of drug actions, systems pharmacology requires integration of models across data modalities, methodologies, organismal hierarchies, and species. This imposes a great challenge on model management, integration, and translation. Here, we discuss several upcoming issues in systems pharmacology and potential solutions to them using big data technology. It will allow systems pharmacology modeling to be findable, accessible, interoperable, reusable, reliable, interpretable, and actionable.

Pharmacology and Toxicology

Continuous rearrangement of the postsynaptic gephyrin scaffolding domain: a super-resolution quantified and energetic approach

Synaptic function and plasticity requires a delicate balance between overall structural stability and the continuous rearrangement of the components that make up the presynaptic active zone and the postsynaptic density (PSD). Photoactivated localization microscopy (PALM) has provided a detailed view of the nanoscopic structure and organization of some of these synaptic elements. Still lacking, are tools to address the morphing and stability of such complexes at super-resolution. We describe an approach to quantify morphological changes and energetic states of multimolecular assemblies over time. With this method, we studied the scaffold protein gephyrin, which forms postsynaptic clusters that play a key role in the stabilization of receptors at inhibitory synapses. Postsynaptic gephyrin clusters exhibit an internal microstructure composed of nanodomains. We found, that within the PSD, gephyrin molecules continuously undergo spatial reorganization. This dynamic behavior depends on neuronal activity and cytoskeleton integrity. The proposed approach also allowed access to the effective energy responsible for the tenacity of the PSD despite molecular instability.\n\nSignificant statementSuper-resolution microscopy has become an important tool for the study of biological systems, allowing detailed, nano-scale structural reconstruction, single molecule tracking, particle counting, and interaction studies. However, quantification tools that take full advantage of the information provided by this technology are still lacking. We describe a novel quantification method to obtain information related to the size, directionality, dynamics, and stability of clustered structures from super-resolution microscopy. With this method, we studied the stability of gephyrin clusters, the main inhibitory scaffold protein. We found that gephyrin molecules continuously undergo reorganization based on neuronal activity and changes in the cytoskeleton.

biophysics

De novo transcriptomic characterization of Betta splendens for identifying sex-biased genes potentially involved in aggressive behavior modulation and EST-SSR maker development

Betta splendens is not only a commercially important labyrinth fish but also a nice research model for understanding the biological underpinnings of aggressive behavior. However, the shortage of basic genetic resource severely inhibits investigations on the molecular mechanism in sexual dimorphism of aggressive behavior typicality, which are essential for further behavior-related studies. There is a lack of knowledge regarding the functional genes involved in aggression expression. The scarce marker resource also impedes research progress of population genetics and genomics. In order to enrich genetic data and sequence resources, transcriptomic analysis was conducted for mature B. splendens using a multiple-tissues mixing strategy. A total of 105,505,486 clean reads were obtained and by de novo assembly, 69,836 unigenes were generated. Of which, 35,751 unigenes were annotated in at least one of queried databases. The differential expression analysis resulted in 17,683 transcripts differentially expressed between males and females. Plentiful sex-biased genes involved in aggression exhibition were identified via a screening from Gene Ontology terms and Kyoto Encyclopedia of Genes and Genomes pathways, such as htr, drd, gabr, cyp11a1, cyp17a1, hsd17b3, dax1, sf-1, hsd17b7, gsdf1 and fem1c. These putative genes would make good starting points for profound mechanical exploration on aggressive behavioral regulation. Moreover, 12,751 simple sequence repeats were detected from 9,617 unigenes for marker development. Nineteen of the 100 randomly selected primer pairs were demonstrated to be polymorphic. The large amount of transcript sequences will considerably increase available genomic information for gene mining and function analysis, and contribute valuable microsatellite marker resources to in-depth studies on molecular genetics and genomics in the future.

zoology

The rate and molecular spectrum of spontaneous mutations in the GC-rich multi-chromosome genome of Burkholderia cenocepacia

Spontaneous mutations are ultimately essential for evolutionary change and are also the root cause of many diseases. However, until recently, both biological and technical barriers have prevented detailed analyses of mutation profiles, constraining our understanding of the mutation process to a few model organisms and leaving major gaps in our understanding of the role of genome content and structure on mutation. Here, we present a genome-wide view of the molecular mutation spectrum in Burkholderia cenocepacia, a clinically relevant pathogen with high %GC-content and multiple chromosomes. We find that B. cenocepacia has low genome-wide mutation rates with insertion-deletion mutations biased towards deletions, consistent with the idea that deletion pressure reduces prokaryotic genome sizes. Unlike prior studies of other organisms, mutations in B. cenocepacia are not AT-biased, which suggests that at least some genomes with high %GC-content experience unusual base-substitution mutation pressure. Importantly, we also observe variation in both the rates and spectra of mutations among chromosomes and elevated G:C>T:A transversions in late-replicating regions. Thus, although some patterns of mutation appear to be highly conserved across cellular life, others vary between species and even between chromosomes of the same species, potentially influencing the evolution of nucleotide composition and genome architecture.

Evolutionary Biology

Strong gene activation with genome-wide specificity using a new orthogonal CRISPR/Cas9-based Programmable Transcriptional Activator.

Synthetic Biology (SynBio) aims at rewiring plant metabolic and developmental programs with orthogonal regulatory circuits. This endeavour requires new molecular tools able to interact with endogenous factors in a potent yet at the same time highly specific manner. A promising new class of SynBio tools that could play this function are the synthetic transcriptional activators based on CRISPR/Cas9 architecture, which combine autonomous activation domains (ADs) capable of recruiting the cells transcription machinery, with the easily customizable DNA-binding activity of nuclease-inactivated Cas9 protein (dCas9), creating so-called Programmable Transcriptional Activators (PTAs). In search for optimized dCas9-PTAs we performed a combinatorial analysis with seven different ADs arranged in four different protein/RNA architectures. This analysis resulted in the selection of a new dCas9-PTA with improved features as compared with previously reported activators. The new synthetic riboprotein, named dCasEV2.1, combines EDLL and VPR ADs using a multiplexable mutated version (v2.1) of the previously described aptamer-containing guide RNA2.0. We show here that dCasEV2.1 is a strong and wide spectrum activator, displaying variable activation levels depending on the basal activity of the target promoter. Maximum activation rates reaching up to 10000 fold were observed when targeting the NbDFR gene. Most remarkably, RNAseq analysis of dCasEV2.1-transformed N. benthamiana leaves revealed that the topmost activation capacity of dCasEV2.1 on target genes is accompanied with strict genome-wide specificity, making dCasEV2.1 an attractive tool for rewiring plant metabolism and regulatory networks.

synthetic biology

Reference trait analysis reveals correlations between gene expression and quantitative traits in disjoint samples

Systems genetic analysis of complex traits involves the integrated analysis of genetic, genomic, and disease related measures. However, these data are often collected separately across multiple study populations, rendering direct correlation of molecular features to complex traits impossible. Recent transcriptome-wide association studies (TWAS) have harnessed gene expression quantitative trait loci (eQTL) to associate unmeasured gene expression with a complex trait in genotyped individuals, but this approach relies primarily on strong eQTLs. We propose a simple and powerful alternative strategy for correlating independently obtained sets of complex traits and molecular features. In contrast to TWAS, our approach gains precision by correlating complex traits through a common set of continuous phenotypes instead of genetic predictors, and can identify transcript-trait correlations for which the regulation is not genetic. In our approach, a set of multiple quantitative \"reference\" traits is measured across all individuals, while measures of the complex trait of interest and transcriptional profiles are obtained in disjoint sub-samples. A conventional multivariate statistical method, canonical correlation analysis, is used to relate the reference traits and traits of interest in order to identify gene expression correlates. We evaluate power and sample size requirements of this methodology, as well as performance relative to other methods, via extensive simulation and analysis of a behavioral genetics experiment in 258 Diversity Outbred mice involving two independent sets of anxiety-related behaviors and hippocampal gene expression. After splitting the dataset and hiding one set of anxiety-related traits in half the samples, we identified transcripts correlated with the hidden traits using the other set of anxiety-related traits and exploiting the highest canonical correlation (R = 0.69) between the trait datasets. We demonstrate that this approach outperforms TWAS in identifying associated transcripts. Together, these results demonstrate the validity, reliability, and power of the reference trait method for identifying relations between complex traits and their molecular substrates.\n\nAUTHOR SUMMARYSystems genetics exploits natural genetic variation and high-throughput measurements of molecular intermediates to dissect genetic contributions to complex traits. An important goal of this strategy is to correlate molecular features, such as transcript or protein abundance, with complex traits. For practical, technical, or financial reasons, it may be impossible to measure complex traits and molecular intermediates on the same individuals. Instead, in some cases these two sets of traits may be measured on independent cohorts. We outline a method, reference trait analysis, for identifying molecular correlates of complex traits in this scenario. We show that our method powerfully identifies complex trait correlates across a wide range of parameters that are biologically plausible and experimentally practical. Furthermore, we show that reference trait analysis can identify transcripts correlated to a complex trait more accurately than approaches such as TWAS that use genetic variation to predict gene expression. Reference trait analysis will contribute to furthering our understanding of variation in complex traits by identifying molecular correlates of complex traits that are measured in different individuals.

genomics

Protein-Protein Interaction Network Analysis and Identification of Key Players in N-hydroxy-nor-L-Arg (nor-NOHA) and N(omega)-hydroxy-L-arginine (NOHA) Mediated Pathways for Treatment of Cancer Through Arginase Inhibiton: Insights from Systems Biology

L-arginine is involved in a number of biological processes in our bodies. Metabolism of L-arginine by the enzyme arginase has been found to be associated with cancer cell proliferation. Arginase inhibition has been proposed as a potential therapeutic means to inhibit this process. N-hydroxy-nor-L-Arg (nor-NOHA) and N (omega)-hydroxy-L-arginine (NOHA) has shown promise in inhibiting cancer progression through arginase inhibition. In this study, nor-NOHA and NOHA-associated genes and proteins were analyzed with several Bioinformatics and Systems Biology tools to identify the associated pathways and the key players involved so that a more comprehensive view of the molecular mechanisms including the regulatory mechanisms can be achieved and more potential targets for treatment of cancer can be discovered. Based on the analyses carried out, 3 significant modules have been identified from the PPI network. Five pathways/processes have been found to be significantly associated with nor-NOHA and NOHA associated genes. Out of the 1996 proteins in the PPI network, 4 have been identified as hub proteins-SOD, SOD1, AMD1, and NOS2. These 4 proteins have been implicated in cancer by other studies. Thus, this study provided further validation into the claim of these 4 proteins being potential targets for cancer treatment.

systems biology

Drosophila embryogenesis scales uniformly across temperature and developmentally diverse species

Temperature affects both the timing and outcome of animal development, but the detailed effects of temperature on the progress of early development have been poorly characterized. To determine the impact of temperature on the order and timing of events during Drosophila melanogaster embryogenesis, we used time-lapse imaging to track the progress of embryos from shortly after egg laying through hatching at seven precisely maintained temperatures between 17.5{degrees}C and 32.5{degrees}C. We employed a combination of automated and manual annotation to determine when 36 milestones occurred in each embryo. D. melanogaster embryogenesis takes 33 hours at 17.5{degrees}C, and accelerates with increasing temperature to a low of 16 hours at 27.5{degrees}C, above which embryogenesis slows slightly. Remarkably, while the total time of embryogenesis varies over two fold, the relative timing of events from cellularization through hatching is constant across temperatures. To further explore the relationship between temperature and embryogenesis, we expanded our analysis to cover ten additional Drosophila species of varying climatic origins. Six of these species, like D. melanogaster, are of tropical origin, and embryogenesis time at different temperatures was similar for them all. D. mojavensis, a sub-tropical fly, develops slower than the tropical species at lower temperatures, while D. virilis, a temperate fly, exhibits slower development at all temperatures. The alpine sister species D. persimilis and D. pseudoobscura develop as rapidly as tropical flies at cooler temperatures, but exhibit diminished acceleration above 22.5{degrees}C and have drastically slowed development by 30{degrees}C. Despite ranging from 13 hours for D. erecta at 30{degrees}C to 46 hours for D. virilis at 17.5{degrees}C, the relative timing of events from cellularization through hatching is constant across all of the species and temperatures examined here, suggesting the existence of a previously unrecognized timer controlling the progress of embryogenesis that has been tuned by natural selection in response to the thermal environment in which each species lives.\n\nAuthor SummaryTemperature profoundly impacts the rate of development of \"cold-blooded\" animals, which proceeds far faster when it is warm. There is, however, no universal relationship. Closely related species can develop at markedly different speeds at the same temperature, likely resulting from environmental adaptation. This creates a major challenge when comparing development among species, as it is unclear whether they should be compared at the same temperature or under different conditions to maintain the same developmental rate. Facing this challenge while working with flies (Drosophila species), we found there was little data to inform this decision. So, using time-lapse imaging, precise temperature-control, and computational and manual video-analysis, we tracked the complex process of embryogenesis in 11 species at seven different temperatures. There was over a three-fold difference in developmental rate between the fastest species at its fastest temperature and the slowest species at its slowest temperature. However, our finding that the timing of events within development all scaled uniformly across species and temperatures astonished us. This is good news for developmental biologists, since we can induce species to develop nearly identically by growing them at different temperatures. But it also means flies must possess some unknown clock-like molecular mechanism driving embryogenesis forward.

Developmental Biology