Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Bioinformatics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Cytokinin perception in potato: New features of canonic players

Potato is the most economically important non-cereal food crop. Tuber formation in potato is regulated by phytohormones, cytokinins (CKs) in particular. The present work was aimed to study CK signal perception in potato. The sequenced potato genome of doubled monoploid Phureja was used for bioinformatic analysis and as a tool for identification of putative CK receptors from autotetraploid potato cv. Desiree. All basic elements of multistep phosphorelay (MSP) required for CK signal transduction were identified in Phureja genome, including three genes orthologous to three CK receptor genes (AHK 2-4) of Arabidopsis. As distinct from Phureja, autotetraploid potato contains at least two allelic isoforms of each receptor type. Putative receptor genes from Desiree plants were cloned, sequenced and expressed, and main characteristics of encoded proteins, firstly their consensus motifs, structure models, ligand-binding properties, and the ability to transmit CK signal, were determined. In all studied aspects the predicted sensor histidine kinases met the requirements for genuine CK receptors. Expression of potato CK receptors was found to be organ-specific and sensitive to growth conditions, particularly to sucrose content. Our results provide a solid basis for further in-depth study of CK signaling system and biotechnological improvement of potato.

plant biology

A pseudogene of caffeic acid-o-methyltransferase (COMT) in Acacia mangium: Comparative analysis with other COMT plant promoters

Acacia mangium is a prominent tree species in the forest plantation industry of Southeast Asia, grown mainly to produce pulp and paper, and to a lesser extent wood chips and solid wood products. Lignin, a natural complex polymer used by plants for structural support and defence, has to be chemically removed during the production of quality paper. Delignification is very expensive and moreover, is an environmental pollutant. Understanding the complex mechanisms that underlie the regulation of lignin biosynthetic genes requires in-depth knowledge of not only the genes involved but also their regulatory elements. Using Thermal Asymmetric Interlaced PCR, a 770 bp promoter sequence with 93% identity with COMT1 gene from Acacia auriculiformis x A. mangium hybrid was isolated from A. mangium. Bioinformatics analysis revealed the presence of cis acting elements commonly found in other lignin biosynthesis genes such as TATA box, CAAT box, W box, AC-I and AC-11 elements. However, a nonsense mutation that created a premature stop codon was found on the first exon. Modelling of MYB transcription factor binding site on this newly isolated pseudogene shows it has binding sites for important transcription factors involved in lignin biosynthesis both in Arabidopsis thaliana and Eucalyptus grandis. Given the remarkable structures of its regulatory region, the possible structure of its transcript was detected using Mfold. Results show the transcript are capable of forming stem loop structures, a characteristic commonly attributed to presence of miRNA. Possible functions of pseudoAmCOMT1 were discussed.

genetics

Optimization and uncertainty analysis of ODE models using second order adjoint sensitivity analysis

MotivationParameter estimation methods for ordinary differential equation (ODE) models of biological processes can exploit gradients and Hessians of objective functions to achieve convergence and computational efficiency. However, the computational complexity of established methods to evaluate the Hessian scales linearly with the number of state variables and quadratically with the number of parameters. This limits their application to low-dimensional problems.\n\nResultsWe introduce second order adjoint sensitivity analysis for the computation of Hessians and a hybrid optimization-integration based approach for profile likelihood computation. Second order adjoint sensitivity analysis scales linearly with the number of parameters and state variables. The Hessians are effectively exploited by the proposed profile likelihood computation approach. We evaluate our approaches on published biological models with real measurement data. Our study reveals an improved computational efficiency and robustness of optimization compared to established approaches, when using Hessians computed with adjoint sensitivity analysis. The hybrid computation method was more than two-fold faster than the best competitor. Thus, the proposed methods and implemented algorithms allow for the improvement of parameter estimation for medium and large scale ODE models.\n\nAvailabilityThe algorithms for second order adjoint sensitivity analysis are implemented in the Advance MATLAB Interface CVODES and IDAS (AMICI, https://github.com/ICB-DCM/AMICI/). The algorithm for hybrid profile likelihood computation is implemented in the parameter estimation toolbox (PESTO, https://github.com/ICB-DCM/PESTO/). Both toolboxes are freely available under the BSD license.\n\nContactjan.hasenauer@helmholtz-muenchen.de\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

Rare Variant Pathogenicity Triage and Inclusion of Synonymous Variants Improves Analysis of Disease Associations

Many G protein-coupled receptors (GPCRs) lack common variants that lead to reproducible genome-wide disease associations. Here we used rare variant approaches to assess the disease associations of 85 orphan or understudied GPCRs in an unselected cohort of 51,289 individuals. Rare loss-of-function variants, missense variants predicted to be pathogenic or likely pathogenic, and a subset of rare synonymous variants were used as independent data sets for sequence kernel association testing (SKAT). Strong, phenome-wide disease associations shared by two or more variant categories were found for 39% of the GPCRs. Validating the bioinformatics and SKAT analyses, functional characterization of rare missense and synonymous variants of GPR39, a Family A GPCR, showed altered expression and/or Zn2+-mediated signaling for members of both variant classes. Results support the utility of rare variant analyses for identifying disease associations for genes that lack common variants, while also highlighting the functional importance of rare synonymous variants.\n\nAuthor summaryRare variant approaches have emerged as a viable way to identify disease associations for genes without clinically important common variants. Rare synonymous variants are generally considered benign. We demonstrate that rare synonymous variants represent a potentially important dataset for deriving disease associations, here applied to analysis of a set of orphan or understudied GPCRs. Synonymous variants yielded disease associations in common with loss-of-function or missense variants in the same gene. We rationalize their associations with disease by confirming their impact on expression and agonist activation of a representative example, GPR39. This study highlights the importance of rare synonymous variants in human physiology, and argues for their routine inclusion in any comprehensive analysis of genomic variants as potential causes of disease.

genetics

The Staphylococcus aureus Two-Component System AgrAC Displays Four Distinct Genomic Arrangements That Delineate Genomic Virulence Factor Signatures

Two-component systems (TCSs) consist of a histidine kinase and a response regulator. Here, we evaluated the conservation of the AgrAC TCS among 149 completely sequenced S. aureus strains. It is composed of four genes: agrBDCA. We found that: i) AgrAC system (agr) was found in all but one of the 149 strains; ii) The agr positive strains were further classified into four agr types based on AgrD protein sequences, iii) the four agr types not only specified the chromosomal arrangement of the agr genes but also the sequence divergence of AgrC histidine kinase protein, which confers signal specificity, iv) the sequence divergence was reflected in distinct structural properties especially in the transmembrane region and second extracellular binding domain, and v) there was a strong correlation between the agr type and the virulence genomic profile of the organism. Taken together, these results demonstrate that bioinformatic analysis of the agr locus leads to a classification system that correlates with the presence of virulence factors and protein structural properties.

microbiology

A novel GATA-binding protein 4 gene variation associated with familial atrial septal defect

Atrial septal defect (ASD) is the most common congenital heart defect. Part of ASD exhibits familial predisposition, but the genetic mechanism remains largely unknown. In the current study, we use multiple methods to identify and confirm the gene associated with a familial ASD. Chromosomal microarray analyses, whole exome sequencing, Sanger sequencing, multiple bioinformatics programs, in silico protein structure modeling and molecular dynamics simulation were performed to predict the pathogenic of the variant gene. Dual-Luciferase reporter gene assay was performed to evaluate the influence of downstream target gene of the target variation. A novel, heterozygous, missense variant GATA-binding protein 4 (GATA4):c.958C>T, p.R320W was identified. An autosomal dominant inheritance pattern with incomplete penetrance was observed in the family. Multiple prediction indicate the variant in GATA4 to be deleterious. Molecular dynamics simulation further revealed that the variation of p.R320W could prevent the zinc finger of GATA4 from interacting with the DNA. Dual-Luciferase reporter assay demonstrated a significant decrease in transcriptional activity (0.90{+/-}0.099 vs 1.50{+/-}0.079, p = 0.001) of the variant GATA4 compared with the wild type. We believe the novel variation of GATA4 (c.958C>T, p.R320W) with a pattern of incomplete inheritance that may be highly associated with this familial ASD. The finding enriched our knowledge of variations that may associated with ASD.

genetics

25 Years of Molecular Biology Databases: A Study of Proliferation, Impact, and Maintenance

Online resources enable unfettered access to and analysis of scientific data and are considered crucial for the advancement of modern science. Despite the clear power of online data resources, including web-available databases, proliferation can be problematic due to challenges in sustainability and long-term persistence. As areas of research become increasingly dependent on access to collections of data, an understanding of the scientific communitys capacity to develop and maintain such resources is needed.\n\nThe advent of the Internet coincided with expanding adoption of database technologies in the early 1990s, and the molecular biology community was at the forefront of using online databases to broadly disseminate data. The journal Nucleic Acids Research has long published articles dedicated to the description of online databases, as either debut or update articles. Snapshots throughout the entire history of online databases can be found in the pages of Nucleic Acids Research s \"Database Issue.\" Given the prominence of the Database Issue in the molecular biology and bioinformatics communities and the relative rarity of consistent historical documentation, database articles published in Database Issues provide a particularly unique opportunity for longitudinal analysis.\n\nTo take advantage of this opportunity, the study presented here first identifies each unique database described in 3055 Nucleic Acids Research Database Issue articles published between 1991-2016 to gather a rich dataset of databases debuted during this time frame, regardless of current availability. In total, 1727 unique databases were identified and associated descriptive statistics were gathered for each, including year debuted in a Database Issue and the number of all associated Database Issue publications and accompanying citation counts. Additionally, each database identified was assessed for current availability through testing of all associated URLs published. Finally, to assess maintenance, database websites were inspected to determine the last recorded update. The resulting work allows for an examination of the overall historical trends, such as the rate of database proliferation and attrition as well as an evaluation of citation metrics and on-going database maintenance.

molecular biology

Evaluation of Whole Exome Sequencing as an Alternative of BeadChip and Whole Genome Sequencing in Human Population Genetic Analysis

Understanding the underlying genetic structure of human populations is of fundamental interest to both biological and social sciences. Advances in high-throughput genotyping technology have markedly improved our understanding of global patterns of human genetic variation. The most widely used methods for collecting variant information at the DNA-level include whole genome sequencing, which continues to remain costly, and the more economical solution of array-based techniques, as these are capable of simultaneously genotyping a pre-selected set of variable DNA sites in the human genome. The largest publicly accessible set of human genomic sequence data available today originates from exome sequencing that comprises around 1.2% of the whole genome (approximately 30 million base pairs). In this study, we compared the application of the exome dataset to the array-based dataset and to the gold standard whole genome dataset using the same population genetic analysis methods. Our results draw attention to some of the inherent problems that arise from using pre-selected SNP sets for population genetic analysis. Additionally, we demonstrate that exome sequencing provides a better alternative to the array-based methods for population genetic analysis. In this study, we propose a strategy for unbiased variant collection from exome data and offer a bioinformatics protocol for proper data processing.

genomics

Reproducible integration of multiple sequencing datasets to form high-confidence SNP, indel, and reference calls for five human genome reference materials

Benchmark small variant calls from the Genome in a Bottle Consortium (GIAB) for the CEPH/HapMap genome NA12878 (HG001) have been used extensively for developing, optimizing, and demonstrating performance of sequencing and bioinformatics methods. Here, we develop a reproducible, cloud-based pipeline to integrate multiple sequencing datasets and form benchmark calls, enabling application to arbitrary human genomes. We use these reproducible methods to form high-confidence calls with respect to GRCh37 and GRCh38 for HG001 and 4 additional broadly-consented genomes from the Personal Genome Project that are available as NIST Reference Materials. These new genomes broad, open consent with few restrictions on availability of samples and data is enabling a uniquely diverse array of applications. Our new methods produce 17% more high-confidence SNPs, 176% more indels, and 12% larger regions than our previously published calls. To demonstrate that these calls can be used for accurate benchmarking, we compare other high-quality callsets to ours (e.g., Illumina Platinum Genomes), and we demonstrate that the majority of discordant calls are errors in the other callsets, We also highlight challenges in interpreting performance metrics when benchmarking against imperfect high-confidence calls. We show that benchmarking tools from the Global Alliance for Genomics and Health can be used with our calls to stratify performance metrics by variant type and genome context and elucidate strengths and weaknesses of a method.

genomics

Protein-Protein Interaction Network Analysis and Identification of Key Players in N-hydroxy-nor-L-Arg (nor-NOHA) and N(omega)-hydroxy-L-arginine (NOHA) Mediated Pathways for Treatment of Cancer Through Arginase Inhibiton: Insights from Systems Biology

L-arginine is involved in a number of biological processes in our bodies. Metabolism of L-arginine by the enzyme arginase has been found to be associated with cancer cell proliferation. Arginase inhibition has been proposed as a potential therapeutic means to inhibit this process. N-hydroxy-nor-L-Arg (nor-NOHA) and N (omega)-hydroxy-L-arginine (NOHA) has shown promise in inhibiting cancer progression through arginase inhibition. In this study, nor-NOHA and NOHA-associated genes and proteins were analyzed with several Bioinformatics and Systems Biology tools to identify the associated pathways and the key players involved so that a more comprehensive view of the molecular mechanisms including the regulatory mechanisms can be achieved and more potential targets for treatment of cancer can be discovered. Based on the analyses carried out, 3 significant modules have been identified from the PPI network. Five pathways/processes have been found to be significantly associated with nor-NOHA and NOHA associated genes. Out of the 1996 proteins in the PPI network, 4 have been identified as hub proteins-SOD, SOD1, AMD1, and NOS2. These 4 proteins have been implicated in cancer by other studies. Thus, this study provided further validation into the claim of these 4 proteins being potential targets for cancer treatment.

systems biology

Long-read DNA metabarcoding of ribosomal rRNA in the analysis of fungi from aquatic environments

DNA metabarcoding is now widely used to study prokaryotic and eukaryotic microbial diversity. Technological constraints have limited most studies to marker lengths of ca. 300-600 bp. Longer sequencing reads of several 5 thousand bp are now possible with third-generation sequencing. The increased marker lengths provide greater taxonomic resolution and enable the use of phylogenetic methods of classifcation, but longer reads may be subject to higher rates of sequencing error and chimera formation. In addition, most well-established bioinformatics tools for DNA metabarcoding were originally 10 designed for short reads and are therefore not suitable. Here we used Pacifc Biosciences circular consensus sequencing (CCS) to DNA-metabarcode environmental samples using a ca. 4,500 bp marker that included most of the eukaryote ribosomal SSU and LSU rRNA genes and the ITS spacer region. We developed a long-read analysis pipeline that reduced error rates to levels 15 comparable to short-read platforms. Validation using fungal isolates and a mock community indicated that our pipeline detected 98% of chimeras de novo i.e., even in the absence of reference sequences. We recovered 947 OTUs from water and sediment samples in a natural lake, 848 of which could be classifed to phylum, 486 to family, 397 to genus and 330 to species. By 20 allowing for the simultaneous use of three global databases (Unite, SILVA, RDP LSU), long-read DNA metabarcoding provided better taxonomic resolution than any single marker. We foresee the use of long reads enabling the cross-validation of reference sequences and the synthesis of ribosomal rRNA gene databases. The universal nature of the rRNA operon and our recovery of >100 25 non-fungal OTUs indicate that long-read DNA metabarcoding holds promise for the study of eukaryotic diversity more broadly.

ecology

Rapid Paediatric Sequencing (RaPS): Comprehensive real-life workflow for rapid diagnosis of critically ill children

BackgroundRare genetic conditions are frequent risk factors for, or direct causes of, organ failure requiring paediatric intensive care unit (PICU) support. Such conditions are frequently suspected but unidentified at PICU admission. Compassionate and effective care is greatly assisted by definitive diagnostic information. There is therefore a need to provide a rapid genetic diagnosis to inform clinical management.\n\nTo date, Whole Genome Sequencing (WGS) approaches have proved successful in diagnosing a proportion of children with rare diseases, but results may take months to report or require the use of equipment and practices not compatible with a clinical diagnostic setting. We describe an end-to-end workflow for the use of rapid WGS for diagnosis in critically ill children in a UK National Health Service (NHS) diagnostic setting.\n\nMethodsWe sought to establish a multidisciplinary Rapid Paediatric Sequencing (RaPS) team for case selection, trio WGS, a rapid bioinformatics pipeline for sequence analysis and a phased analysis and reporting system to prioritise genes with a high likelihood of being causal. Our workflow was iteratively developed prospectively during the analysis of the first 10 children and applied to the following 14 to assess its utility.\n\nFindingsTrio WGS in 24 critically ill children led to a molecular diagnosis in ten (42%) through the identification of causative genetic variants. In three of these ten individuals (30%) the diagnostic result had an immediate impact on the individuals clinical management. For the last 14 trios, the shortest time taken to reach a provisional diagnosis was four days (median 7 days).\n\nInterpretationRapid WGS can be used to diagnose and inform management of critically ill children using widely available off the shelf products within the constraints of an NHS clinical diagnostic setting. We provide a robust workflow that will inform and facilitate the rollout of rapid genome sequencing in the NHS and other healthcare systems globally.\n\nFundingThe study was funded by NIHR GOSH/UCL BRC: ormbrc-2012-1

genomics

Genome-specific histories of divergence and introgression between an allopolyploid unisexual salamander lineage and two sexual species

Quantifying genetic introgression between sexual species and polyploid lineages traditionally thought to be asexual is an important step in understanding what factors drive the longevity of putatively asexual groups. However, the presence of multiple distinct subgenomes within a single lineage provides a significant logistical challenge to evaluating the origin of genetic variation in most polyploids. Here, we capitalize on three recent innovations--variation generated from ultraconserved elements (UCEs), bioinformatic techniques for assessing variation in polyploids, and model-based methods for evaluating historical gene flow--to measure the extent and tempo of introgression over the evolutionary history of an allopolyploid lineage of all-female salamanders and two ancestral sexual species. We first analyzed variation from more than a thousand UCEs using a reference mapping method developed for polyploids to infer subgenome specific patterns of variation in the all-female lineage. We then used PHRAPL to choose between sets of historical models that reflected different patterns of introgression and divergence between the genomes of the parental species and the same genomes found within the polyploids. Our analyses support a scenario in which the genomes sampled in unisexuals salamanders were present in the lineage [~]3.4 million years ago, followed by an extended period of divergence from their parental species. Recent secondary introgression has occurred at different times between each sexual species and their representative genomes within the unisexuals during the last 500,000 years. Sustained introgression of sexual genomes into the unisexual lineage has been the defining characteristic of their reproductive mode, but this study provides the first evidence that unisexual genomes have also undergone long periods of divergence without introgression. Unlike other unisexual, sperm-dependent taxa in which introgression is rare, the alternating periods of divergence and introgression between unisexual salamanders and their sexual relatives could reveal the scenarios in which the influx of novel genomic material is favored and potentially explain why these salamanders are among the oldest described unisexual animals.

evolutionary biology

Single-molecule optical mapping enables accurate molecular diagnosis of facioscapulohumeral muscular dystrophy (FSHD)

Facioscapulohumeral Muscular Dystrophy (FSHD) is a common adult muscular dystrophy in which the muscles of the face, shoulder blades and upper arms are among the most affected. FSHD is the only disease in which \"junk\" DNA is reactivated to cause disease, and the only known repeat array-related disease where fewer repeats cause disease. More than 95% of FSHD cases are associated with copy number loss of a 3.3kb tandem repeat (D4Z4 repeat) at the subtelomeric chromosomal region 4q35, of which the pathogenic allele contains less than 10 repeats and has a specific genomic configuration called 4qA. Currently, genetic diagnosis of FSHD requires pulsed-field gel electrophoresis followed by Southern blot, which is labor-intensive, semi-quantitative and requires long turnaround time. Here, we developed a novel approach for genetic diagnosis of FSHD, by leveraging Bionano Saphyr single-molecule optical mapping platform. Using a bioinformatics pipeline developed for this assay, we found that the method gives direct quantitative measurement of repeat numbers, can differentiate 4q35 and the highly paralogous 10q26 regions, can determine the 4qA/4qB allelic configuration, and can quantitate levels of post-zygotic mosaicism. We evaluated this approach on 5 patients (including two with post-zygotic mosaicism) and 2 patients (including one with post-zygotic mosaicism) from two separate cohorts, and had complete concordance with Southern blots, but with improved quantification of repeat numbers resolved between haplotypes. We concluded that single-molecule optical mapping is a viable approach for molecular diagnosis of FSHD and may be applied in clinical diagnostic settings once more validations are performed.

genomics

Controlling false discoveries in Bayesian gene networks with lasso regression p-values

MotivationBayesian networks can represent directed gene regulations and therefore are favored over co-expression networks. However, hardly any Bayesian network study concerns the false discovery control (FDC) of network edges, leading to low accuracies due to systematic biases from inconsistent false discovery levels in the same study.\n\nResultsWe design four empirical tests to examine the FDC of Bayesian networks from three p-value based lasso regression variable selections -- two existing and one we originate. Our method, lassopv, computes p-values for the critical regularization strength at which a predictor starts to contribute to lasso regression. Using null and Geuvadis datasets, we find that lassopv obtains optimal FDC in Bayesian gene networks, whilst existing methods have defective p-values. The FDC concept and tests extend to most network inference scenarios and will guide the design and improvement of new and existing methods. Our novel variable selection method with lasso regression also allows FDC on other datasets and questions, even beyond network inference and computational biology.\n\nAvailabilityLassopv is implemented in R and freely available at https://github.com/lingfeiwang/lassopv and https://cran.r-project.org/package=lassopv.\n\nContactLingfei.Wang@roslin.ed.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

systems biology

THE MITOCHONDRIAL DNA CONTROL REGION MIGHT HAVE USEFUL DIAGNOSTIC AND PROGNOSTIC BIOMARKERS FOR THYROID TUMORS

BackgroundIt is currently present in the literature that mitochondrial DNA (mtDNA) defects are associated with a great number of diseases including cancers. The role of mitochondrial DNA (mtDNA) variations in the development of thyroid cancer is a highly controversial topic. In this study, we investigated the role of mt-DNA control region (CR) variations in thyroid tumor progression and the influence of mtDNA haplogroups on susceptibility to thyroid tumors.\n\nMaterial & methodFor this purpose, totally 108 hot thyroid nodules (HTNs), 95 cold thyroid nodules (CTNs), 48 papillary thyroid carcinoma (PTC) samples with their surrounding tissues and 104 healthy control subjects blood samples were screened for entire mtDNA CR variations by using Sanger sequencing. The obtained DNA sequences were anaysed with the mistomaster, a web-based bioinformatics tool.\n\nResultsMtDNA haplogroup U was significantly associated with susceptibility to benign and malign thyroid entities on the other hand J haplogroup was associated with a protective role for benign thyroid nodules. Besides, 8 SNPs (T146C, G185A, C194T, C295T, G16129A, T16304C, A16343G and T16362C) in mtDNA CR region were associated with the occurrence of benign and malign thyroid nodules in Turkish population. By contrast with the healthy Turkish population and HTNs, frequency of C7 repeats in D310 polycytosine sequence was found higher in cold thyroid nodules and PTC samples. Beside this, the frequency of somatic mutations in mtMSI regions including T16189C and D514 CA dinucleotide repeats were found higher in PTC samples than the benign thyroid nodules. Conversely, the frequency of somatic mutations in D310 was detected higher in HTNs than CTNs and PTCs.\n\nConclusionmtDNA D310 instability do not play a role in tumorogenesis of the PTC but the results indicates that it might be used as a diagnostic clonal expansion biomarker for premalignant thyroid tumor cells. Beside this, D514 CA instability might be used as prognostic biomarker in PTCs. Also, we showed that somatic mutation rate is less frequent in more aggressive tumors when we examined micro- and macro carcinomas as well as BRAFV600E mutation.

cancer biology

Natural selection on gene-specific codon usage bias is common across eukaryotes

Although the actual molecular evolutionary forces that shape differences in codon usage across species remain poorly understood, majority of synonymous mutations are assumed to be functionally neutral because they do not affect protein sequences. However, empirical studies suggest that some synonymous mutations can have phenotypic consequences. Here we show that in contrast to the current dogma, natural selection on gene-specific codon usage bias is common across Eukaryota. Furthermore, by using bioinformatic and experimental approaches, we demonstrate that specific combinations of rare codons contribute to the spatial and sex-related regulation of some protein-coding genes in Drosophila melanogaster. Together, these data indicate that natural selection can shape gene-specific codon usage bias, which therefore, represents an overlooked genomic feature that is likely to play an important role in the spatial and temporal regulation of gene functions. Hence, the broadly accepted dogma that synonymous mutations are in general functionally neutral should be reconsidered.

evolutionary biology

LIkelihood-based Fits of Folding Transitions (LIFFT) for Biomolecule Mapping Data

SummaryBiomolecules shift their structures as a function of temperature and concentrations of protons, ions, small molecules, proteins, and nucleic acids. These transitions impact or underlie biological function and are being monitored at increasingly high throughput. For example, folding transitions for large collections of RNAs can now be monitored at single residue resolution by chemical mapping techniques. LIkelihood-based Fits of Folding Transitions (LIFFT) quantifies these data through well-defined thermodynamic models. LIFFT implements a Bayesian framework that takes into account data at all measured residues and enables visual assessment of modeling uncertainties that can be overlooked in least-squares fits. The framework is appropriate for multimodal techniques ranging from chemical mapping including multi-wavelength spectroscopy.\n\nAvailabilityFreely available MATLAB package at https://ribokit.stanford.edu/LIFFT/.\n\nContactrhiju@stanford.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

biochemistry