Search bioRxivSearch

SEARCH · Search bioRxiv

Search Search bioRxiv

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,441 records · Page 80Linked to original sources

Quantitative network organization of interactions emerging from the evolution of sequence and structure space of antibodies: the RADARS model

Adaptive immunity in vertebrates represents a complex self-organizing network of protein interactions that develops throughout the lifetime of an individual. While deep sequencing of the immune-receptor repertoire may reveal clonal relationships, functional interpretation of such data is hampered by the inherent limitations of converting sequence to structure to function. In this paper a novel model of antibody interaction space and network, termed radial adjustment of system resolution, RADARS, is proposed. The model is based on the radial growth of interaction affinity of antibodies towards an infinity of directions in structure space, each direction representing particular shapes of antigen epitopes. Levels of interaction affinity appear as free energy shells of the system, where hierarchical B-cell development and differentiation takes place. Equilibrium in this immunological thermodynamic system can be described by a power-law distribution of antibody free energies with an ideal network degree exponent of phi square, representing a scale-free fractal network of antibody interactions. Plasma cells are network hubs, memory B cells are nodes with intermediate degrees and B1 cells represent nodes with minimal degree. Thus, the RADARS model implies that antibody structure space develops against an infinite antigen structure space via interactions that are individually immunologically controlled, but on a systems level are organized by thermodynamic probability distributions. The network of interactions, which control B-cell development and differentiation, represent pathways of antigen removal on systems level. Understanding such quantitative network properties of the system should help the organization of sequence-derived structural data, offering the possibility to relate sequence to function in a complex, self-organizing biological system. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=150 SRC="FIGDIR/small/438804v7_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@883eb7org.highwire.dtl.DTLVardef@cd7d5forg.highwire.dtl.DTLVardef@e225eeorg.highwire.dtl.DTLVardef@1286b49_HPS_FORMAT_FIGEXP M_FIG Graphical abstract C_FIG

immunology

Genome-wide signatures of local adaptation among seven stoneflies species along a nationwide latitudinal gradient in Japan

BackgroundEnvironmental heterogeneity continuously produces a selective pressure that results in genomic variation among organisms; understanding this relationship remains a challenge in evolutionary biology. Here, we evaluated the degree of genome-environmental association of seven stonefly species across a wide geographic area in Japan and additionally identified putative environmental drivers and their effect on co-existing multiple stonefly species. Double-digest restriction-associated DNA (ddRAD) libraries were independently sequenced for 219 individuals from 23 sites across four geographical regions along a nationwide latitudinal gradient in Japan.\n\nResultsA total of 4,251 candidate single nucleotide polymorphisms (SNPs) strongly associated with local adaptation were discovered using Latent mixed models; of these, 294 SNPs showed strong correlation with environmental variables, specifically precipitation and altitude, using distance-based redundancy analysis. Genome-genome comparison among the seven species revealed a high sequence similarity of candidate SNPs within a geographical region, suggesting the occurrence of a parallel evolution process.\n\nConclusionsOur results revealed genomic signatures of local adaptation and their influence on multiple, co-occurring species. These results can be potentially applied for future studies on river management and climatic stressor impacts.

genomics

The allometry of brain size in mammals

Why some animals have big brains and others do not has intrigued scholars for millennia. Yet, the taxonomic scope of brain size research is limited to a few mammal lineages. Here we present a brain size dataset compiled from the literature for 1552 species with representation from 28 extant taxonomic orders. The brain-body size allometry across all mammals is (Brain) = -1.26 (Body)0.75. This relationship shows strong phylogenetic signal as expected due to shared evolutionary histories. Slopes using median species values for each order, family, and genus, to ensure evolutionary independence, approximate [~]0.75 scaling. Why brain size scales to the [3/4] power to body size across mammals is, to our knowledge, unknown. Slopes within taxonomic orders exhibiting smaller size ranges are often shallower than 0.75 and range from 0.24 to 0.81 with a median slope of 0.64. Published brain size data is lacking for the majority of extant mammals (>70% of species) with strong bias in representation from Primates, Carnivores, Perrisodactyla, and Australidelphian marsupials (orders Dasyuromorphia, Diprotodontia, Peramelemorphia). Several orders are particularly underrepresented. For example, brain size data are available for less than 20% of species in each of the following speciose lineages: Soricomorpha, Rodentia, Lagomorpha, Didelphimorphia, and Scandentia. Use of museum collections can decrease the current taxonomic bias in mammal brain size data and tests of hypothesis.

animal behavior and cognition

An averaging strategy to reduce variability in target-decoy estimates of false discovery rate

Decoy database search with target-decoy competition (TDC) provides an intuitive, easy-to-implement method for estimating the false discovery rate (FDR) associated with spectrum identifications from shotgun proteomics data. However, the procedure can yield different results for a fixed dataset analyzed with different decoy databases, and this decoy-induced variability is particularly problematic for smaller FDR thresholds, datasets or databases. In such cases, the nominal FDR might be 1% but the true proportion of false discoveries might be 10%. The averaged TDC protocol combats this problem by exploiting multiple independently shuffled decoy databases to provide an FDR estimate with reduced variability. We provide a tutorial introduction to aTDC, describe an improved variant of the protocol that offers increased statistical power, and discuss how to deploy aTDC in practice using the Crux software toolkit.

bioinformatics

Algorithmic biosynthesis of eukaryotic glycans

An algorithm converts inputs to corresponding unique outputs through a sequence of actions. Algorithms are used as metaphors for complex biological processes such as organismal development. Here we make this metaphor rigorous for glycan biosynthesis. Glycans are branched sugar oligomers that are attached to cell-surface proteins and convey cellular identity. Eukaryotic O-glycans are synthesized by collections of enzymes in Golgi compartments. A compartment can stochastically convert a single input oligomer to a heterogeneous set of possible output oligomers; yet a given type of protein is invariably associated with a narrow and reproducible glycan oligomer profile. Here we resolve this paradox by borrowing from the theory of algorithmic self-assembly. We rigorously enumerate the sources of glycan microheterogeneity: incomplete oligomers via early exit from the reaction compartment; tandem repeat oligomers via runaway reactions; and competing oligomer fates via divergent reactions. We demonstrate how to diagnose and eliminate each of these, thereby obtaining \"algorithmic compartments\" that convert inputs to corresponding unique outputs. Given an input and a target output we either prove that the output cannot be algorithmically synthesized from the input, or explicitly construct an ordered series of algorithmic compartments that achieves this synthesis. Our theoretical analysis allows us to infer the causes of non-algorithmic microheterogeneity and species-specific diversity in real glycan datasets.

systems biology

Brain responses to anticipating and receiving beer: Comparing light, at-risk, and dependent alcohol users

BackgroundImpaired brain processing of alcohol-related rewards has been suggested to play a central role in alcohol use disorder. Yet, evidence remains inconsistent, and mainly originates from studies in which participants passively observe alcohol cues or taste alcohol. Here we designed a protocol in which beer consumption was predicted by incentive cues and contingent on instrumental action, closer to real life situations. We predicted that anticipating and receiving beer (compared with water) would elicit activity in the brain reward network, and that this activity would correlate with drinking level across participants.\n\nMethodsThe sample consisted of 150 beer-drinking males, aged 18-25 years. Three groups were defined based on AUDIT scores: light drinkers (n=40), at-risk drinkers (n=63), and dependent drinkers (n=47). fMRI measures were obtained while participants engaged in the Beer Incentive Delay task involving beer- and water-predicting cues, followed by real sips of beer or water.\n\nResultsDuring anticipation, outcome notification and delivery of beer compared with water, higher activity was found in a reward-related brain network including the medial prefrontal cortex, orbitofrontal cortex and amygdala. Yet, no activity was observed in the striatum, and no differences were found between the groups.\n\nConclusionsOur results reveal that anticipating, obtaining and tasting beer activates parts of the brain reward network, but that these brain responses do not differentiate between different drinking levels. We speculate that other factors, such as cognitive control or sensitivity to social context, may be more discriminant predictors of drinking behaviour in young adults.

neuroscience

Inferring linguistic transmission between generations at the scale of individuals

Historical linguistics strongly benefited from recent methodological advances inspired by phylogenetics. Nevertheless, no available method uses contemporaneous within-population linguistic diversity to reconstruct the history of human populations. Here, we developed an approach inspired from population genetics to perform historical linguistic inferences from linguistic data sampled at the individual scale, within a population. We built four within-population demographic models of linguistic transmission over generations, each differing by the number of teachers involved during the language acquisition and the relative roles of the teachers. We then compared the simulated data obtained with these models with real contemporaneous linguistic data sampled from Tajik speakers from Central Asia, an area known for its large within-population linguistic diversity, using approximate Bayesian computation methods. Under this statistical framework, we were able to select the models that best explained the data, and infer the best-fitting parameters under the selected models. This demonstrates the feasibility of using contemporaneous within-population linguistic diversity to infer historical features of human cultural evolution.

bioinformatics

Estimation of allele-specific fitness effects across human protein-coding sequences and implications for disease

A central challenge in human genomics is to understand the cellular, evolutionary, and clinical significance of genetic variants. Here we introduce a unified population-genetic and machine-learning model, called Linear Allele-Specific Selection InferencE (LASSIE), for estimating the fitness effects of all potential single-nucleotide variants, based on polymorphism data and predictive genomic features. We applied LASSIE to 51 high-coverage genome sequences annotated with 33 genomic features, and constructed a map of allele-specific selection coefficients across all protein-coding sequences in the human genome. We show that this map is informative about both human evolution and disease.

genomics

Algorithmic improvements for discovery of germline copy number variants in next-generation sequencing data

Copy number variants (CNVs) play a significant role in human heredity and disease, however sensitive and specific characterization of CNVs from NGS data has remained challenging. Detection is especially problematic for hybridization-capture data in which read counts are the sole source of copy number information. We describe two algorithmic adaptations that improve CNV detection accuracy in a Hidden Markov Model (HMM) context. First, we present a method for com puting target- and copy number state-specific emission distributions. Second, we demonstrate that the Pointwise Maximum a posteriori (PMAP) HMM decoding procedure yields improved sensitivity for small CNV calls compared to the more common Viterbi HMM decoder. We develop a prototype implementation, called Cobalt, and compare it to other CNV detection tools using sets of simulated and previously detected CNVs with sizes spanning a single exon up to a full chromosome. In both the simulation and previously detected CNV studies Cobalt shows similar sensitivity but significantly improved positive predictive value (PPV) compared to other callers. Overall sensitivity is 80%-90% for deletion CNVs spanning 1-4 targets and 90%-100% for larger deletion events, while sensitivity is somewhat lower for small duplication CNVs. Cobalt demonstrates significantly improved positive predictive value (PPV) compared to other callers with similar sensitivity, typically making 5X fewer total calls overall.

bioinformatics

Serum triglycerides in Alzheimer’s disease: Relation to neuroimaging and CSF biomarkers

ObjectiveTo investigate the association of triglyceride (TG) principal component scores with Alzheimers disease (AD) and the \"A/T/N/V\" (Amyloid, Tau, Neurodegeneration, and Cerebrovascular disease) biomarkers for AD.\n\nMethodsSerum levels of 84 TG species were measured using untargeted lipid profiling of 689 participants from the Alzheimers Disease Neuroimaging Initiative (ADNI) cohort including 190 cognitively normal older adults (CN) and 339 mild cognitive impairment (MCI) and 160 AD. Principal component analysis with factor rotation was used for dimension reduction of TG species. Differences in principal components between diagnostic groups and associations between principal components and AD biomarkers (including CSF, MRI and [18F]FDG-PET) were assessed using a multivariate generalized linear model (GLM) approach. In both cases, the Bonferroni method of adjustment was employed to correct for multiple comparisons.\n\nResultsThe 84 TGs yielded 9 principal components, two of which consisting of long-chain, polyunsaturated fatty acid-containing TGs (PUTGs), were significantly associated with MCI and AD. Lower levels of PUTGs were observed in MCI and AD compared to CN. PUTG principal component scores were also significantly associated with hippocampal volume and entorhinal cortical thickness. In participants carrying APOE {varepsilon}4 allele, these principal components were significantly associated with CSF amyloid-{beta}1-42 values and entorhinal cortical thickness.\n\nConclusionsThis study shows PUTG component scores significantly associated with diagnostic group and AD biomarkers, a finding that was more pronounced in APOE {varepsilon}4 carriers. Replication in independent larger studies and longitudinal follow-up are warranted.

neuroscience

A Novel Dual And Triple RSVP Paradigm For P300 Speller

ObjectiveA speller system enables disabled people, specifically those with spinal cord injuries, to visually select and spell characters. A problem of primary speller systems is that they are gaze shift dependent. To overcome this problem, a single RSVP paradigm was introduced in which characters are displayed one by one at the center of a screen. In this paper, two new protocols named Dual and Triple RSVP paradigms are introduced and their results are compared against the single paradigm.\n\nMethodsIn the Dual and Triple paradigms, two and three characters are displayed at the center of the screen simultaneously, therefore holding the advantage of displaying the target character twice and three times respectively, compared to the one-time appearance in the single paradigm. Subsequently, by reducing the number of repetitions in the Dual and Triple paradigms, it is expected that ITR decreases. To compare the results of these three paradigms, three subjects participated in experiments using all three paradigms.\n\nResultsThe offline results demonstrate an average character detection accuracy of 97% for the single and double protocols, and 80% for the Triple paradigm. In addition, average ITR is calculated to be 5.45, 7.62 and 7.90 bit/min for the single, Dual and Triple paradigms respectively. Results demonstrate an equally good character detection accuracy for the single and Dual paradigms, and a significant increase in ITR in the Dual paradigm compared to the single. The Triple RSVP paradigm demonstrates an almost equal ITR to that of the Dual paradigm, while decreasing character detection accuracy significantly.\n\nConclusionsResults demonstrate that the Dual RSVP paradigm can be recognized as the most suitable approach, by providing the best balance between ITR and character detection accuracy.\n\nSignificanceThis research demonstrates the improved performance of a newly proposed speller system (the Dual RSVP paradigm). By replacing existing methods with this new approach, the performance of speller matrices will be enhanced, and in addition the gaze dependency issue that caused limitations for users suffering from unimpaired oculomotor control will be overcome.

neuroscience

Predicted asymmetrical effects of warming on nocturnal and diurnal ectotherms

Many ectotherms restrict activity to times and places with favorable temperatures. This widespread pattern of habitat use in fluctuating environments may alter predictions of how climate change will affect ectotherms. By considering time elapsed within a range of suitable temperatures as a resource, I demonstrate that warming is expected to affect thermally restricted nocturnal and diurnal activity windows asymmetrically. Under warming scenarios, thermally restricted nocturnal activity windows lengthen while diurnal activity windows contract. This divergent prediction results from the shape of the function relating time to temperature within a day, which is typically concave during the day and convex during the night. This characteristic shape is nearly universal across terrestrial environments due to the changing angle of the sun throughout each day and exponential decay of overnight temperatures. These predicted asymmetries are exacerbated by expectations of diurnally asymmetric warming (more warming during the night compared to the day). Using example data from a montane ant community, I demonstrate that, as predicted, moderate simulated warming expands activity time available to cool active species and reduces activity time available to warm active species. Together these results suggest that the time of day during which an ectotherms optimal temperature occurs can be an important factor in determining response to warming.

ecology

Overdispersed gene expression characterizes schizophrenic brains

Schizophrenia (SCZ) is a severe, highly heterogeneous psychiatric disorder with varied clinical presentations. The polygenic genetic architecture of SCZ makes identification of causal variants daunting. Gene expression analyses have shown that SCZ may result in part from transcriptional dysregulation of a number of genes. However, most of these studies took the commonly used approach--differential gene expression analysis, assuming people with SCZ are a homogenous group, all with similar expression levels for any given gene. Here we show that the overall gene expression variability in SCZ is higher than that in an unaffected control (CTL) group. Specifically, we applied the test for equality of variances to the normalized expression data generated by the CommonMind Consortium (CMC) and identified 87 genes with significantly higher expression variances in the SCZ group than the CTL group. One of the genes with differential variability, VEGFA, encodes a vascular endothelial growth factor, supporting a vascular-ischemic etiology of SCZ. We also applied a Mahalanobis distance-based test for multivariate homogeneity of group dispersions to gene sets and identified 19 functional gene sets with higher expression variability in the SCZ group than the CTL group. Several of these gene sets are involved in brain development (e.g., development of cerebellar cortex, cerebellar Purkinje cell layer and neuromuscular junction), supporting that structural and functional changes in the cortex cause SCZ. Finally, using expression variability QTL (evQTL) analysis, we show that common genetic variants contribute to the increased expression variability in SCZ. Our results reveal that SCZ brains are characterized by overdispersed gene expression, resulting from dysregulated expression of functional gene sets pertaining to brain development, necrotic cell death, folic acid metabolism, and several other biological processes. Using SCZ as a model of complex genetic disorders with a heterogeneous etiology, our study provides a new conceptual framework for variability-centric analyses. Such a framework is likely to be important in the era of personalized medicine. (313 words)

genetics

Omnivory does not preclude strong trophic cascades

Omnivory has been cited as an explanation for why trophic cascades are weak in many ecosystems, but empirical support for this prediction is equivocal. Compared to predators that feed only on herbivores, top omnivores -- species that feed on both herbivores and primary producers -- have been observed generating cascades ranging from strong, to moderate, null, and negative. To gain intuition about the sensitivity of cascades to omnivory, we analyzed models describing systems with top omnivores that display either fixed or flexible diets, two foraging strategies that are supported by empirical observations. We identified regions of parameter space wherein omnivores following a fixed foraging strategy, with herbivores and producers comprising a constant proportion of the diet, non-intuitively generate stronger cascades than predators that are otherwise demographically identical: (i) high productivity relative to herbivore mortality, and (ii) small discrepancies in producer versus herbivore reward create conditions in which cascades are stronger with moderate omnivory. In contrast, flexible omnivores that attempt to optimize per capita growth rates during search never induce cascades that are stronger than the case of predators. Although we focus on simple models, the consistency of these general patterns together with prior empirical evidence suggests that omnivores should not be uniformly ruled out as agents of strong trophic cascades.

ecology

Sample-based regulatory intervention for managing risk within heterogeneous populations, with phytosanitary inspection of mixed consignments as a case study

Inspection of consignments of imported goods is commonly undertaken at national borders in order to prevent incursions of pests and diseases and deter malefactors. Inspection of the whole consignment is usually either impossible or inefficient, so inspection of a random sample is used instead.\n\nThe size of the random sample for plant products is justified by appeal to International Standards for Phytosanitary Measures No. 31, \"Methodologies for Sampling of Consignments\". ISPM 31 notes that \"A lot to be sampled should be a number of units of a single commodity identifiable by its homogeneity [...]\" and \"Treating multiple commodities as a single lot for convenience may mean that statistical inferences can not be drawn from the results of the sampling.\"\n\nHowever, commonly consignments are heterogeneous, either because the same commodities have multiple sources or because there are several different commodities. The ISPM 31 prescription creates a substantial impost on border inspection because it suggests that heterogeneous populations must be split into homogeneous sub-populations from which separate samples of nominal size must be taken.\n\nWe demonstrate that if consignments with known heterogeneity are treated as stratified populations and the random sample of units is allocated proportionally based on the number of units in each stratum, then the nominal sensitivity at the consignment level is achieved if our concern is the level of contamination in the entire consignment taken as a whole. We argue that unknown heterogeneity is no impediment to appropriate statistical inference. We conclude that the international standard is unnecessarily restrictive.

ecology

Development of a novel signature of long noncoding RNAs as a prognostic biomarker for esophageal cancer

ObjectivesThis study aims to develop a lncRNA signature based on RNA-Seq data to predict overall survival in esophageal cancer patients.\n\nMethodsThe lncRNA expression profiles and clinical data were downloaded from The Cancer Genome Atlas (TCGA) database on August 30, 2017. Differentially expressed lncRNAs were screened out between tumor tissues and adjacent normal tissues. The univariate and multivariate Cox regression models were used to develop a prognostic signature for all esophageal cancer patients. The receiver operating curve (ROC) was used to test the sensitivity and specificity of lncRNA signature. Survivals were compared via log-rank test. GO and KEGG enrichment analyses were used to explore the potential functions of prognostic lncRNAs.\n\nResultsWe identified two lncRNAs (RPL34-AS1 and GK3P) were significantly associated with the overall survival of the total 150 esophageal cancer patients. A novel two-lncRNA signature was constructed by Cox regression models. Signature low-risk cases showed better overall survival (median 625.560 days vs. 478.000 days, p = 0.002) than high-risk cases. Further analysis suggested that this two-lncRNA signature was independent of clinical characteristics. GO functional and KEGG pathway enrichment analyses revealed potential functional roles of the two prognostic lncRNAs in tumorigenesis.\n\nConclusionsOur findings suggest that the two-lncRNA signature may be a useful prognostic biomarker for predicting overall survival in esophageal cancer patients.

bioinformatics

Nomenclature Errors in Public 16S rRNA Gene Reference Databases

BackgroundTargeted gene surveys of the 16S rRNA gene have become a standard method for profiling the membership and biodiversity of microbial communities. These studies rely upon specialized databases that provide reference sequences and their corresponding taxonomic classifications, but few independent evaluations of the nomenclature used in the taxonomic classifications have been performed.\n\nResultsNomenclature data collected from the List of Prokaryotic names with Standing in Nomenclature, Prokaryotic Nomenclature Up-to-Date, and CyanoDB databases were used to validate the nomenclature contained in the taxonomic classifications in the Greengenes, RDP, and SILVA 16S rRNA gene reference databases. Between 82% and 97% of the genus annotations assigned to 16S rRNA gene reference sequences were deemed valid in the reference databases. Between 18% and 97% of the species annotations in Greengenes and SILVA were deemed valid. Misannotations included the use of metadata in place of taxonomic classifications, non-adherence to the binomial nomenclature, and sequences classified as eukaryote organelles or taxa.\n\nConclusionsThe misannotations identified in public 16S rRNA gene databases call into question the reliability of research made using these resources. As targeted gene surveys depend on high quality marker gene databases, imed nomenclature accuracy will be necessary.

microbiology

Density dependent enhancement effect of Wolbachia and the host RNAi response to a densovirus in Aedes cells

The endosymbiotic bacterium Wolbachia pipientis has been shown to restrict a range of RNA viruses in Drosophila melanogaster and transinfected dengue mosquito, Aedes aegypti. Here, we show that Wolbachia infection enhances replication of Aedes albopictus densovirus (AalDNV-1), a single stranded DNA virus, in Aedes cell lines in a density-dependent manner. Analysis of previously produced small RNAs of Aag2 cells showed that Wolbachia-infected cells produced greater proportions of viral derived short interfering RNAs as compared to uninfected cells. Additionally, we found production of viral derived PIWI-like RNAs (vpiRNA) produced in response to AalDNV-1 infection. Nuclear fractions of Aag2 cells produced a primary vpiRNA signature U1 bias whereas the typical \"ping-pong\" signature (U1 - A10) was evident in the cytoplasmic fraction. This is the first report of the density-dependent enhancement of DNA viruses by Wolbachia. Further, we report the generation of vpiRNAs in a DNA virus-host interaction for the first time.

microbiology