Search bioRxivSearch

SEARCH · Search bioRxiv

Search Search bioRxiv

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

Measurement of the Topological Dimension of Hippocampal Place Cell Activity

The fundamental property of topological dimension of neural activity has never been measured before. We measured the topological dimension of the activity of 89 rat hippocampal place cells recorded during foraging using the neighborhood boundary countdown method. Points were rate vectors forming a manifold in an 89 dimensional rate space Since a boundary has one less dimension than the neighborhood it encloses, reducing the dimension in sequential radial cuts finally arrives at boundary that is two points. The number of cuts required is the topological dimension of the set. Due to sparsity, we used a shell-like boundary with thickness. As expected, we found the large inductive dimension of this set of points to be two. To examine the robustness of the method, we used both real and modeled place cells and varied the number of neurons, duration of the time step and smoothing, radius and thickness for the boundary, decrement of the radius with each cut, requirements for the center point, and number of points in the final cluster. With the exception of shell thickness, the result was insensitive to these variations. Knowing the topological dimension allows application of rigorous topological principles to the activity of neuron populations. The method potentially can be applied to other neural classes and other behaviors, to enumerate firing variables when they are unknown, even count the factors influencing neural activity in sleep.

neuroscience

A novel signature derived from immunoregulatory and hypoxia genes predicts prognosis in liver and five other cancers

BackgroundDespite much progress in cancer research, its incidence and mortality continue to rise. A robust biomarker that would predict tumor behavior is highly desirable and could improve patient treatment and prognosis.\n\nMethodsIn a retrospective bioinformatics analysis involving patients with liver cancer (n=839), we developed a prognostic signature consisting of 45 genes associated with tumor-infiltrating lymphocytes and cellular responses to hypoxia. From this gene set, we were able to identify a second prognostic signature comprised of 8 genes. Its performance was further validated in five other cancers: head and neck (n=520), renal papillary cell (n=290), lung (n=515), pancreas (n=178) and endometrial (n=370).\n\nFindingsThe 45-gene signature predicted overall survival in three liver cancer cohorts: hazard ratio (HR)=1.82, P=0.006; HR=1.84, P=0.008 and HR=2.67, P=0.003. Additionally, the reduced 8-gene signature was sufficient and effective in predicting survival in liver and five other cancers: liver (HR=2.36, P=0.0003; HR=2.43, P=0.0002 and HR=3.45, P=0.0007), head and neck (HR=1.64, P=0.004), renal papillary cell (HR=2.31, P=0.04), lung (HR=1.45, P=0.03), pancreas (HR=1.96, P=0.006) and endometrial (HR=2.33, P=0.003). Receiver operating characteristic analyses demonstrated both signatures superior performance over current tumor staging parameters. Multivariate Cox regression analyses revealed that both 45-gene and 8-gene signatures were independent of other clinicopathological features in these cancers. Combining the gene signatures with somatic mutation profiles increased their prognostic ability.\n\nConclusionsThis study, to our knowledge, is the first to identify a gene signature uniting both tumor hypoxia and lymphocytic infiltration as a prognostic determinant in six cancer types (n=2,712). The 8-gene signature can be used for patient risk stratification by incorporating hypoxia information to aid clinical decision making.

cancer biology

Prognostic signatures of oxygen-sensing genes predict patient survival in ten cancers including those of the liver, pancreas and stomach

ObjectivesTumor hypoxia is associated with metastasis and resistance to chemotherapy and radiotherapy. Genes involved in oxygen-sensing are clinically relevant and have significant implications on prognosis.\n\nMethodsWe identified of two signatures, signature 1 (good prognosis) and signature 2 (adverse prognosis), each consisting of 5 genes using three pancreatic cancer cohorts (n=681). We validated the signatures performance in predicting survival in ten cancers using Cox regression and receiver operating characteristic (ROC) analyses.\n\nResultsSignature 1 and signature 2 were associated with good and poor overall survival respectively. Prognosis of signature 1 in 8 cohorts representing 6 cancers (n=2,627): bladder (hazard ratio [HR]=0.68, P=0.039), papillary renal cell (HR=0.35, P=0.013), liver (HR=0.64, P=0.033 and HR=0.49, P=0.025), lung (HR=0.66, P=0.014) and pancreatic (HR=0.42, P<0.001 and HR=0.64, P=0.04) and endometrial (HR=0.40, P<0.001). Prognosis of signature 2 in 12 cohorts representing 9 cancers (n=4,134): bladder (HR=1.46, P=0.039), cervical (HR=1.97, P=0.035), head and neck (HR=1.39, P=0.038), renal clear cell (HR=1.47, P=0.012), papillary renal cell (HR=3.89, P=0.0015), liver (HR=5.10, P<0.0001 and HR=2.26, P<0.001), lung (HR=1.54, P=0.011), pancreatic (HR=2.09, P=0.002, HR=1.46, P=0.018, and HR=1.99, P<0.0001) and stomach (HR=1.78, P=0.004). Multivariate Cox regression confirmed independent clinical relevance of signatures in these cancers. ROC analyses confirmed superior performance of signatures to current tumor staging benchmarks. KDM8 is a potential tumor suppressor downregulated in liver and pancreatic cancers and is an independent prognostic factor. KDM8 expression negatively correlated with cell cycle regulators. Low KDM8 in tumors was associated with loss of cell adhesion phenotype through HNF4A signaling.\n\nConclusionsPan-cancer signatures of oxygen-sensing genes used for risk assessment in 10 cancers (n=6,761) could guide individualized treatment plans.

cancer biology

Synaptic circuits for irradiance coding by intrinsically photosensitive retinal ganglion cells

We have explored the synaptic networks responsible for the unique capacity of intrinsically photosensitive retinal ganglion cells (ipRGCs) to encode overall light intensity. This luminance signal is crucial for circadian, pupillary and related reflexive responses light. By combined glutamate-sensor imaging and patch recording of postsynaptic RGCs, we show that the capacity for intensity-encoding is widespread among cone bipolar types, including OFF types.\n\nNonetheless, the bipolar cells that drive ipRGCs appear to carry the strongest luminance signal. By serial electron microscopic reconstruction, we show that Type 6 ON cone bipolar cells are the dominant source of such input, with more modest input from Types 7, 8 and 9 and virtually none from Types 5i, 5o, 5t or rod bipolar cells. In conventional RGCs, the excitatory drive from bipolar cells is high-pass temporally filtered more than it is in ipRGCs. Amacrine-to-bipolar cell feedback seems to contribute surprisingly little to this filtering, implicating mostly postsynaptic mechanisms. Most ipRGCs sample from all bipolar terminals costratifying with their dendrites, but M1 cells avoid all OFF bipolar input and accept only ectopic ribbon synapses from ON cone bipolar axonal shafts. These are remarkable monad synapses, equipped with as many as a dozen ribbons and only one postsynaptic process.

neuroscience

Asi1 regulates the distribution of proteins at the inner nuclear membrane in Saccharomyces cerevisiae

Inner nuclear membrane (INM) protein composition regulates nuclear function, affecting processes such as gene expression, chromosome organization, nuclear shape and stability. Mechanisms that drive changes in the INM proteome are poorly understood in part because it is difficult to definitively assay INM composition rigorously and systematically. Using a split-GFP complementation system to detect INM access, we examined the distribution of all C-terminally tagged Saccharomyces cerevisiae membrane proteins in wild-type cells and in mutants affecting protein quality control pathways, such as INM-associated degradation (INMAD), ER-associated degradation (ERAD) and vacuolar proteolysis. Deletion of the E3 ligase Asi1 had the most pronounced effect on the INM compared to mutants in vacuolar or ER-associated degradation pathways, consistent with a role for Asi1 in the INMAD pathway. Our data suggests that Asi1 not only removes mis-targeted proteins at the INM, but it also controls the levels and distribution of native INM components, such as the membrane nucleoporin Pom33. Interestingly, loss of Asi1 does not affect Pom33 protein levels but instead alters Pom33 distribution in the NE through Pom33 ubiquitination, which drives INM redistribution. Taken together, our data demonstrate that the Asi1 E3 ligase has a novel function in INM protein regulation in addition to protein turnover.

cell biology

The influenza A virus endoribonuclease PA-X usurps host mRNA processing machinery to limit host gene expression

Many viruses globally shut off host gene expression to inhibit activation of cell-intrinsic antiviral responses. However, host shutoff is not indiscriminate, since viral proteins and host proteins required for viral replication are still synthesized during shutoff. The molecular determinants of target selectivity in host shutoff remain incompletely understood. Here, we report that the influenza A virus shutoff factor PA-X usurps RNA splicing to selectively target host RNAs for destruction. PA-X preferentially degrades spliced mRNAs, both transcriptome-wide and in reporter assays. Moreover, proximity-labeling proteomics revealed that PA-X interacts with cellular proteins involved in RNA splicing. The interaction with splicing contributes to target discrimination and is unique among viral host shutoff nucleases. This novel mechanism sheds light on the specificity of viral control of host gene expression and may provide opportunities for development of new host-targeted antivirals.

microbiology

ADAR2-mediated Q/R editing of GluK2 regulates kainate receptor upscaling in response to suppression of synaptic activity

Kainate receptors (KARs) regulate neuronal excitability and network function. Most KARs contain the subunit GluK2 and the properties of these receptors are determined in part by ADAR2-mediated mRNA editing of GluK2 that changes a genomically encoded glutamine (Q) to arginine (R). Suppression of synaptic activity reduces ADAR2-dependent Q/R editing of GluK2 with a consequential increase in GluK2-containing KAR surface expression. However, the mechanism underlying this reduction in GluK2 editing has not been addressed. Here we show that induction of KAR upscaling results in proteasomal degradation of ADAR2, which reduces GluK2 Q/R editing. Because KARs incorporating unedited GluK2(Q) assemble and exit the ER more efficiently this leads to an upscaling of KAR surface expression. Consistent with this, we demonstrate that partial ADAR2 knockdown phenocopies and occludes KAR upscaling. Moreover, we show that although the AMPAR subunit GluA2 also undergoes ADAR2-dependent Q/R editing, this process does not mediate AMPAR upscaling. These data demonstrate that activity-dependent regulation of ADAR2 proteostasis and GluK2 Q/R editing are key determinants of KAR, but not AMPAR, trafficking and upscaling.\n\nSummary statementSynaptic suppression promotes proteasomal degradation of the mRNA-editing enzyme ADAR2. Decreased ADAR2 levels reduce Q/R editing of the kainate receptor subunit GluK2 leading to enhanced surface expression and homeostatic upscaling.

neuroscience

The whale shark genome reveals how genomic and physiological properties scale with body size

The endangered whale shark (Rhincodon typus) is the largest fish on Earth and is a long-lived member of the ancient Elasmobranchii clade. To characterize the relationship between genome features and biological traits, we sequenced and assembled the genome of the whale shark and compared its genomic and physiological features to those of 81 animals and yeast. We examined scaling relationships between body size, temperature, metabolic rates, and genomic features and found both general correlations across the animal kingdom and features specific to the whale shark genome. Among animals, increased lifespan is positively correlated to body size and metabolic rate. Several genomic features also significantly correlated with body size, including intron and gene length. Our large-scale comparative genomic analysis uncovered general features of metazoan genome architecture: GC content and codon adaptation index are negatively correlated, and neural connectivity genes are longer than average genes in most genomes. Focusing on the whale shark genome, we identified multiple features that significantly correlate with lifespan. Among these were very long gene length, due to large introns highly enriched in repetitive elements such as CR1-like LINEs, and considerably longer neural genes of several types, including connectivity, activity, and neurodegeneration genes. The whale sharks genome had an expansion of gene families related to fatty acid metabolism and neurogenesis, with the slowest evolutionary rate observed in vertebrates to date. Our comparative genomics approach uncovered multiple genetic features associated with body size, metabolic rate, and lifespan, and showed that the whale shark is a promising model for studies of neural architecture and lifespan.

genomics

A robust nonlinear low-dimensional manifold for single cell RNA-seq data

Modern developments in single cell sequencing technologies enable broad insights into cellular state. Single cell RNA sequencing (scRNA-seq) can be used to explore cell types, states, and developmental trajectories to broaden understanding of cell heterogeneity in tissues and organs. Analysis of these sparse, high-dimensional experimental results requires dimension reduction. Several methods have been developed to estimate low-dimensional embeddings for filtered and normalized single cell data. However, methods have yet to be developed for unfiltered and unnormalized count data. We present a nonlinear latent variable model with robust, heavy-tailed error and adaptive kernel learning to estimate low-dimensional nonlinear structure in scRNA-seq data. Gene expression in a single cell is modeled as a noisy draw from a Gaussian process in high dimensions from low-dimensional latent positions. This model is called the Gaussian process latent variable model (GPLVM). We model residual errors with a heavy-tailed Students t-distribution to estimate a manifold that is robust to technical and biological noise. We compare our approach to common dimension reduction tools to highlight our models ability to enable important downstream tasks, including clustering and inferring cell developmental trajectories, on available experimental data. We show that our robust nonlinear manifold is well suited for raw, unfiltered gene counts from high throughput sequencing technologies for visualization and exploration of cell states.

genomics

On the unfounded enthusiasm for soft selective sweeps II: examining recent evidence from humans, flies, and viruses

Since the initial description of the genomic patterns expected under models of positive selection acting on standing genetic variation and on multiple beneficial mutations--so-called soft selective sweeps--researchers have sought to identify these patterns in natural population data. Indeed, over the past two years, large-scale data analyses have argued that soft sweeps are pervasive across organisms of very different effective population size and mutation rate--humans, Drosophila, and HIV. Yet, others have evaluated the relevance of these models to natural populations, as well as the identifiability of the models relative to other known population-level processes, arguing that soft sweeps are likely to be rare. Here, we look to reconcile these opposing results by carefully evaluating three recent studies and their underlying methodologies. Using population genetic theory, as well as extensive simulation, we find that all three examples are prone to extremely high false-positive rates, incorrectly identifying soft sweeps under both hard sweep and neutral models. Furthermore, we demonstrate that well-fit demographic histories combined with rare hard sweeps serve as the more parsimonious explanation. These findings represent a necessary response to the growing tendency of invoking parameter-heavy, assumption-laden models of pervasive positive selection, and neglecting best practices regarding the construction of proper demographic null models.

evolutionary biology

Hippocampal State Transitions at the Boundaries between Trial Epochs

The hippocampus encodes distinct environmental and behavioral contexts with unique patterns of activity. Representational shifts with changes in the context, referred to as remapping, have been extensively studied. However, less is known about the nature of transitions between representations. In this study, we leverage a large dataset of 2056 neurons recorded while rats performed an olfactory memory task with a predictable temporal structure involving trials and inter-trial intervals, separated by salient boundaries at the trial start and trial end. We found that trial epochs were associated with stable hippocampal population representations, despite moment to moment variability in stimuli and behavior. Representations of trial and inter-trial interval epochs were far more distinct than spatial factors would predict and the transitions between the two were abrupt, with a sharp boundary suggestive of a dynamic shift in the representational state. This boundary was associated with a large spike in multi-unit activity, with many individual cells specifically active at the start or end of each trial. Both epochs and boundaries were encoded by hippocampal populations, and these representations carried information on orthogonal axes readily identified using principal component analysis. We suggest that the activity spike at trial boundaries might serve to drive hippocampal activity from one stable state to another, and may play a role in segmenting continuous experience into discrete episodic memories.

neuroscience

A framework for space-efficient variable-order Markov models

MotivationMarkov models with contexts of variable length are widely used in bioinformatics for representing sets of sequences with similar biological properties. When models contain many long contexts, existing implementations are either unable to handle genome-scale training datasets within typical memory budgets, or they are optimized for specific model variants and are thus inflexible.\n\nResultsWe provide practical, versatile representations of variable-order Markov models and of interpolated Markov models, that support a large number of context-selection criteria, scoring functions, probability smoothing methods, and interpolations, and that take up to 4 times less space than previous implementations based on the suffix array, regardless of the number and length of contexts, and up to 10 times less space than previous trie-based representations, or more, while matching the size of related, state-of-the-art data structures from Natural Language Processing. We describe how to further compress our indexes to a quantity related to the redundancy of the training data, saving up to 90% of their space on repetitive datasets, and making them become up to 60 times smaller than previous implementations based on the suffix array. Finally, we show how to exploit constraints on the length and frequency of contexts to further shrink our compressed indexes to half of their size or more, achieving data structures that are 100 times smaller than previous implementations based on the suffix array, or more. This allows variable-order Markov models to be trained on bigger datasets and with longer contexts on the same hardware, thus possibly enabling new applications.\n\nAvailability and implementationhttps://github.com/jnalanko/VOMM

bioinformatics

The Cuon Enigma: Genome survey and comparative genomics of the endangered Dhole (Cuon alpinus)

The Asiatic wild dog is an endangered monophyletic canid restricted to Asia; facing threats from habitat fragmentation and other anthropogenic factors. Dholes have unique adaptations as compared to other wolf-like canids for large litter size (larger number of mammae) and hypercarnivory making it evolutionarily notable. Over evolutionary time, dhole and the subsequent divergent wild canids have lost coat patterns found in African wild dog. Here we report the first high coverage genome survey of Asiatic wild dog and mapped it with African wild dog, dingo and domestic dog to assess the structural variants. We generated a total of 124.8 Gb data from 416140921 raw read pairs and retained 398659457 reads with 52X coverage and mapped 99.16% of the clean reads to the three reference genomes. We identified ~13553269 SNVs, ~2858184 InDels, ~41000 SVs, ~1854109 SSRs and about 1000 CNVs. We compared the annotated genome of dingo and domestic dog with dhole genome sequence to understand the role of genes responsible in pelage pattern, dentition and mammary glands. Positively selected genes for these phenotypes were looked for SNP variants and top ranked genes for coat pattern, dentition and mammary glands were found to play a role in signalling and developmental pathways. Mitochondrial genome assembly predicted 35 genes, 11 CDS and 24 tRNA. This genome information will help in understanding the divergence of two monophlyletic canids, Cuon and Lycaon, and the evolutionary adaptations of dholes with respect to other canids.

genomics

Screening performance of abbreviated versions of the UPSIT smell test

BackgroundHyposmia features in several neurodegenerative conditions, including Parkinsons disease (PD). The University of Pennsylvania Smell Identification Test (UPSIT) is a widely used screening tool for detecting hyposmia, but is time-consuming and expensive when used on a large scale.\n\nMethodsWe assessed shorter subsets of UPSIT items for their ability to detect hyposmia in 891 healthy participants from the PREDICT-PD study. Established shorter tests included Versions A and B of both the 4-item Pocket Smell Test (PST) and 12-item Brief Smell Identification Test (BSIT). Using a data-driven approach, we evaluated screening performances of 23,231,378 combinations of 1-7 smell items from the full UPSIT.\n\nResultsPST Versions A and B achieved sensitivity/specificity of 76.8%/64.9% and 86.6%/45.9% respectively, whilst BSIT Versions A and B achieved 83.1%/79.5% and 96.5%/51.8% for detecting hyposmia defined by the longer UPSIT. From the data-driven analysis, two optimised sets of 7 smells surpassed the screening performance of the 12 item BSITs (with validation sensitivity/specificities of 88.2%/85.4% and 100%/53.5%). A set of 4 smells (Menthol, Clove, Gingerbread and Orange) had higher sensitivity for hyposmia than PST-A, -B and even BSIT-A (with validation sensitivity 91.2%). The same 4 smells also featured amongst those most commonly misidentified by 44 individuals with PD compared to 891 PREDICT-PD controls and a screening test using these 4 smells would have identified all hyposmic patients with PD.\n\nConclusionUsing abbreviated smell tests could provide a cost-effective means of screening for hyposmia in large cohorts, allowing more targeted administration of the UPSIT or similar smell tests.

neuroscience

A multi-parent recombinant inbred line population of Caenorhabditis elegans enhances mapping resolution and identification of novel QTLs for complex life-history traits

Local populations of the bacterivorous nematode Caenorhabditis elegans can be genetically almost as diverse as global populations. To investigate the effect of local genetic variation on heritable traits, we developed a new recombinant inbred line (RIL) population derived from four wild isolates. The wild isolates were collected from two closely located sites in France: Orsay and Santeuil. By crossing these four genetically diverse parental isolates a population of 200 RILs was constructed. RNA-seq was used to obtain sequence polymorphisms identifying almost 9000 SNPs variable between the four genotypes with an average spacing of 11 kb, possibly doubling the mapping resolution relative to currently available RIL panels. The SNPs were used to construct a genetic map to facilitate QTL analysis. Life history traits, such as lifespan, stress resistance, developmental speed and population growth were measured in different environments. For most traits substantial variation was found, and multiple QTLs could be detected, including novel QTLs not found in previous QTL analysis, for example for lifespan or pathogen responses. This shows that recombining genetic variation across C. elegans populations that are in geographical close proximity provides ample variation for QTL mapping. Taken together, we show that RNA-seq can be used for genotyping, that using more parents than the classical two parental genotypes to construct a RIL population facilitates the detection of QTLs and that the use of wild isolates permits analysis of local adaptation and life history trade-offs.

genetics

Mutually exclusive locales for N-linked glycans and disorder in glycoproteins

Several post-translational modifications of proteins lie within regions of disorder, stretches of amino acid residues that exhibit a dynamic tertiary structure and resist crystallization. Such localization has been proposed to expand the binding versatility of the disordered regions, and hence, the repertoire of interacting partners for the proteins. However, investigating a dataset of 500 human N-linked glycoproteins, we observed that the sites of N-linked glycosylations, or N-glycosites, lay predominantly within the regions of predicted order rather than their unstructured counterparts. This mutual exclusivity between disordered stretches and N-glycosites could not be reconciled merely through asymmetry in distribution of asparagines, serines or threonines residues, which comprise the minimum-required signature for conjugation by N-linked glycans, but rather by a contextual enrichment of these residues next to each other within the ordered portions. In fact, N-glycosite neighborhoods and disordered stretches showed distinct sets of enriched residues suggesting their individualized roles in protein phenotype. N-glycosite neighborhood residues also showed higher phylogenetic conservation than disordered stretches within amniote orthologs of glycoproteins. However, a universal search for residue-combinations that are putatively domain-constitutive ranked the disordered regions higher than the N-glycosite neighborhoods. We propose that amino acid residue-combinations bias the permissivity for N-glycoconjugation within ordered regions, so as to balance the tradeoff between the evolution of protein stability, and function, contributed by the N-linked glycans and disordered regions respectively.

bioinformatics

Dynamics of spontaneous alpha activity correlate with language ability in young children

Early childhood is a period of tremendous growth in both language ability and brain maturation. To understand the dynamic interplay between neural activity and spoken language development, we used resting-state EEG recordings to explore the relation between alpha oscillations (7-10 Hz) and oral language ability in 4- to 6-year-old children with typical development (N=41). Three properties of alpha oscillations were investigated: a) alpha power using spectral analysis, b) flexibility of the alpha frequency quantified via the oscillation's moment-to-moment fluctuations, and c) scaling behavior of the alpha oscillator investigated via the long-range temporal correlation in the alpha-amplitude time course. All three properties of the alpha oscillator correlated with children's oral language abilities. Higher language scores were correlated with lower alpha power, greater flexibility of the alpha frequency, and longer temporal correlations in the alpha-amplitude time course. Our findings demonstrate a cognitive role of several properties of the alpha oscillator that has largely been overlooked in the literature.

neuroscience

Mis-perception of motion in depth originates from an incomplete transformation of retinal signals

Depth perception requires the use of an internal model of the eye-head geometry to infer distance from binocular retinal images and extraretinal 3D eye-head information, particularly ocular vergence. Similarly for motion in depth perception, gaze angle is required to correctly interpret the spatial direction of motion from retinal images; however, it is unknown whether the brain can make adequate use of extraretinal version and vergence information to correctly interpret binocular retinal motion for spatial motion in depth perception. Here, we tested this by asking participants to reproduce the perceived spatial trajectory of an isolated point stimulus moving on different horizontal-depth paths either peri-foveally or peripherally while participants gaze was oriented at different vergence and version angles. We found large systematic errors in the perceived motion trajectory that reflected an intermediate reference frame between a purely retinal interpretation of binocular retinal motion (ignoring vergence and version) and the spatially correct motion. A simple geometric model could capture the behavior well, revealing that participants tended to underestimate their version by as much as 17%, overestimate their vergence by as much as 22%, and underestimate the overall change in retinal disparity by as much as 64%. Since such large perceptual errors are not observed in everyday viewing, we suggest that other monocular and/or contextual cues are required for accurate real-world motion in depth perception.

neuroscience