Search bioRxivSearch

SEARCH · Search bioRxiv

Search Search bioRxiv

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,117 records · Page 62Linked to original sources

Structure and genome ejection mechanism of Podoviridae phage P68 infecting Staphylococcus aureus

Phages infecting S. aureus have the potential to be used as therapeutics against antibiotic-resistant bacterial infections. However, there is limited information about the mechanism of genome delivery of phages that infect Gram-positive bacteria. Here we present the structures of S. aureus phage P68 in its native form, genome ejection intermediate, and empty particle. The P68 head contains seventy-two subunits of inner core protein, fifteen of which bind to and alter the structure of adjacent major capsid proteins and thus specify attachment sites for head fibers. Unlike in the previously studied phages, the head fibers of P68 enable its virion to position itself at the cell surface for genome delivery. P68 genome ejection is triggered by disruption of the interaction of one of the portal protein subunits with phage DNA. The inner core proteins are released together with the DNA and enable the translocation of phage genome across the bacterial membrane into the cytoplasm.

microbiology

Refacing: reconstructing anonymized facial features using GANs

Anonymization of medical images is necessary for protecting the identity of the test subjects, and is therefore an essential step in data sharing. However, recent developments in deep learning may raise the bar on the amount of distortion that needs to be applied to guarantee anonymity. To test such possibilities, we have applied the novel CycleGAN unsupervised image-to-image translation framework on sagittal slices of T1 MR images, in order to reconstruct facial features from anonymized data. We applied the CycleGAN framework on both face-blurred and face-removed images. Our results show that face blurring may not provide adequate protection against malicious attempts at identifying the subjects, while face removal provides more robust anonymization, but is still partially reversible.

neuroscience

RAxML-NG: A fast, scalable, and user-friendly tool for maximum likelihood phylogenetic inference

MotivationPhylogenies are important for fundamental biological research, but also have numerous applications in biotechnology, agriculture, and medicine. Finding the optimal tree under the popular maximum like-lihood (ML) criterion is known to be NP-hard. Thus, highly optimized and scalable codes are needed to analyze constantly growing empirical datasets.\n\nResultsWe present RAxML-NG, a from scratch re-implementation of the established greedy tree search algorithm of RAxML/ExaML. RAxML- NG offers improved accuracy, flexibility, speed, scalability, and usability compared to RAxML/ExaML. On taxon-rich datasets, RAxML-NG typically finds higher-scoring trees than IQTree, an increasingly popular recent tool for ML-based phylogenetic inference (although IQ-Tree shows better stability). Finally, RAxML-NG introduces several new features, such as the detection of terraces in tree space and a the recently introduced transfer bootstrap support metric.\n\nAvailabilityThe code is available under GNU GPL at https://github.com/amkozlov/raxml-ng.RAxML-NG web service (maintained by Vital- IT) is available at https://raxml-ng.vital-it.ch/.\n\nContactalexey.kozlov@h-its.org

bioinformatics

Vigour/tolerance trade-off in cultivated sunflower (Helianthus annuus) response to salinity stress is linked to leaf elemental composition

Developing more stress-tolerant crops will require greater knowledge of the physiological basis of stress tolerance. Here we explore how the variation among twenty cultivated sunflower (Helianthus annuus) genotypes for biomass decline in response to increasing salinity relates to leaf traits and leaf trait adjustments. Genotypes were grown in the greenhouse under five salinity treatments (0, 50, 100, 150, or 200 mM NaCl) for 21 days and assessed for growth, leaf physiological traits, and leaf elemental composition. Results showed that there was a trade-off in performance such that vigorous genotypes, higher biomass at zero mM NaCl, had both a larger absolute decrease and proportional decrease in biomass due to increased salinity. Contrary to expectation, genotypes with a low increase in leaf Na+ and Na+:K+ were no better at maintaining biomass with increasing salinity. Rather, genotypes with a greater reduction in leaf S and K+ content were better at maintaining biomass in the face of increasing salinity. While we found a trade-off between vigour and tolerance, some genotypes were more tolerant than expected. Further analysis of the traits underlying this trade-off will allow us to identify traits/mechanisms that could be bred into high vigour genotypes in order to increase their tolerance.

plant biology

Evolution of empathetic moral evaluation

Social norms can promote cooperation in human societies by assigning reputations to individuals based on their past actions. A good reputation indicates that an individual is worthy of help and is likely to reciprocate. A large body of research has established the norms of moral assessment that promote cooperation and maximize social welfare, assuming reputations are objective. But if there is no centralized institution to provide objective moral evaluation, then opinions about an individuals reputation may differ across a population. Here we use evolutionary game theory to study the effects of empathy - the capacity to make moral evaluations from the perspective of another person. We find that empathetic moral evaluation tends to foster cooperation by reducing the rate of unjustified defection. The norms of moral evaluation previously considered most socially beneficial depend on high levels of empathy, whereas different norms are required to maximize social welfare in populations unwilling or incapable of empathy. We demonstrate that empathy itself can evolve through social contagion and attain evolutionary stability under most social norms. We conclude that a capacity for empathetic moral evaluation represents a key component to sustaining cooperation in human societies: cooperation requires getting into the mindset of others whose views differ from our own.

evolutionary biology

The molecular genetics of hand preference revisited

Hand preference is a prominent behavioural trait linked to human brain asymmetry. A handful of genetic variants have been reported to associate with hand preference or quantitative measures related to it. Most of these reports were on the basis of limited sample sizes, by current standards for genetic analysis of complex traits. Here we performed a genome-wide association analysis of hand preference in the large, population-based UK Biobank cohort (N=331,037). We used gene-set enrichment analysis to investigate whether genes involved in visceral asymmetry are particularly relevant to hand preference, following one previous report. We found no evidence implicating any specific candidate variants previously reported. We also found no evidence that genes involved in visceral laterality play a role in hand preference. It remains possible that some of the previously reported genes or pathways are relevant to hand preference as assessed in other ways, or else are relevant within specific disorder populations. However, some or all of the earlier findings are likely to be false positives, and none of them appear relevant to hand preference as defined categorically in the general population. Within the UK Biobank itself, a significant association implicates the gene MAP2 in handedness.

genetics

Standardization and validation of a panel of cross-species microsatellites to individually identify the Asiatic wild dog (Cuon alpinus): implications in population estimation and dynamics

BackgroundThe Asiatic wild dog or dhole (Cuon alpinus) is a highly elusive, monophyletic, forest dwelling, social canid distributed across south and Southeast Asia. Severe pressures from habitat loss, prey depletion, disease, human persecution and interspecific competition resulted in global population decline in dholes. Despite a declining population trend, detailed information on population size, ecology, demography and genetics is lacking. Generating reliable information and landscape level for dholes is challenging due to their secretive behaviour and monomorphic physical features. Recent advances in non-invasive DNA-based tools can be used to monitor populations and individuals across large landscapes. In this paper, we describe standardization and validation of faecal DNA-based methods for individual identification of dholes. We tested this method on field-collected dhole faeces in four tiger reserves of the central Indian landscape in the state of Maharashtra, India. Further, we conducted preliminary analyses of dhole population structure and demography in the study area.\n\nResultsWe tested a total of 18 cross-species markers and developed a panel of 12 markers for unambiguous individual identification of dholes. This marker panel identified 101 unique individuals from faecal samples collected across our pilot field study area. These loci showed varied level of amplification success (57-88%), polymorphism (3-9 alleles), heterozygosity (0.23-0.63) and produced a cumulative probability of identity (unbiased) and probability of identity (sibs) value of 4.7x10-10 and 1.5x10-4, respectively. Our preliminary analyses of population structure indicated four genetic subpopulations in dholes. Qualitative analyses of population demography show signal of population decline.\n\nConclusionOur results demonstrated that the selected panel of 12 microsatellite loci can conclusively identify dholes from poor quality, non-invasive biological samples and help in exploring various population parameters. Our methods can be used to estimate dhole populations and assess population trends for this elusive, social carnivore.

genetics

MITRE: predicting host status from microbiota time-series data

Longitudinal studies are crucial for discovering casual relationships between the microbiome and human disease. We present Microbiome Interpretable Temporal Rule Engine (MITRE), the first machine learning method specifically designed for predicting host status from microbiome time-series data. Our method maintains interpretability by learning predictive rules over automatically inferred time-periods and phylogenetically related microbes. We validate MITREs performance on semi-synthetic data, and five real datasets measuring microbiome composition over time in infant and adult cohorts. Our results demonstrate that MITRE performs on par or outperforms \"black box\" machine learning approaches, providing a powerful new tool enabling discovery of biologically interpretable relationships between microbiome and human host.

bioinformatics

A network module for the Perseus software for computational proteomics facilitates proteome interaction graph analysis

Proteomics data analysis strongly benefits from not studying single proteins in isolation but taking their multivariate interdependence into account. We introduce PerseusNet, the new Perseus network module for the biological analysis of proteomics data. Proteomics is commonly used to generate networks, e.g. with affinity purification experiments, but networks are also used to explore proteomics data. PerseusNet supports the biomedical researcher for both modes of data analysis with a multitude of activities. For affinity purification, a volcano plot-based statistical analysis method for network generation is featured which is scalable to large numbers of baits. For posttranslational modifications of proteins, such as phosphorylation, a collection of dedicated network analysis tools helps elucidating cellular signaling events. Co-expression network analysis of proteomics data adopts established tools from transcriptome co-expression analysis. PerseusNet is extensible through a plug-in architecture in a multi-lingual way, integrating analyses in C#, Python and R and is freely available at http://www.perseus-framework.org.

bioinformatics

Parallel systems for sound processing and functional connectivity among layer 5 and 6 auditory corticothalamic neurons

Cortical layers (L) 5 and 6 are populated by a spatially intermingled menagerie of neurons with distinct inputs and downstream targets. Here, we made optogenetically guided recordings from L5 and L6 corticothalamic (CT) neurons in the mouse auditory cortex to discern underlying patterns of functional connectivity and sensory processing in the largest sub-cerebral projection system. Whereas L5 CT neurons showed broad stimulus selectivity with sluggish response latencies and extended temporal non-linearities, L6 CTs exhibited sparse sound feature selectivity and rapid temporal processing. L5 CT spikes lagged behind neighboring units and imposed weak feedforward excitation within the local column. By contrast, L6 CT spikes drove robust and sustained activity in neighboring units. Our findings underscore a duality among CT projection neurons, where L5 CT units are canonical broadcast neurons that integrate sensory inputs for transmission to distributed downstream targets, while L6 CT neurons are positioned to regulate thalamocortical response gain and selectivity.

neuroscience

Conducting social network analysis with animal telemetry data: applications and methods using spatsoc

O_LIWe present spatsoc: an R package for conducting social network analysis with animal telemetry data.\nC_LIO_LIAnimal social network analysis is a method for measuring relationships between individuals to describe social structure. Using animal telemetry data for social network analysis requires functions to generate proximity-based social networks that have flexible temporal and spatial grouping. Data can be complex and relocation frequency can vary so the ability to provide specific temporal and spatial thresholds based on the characteristics of the species and system is required.\nC_LIO_LIspatsoc fills a gap in R packages by providing flexible functions, explicitly for animal telemetry data, to generate gambit-of-the-group data, perform data-stream randomization and generate group by individual matrices.\nC_LIO_LIThe implications of spatsoc are that current users of large animal telemetry or otherwise georeferenced data for movement or spatial analyses will have access to efficient and intuitive functions to generate social networks.\nC_LI

ecology

Quantitative Proteomic Analysis of Prostate Tissue Specimens Identifies Deregulated ProteinComplexes in Primary Prostate Cancer

Prostate cancer (PCa) is the most frequently diagnosed non-skin cancer and a leading cause of mortality among males in developed countries. However, our understanding of the global changes of protein complexes within PCa tissue specimens remains very limited, although it has been well recognized that protein complexes carry out essentially all major processes in living organisms and that their deregulation drives the pathogenesis and progression of various diseases. By coupling tandem mass tagging-synchronous precursor selection-mass spectrometry/mass spectrometry/mass spectrometry (TMT-SPSMS3) with differential expression and co-regulation analyses, the present study compared the differences between protein complexes in normal prostate, low-grade PCa, and high-grade PCa tissue specimens. Globally, a large downregulated putative protein-protein interaction (PPI) network was detected in both low-grade and high-grade PCa, yet a large upregulated putative PPI network was only detected in high-grade but not low-grade PCa, compared with normal controls. To identify specific protein complexes that are deregulated in PCa, quantified proteins were mapped to protein complexes in CORUM, a collection of experimentally verified mammalian protein complexes. Differential expression analysis suggested that mitochondrial ribosomes and the fibrillin-associated protein complex were significantly overexpressed, whereas the ITGA6-ITGB4-Laminin10/12 and the P2X7 receptor signaling complexes were significantly downregulated, in PCa compared with normal prostate. Moreover, differential co-regulation analysis indicated that the assembly levels of some nuclear protein complexes involved in RNA synthesis and processing were significantly increased in low-grade PCa, and those of mitochondrial complex I and its subcomplexes were significantly increased in high-grade PCa, compared with normal prostate. In summary, the study represents the first global and quantitative comparison of protein complexes in prostate tissue specimens. It is expected to enhance our understanding of the molecular mechanisms underlying PCa development and progression in human patients, as well as lead to the discovery of novel biomarkers and therapeutic targets for precision management of PCa.

cancer biology

Evolutionary couplings detect side-chain interactions

Patterns of amino acid covariation in large protein sequence alignments can inform the prediction of de novo protein structures, binding interfaces, and mutational effects. While algorithms that detect these so-called evolutionary couplings between residues have proven useful for practical applications, less is known about how and why these methods perform so well, and what insights into biological processes can be gained from their application. Evolutionary coupling algorithms are commonly benchmarked by comparison to true structural contacts derived from solved protein structures. However, the methods used to determine true structural contacts are not standardized and different definitions of structural contacts may have important consequences for interpreting the results from evolutionary coupling analyses and understanding their overall utility. Here, we show that evolutionary coupling analyses are significantly more likely to identify structural contacts between side-chain atoms than between backbone atoms. We use both simulations and empirical analyses to highlight that purely backbone-based definitions of true residue-residue contacts (i.e., based on the distance between C atoms) may underestimate the accuracy of evolutionary coupling algorithms by as much as 40% and that a commonly used reference point (C{beta} atoms) underestimates the accuracy by 10-15%. These findings show that co-evolutionary outcomes differ according to which atoms participate in residue-residue interactions and suggest that accounting for different interaction types may lead to further improvements to contact-prediction methods.\n\nSignificance StatementEvolutionary couplings between residues within a protein can provide valuable information about protein structures, protein-protein interactions, and the mutability of individual residues. However, the mechanistic factors that determine whether two residues will co-evolve remains unknown. We show that structural proximity by itself is not sufficient for co-evolution to occur between residues. Rather, evolutionary couplings between residues are specifically governed by interactions between side-chain atoms. By contrast, intramolecular contacts between atoms in the protein backbone display only a weak signature of evolutionary coupling. These findings highlight that different types of stabilizing contacts exist within protein structures and that these types have a differential impact on the evolution of protein structures that should be considered in co-evolutionary applications.

biophysics

Unsupervised Machine learning to subtype Sepsis-Associated Acute Kidney Injury

ObjectiveAcute kidney injury (AKI) is highly prevalent in critically ill patients with sepsis. Sepsis-associated AKI is a heterogeneous clinical entity, and, like many complex syndromes, is composed of distinct subtypes. We aimed to agnostically identify AKI subphenotypes using machine learning techniques and routinely collected data in electronic health records (EHRs).\n\nDesignCohort study utilizing the MIMIC-III Database.\n\nSettingICUs from tertiary care hospital in the U.S.\n\nPatientsPatients older than 18 years with sepsis and who developed AKI within 48 hours of ICU admission.\n\nInterventionsUnsupervised machine learning utilizing all available vital signs and laboratory measurements.\n\nMeasurements and Main ResultsWe identified 1,865 patients with sepsis-associated AKI. Ten vital signs and 691 unique laboratory results were identified. After data processing and feature selection, 59 features, of which 28 were measures of intra-patient variability, remained for inclusion into an unsupervised machine-learning algorithm. We utilized k-means clustering with k ranging from 2 - 10; k=2 had the highest silhouette score (0.62). Cluster 1 had 1,358 patients while Cluster 2 had 507 patients. There were no significant differences between clusters on age, race or gender. We found significant differences in comorbidities and small but significant differences in several laboratory variables (hematocrit, bicarbonate, albumin) and vital signs (systolic blood pressure and heart rate). In-hospital mortality was higher in cluster 2 patients, 25% vs. 20%, p=0.008. Features with the largest differences between clusters included variability in basophil and eosinophil counts, alanine aminotransferase levels and creatine kinase values.\n\nConclusionsUtilizing routinely collected laboratory variables and vital signs in the EHR, we were able to identify two distinct subphenotypes of sepsis-associated AKI with different outcomes. Variability in laboratory variables, as opposed to their actual value, was more important for determination of subphenotypes. Our findings show the potential utility of unsupervised machine learning to better subtype AKI.

bioinformatics

Optical Sectioning of Live Mammal with Near-Infrared Light Sheet

Deep-tissue three-dimensional optical imaging of live mammals in vivo with high spatiotemporal resolution in non-invasive manners has been challenging due to light scattering. Here, we developed near-infrared (NIR) light sheet microscopy (LSM) with optical excitation and emission wavelengths up to ~ 1320 nm and ~ 1700 nm respectively, far into the NIR-II (1000-1700 nm) region for 3D optical sectioning through live tissues. Suppressed scattering of both excitation and emission photons allowed one-photon optical sectioning at ~ 2 mm depth in highly scattering brain tissues. NIR-II LSM enabled non-invasive in vivo imaging of live mice, revealing never-before-seen dynamic processes such as highly abnormal tumor microcirculation, and 3D molecular imaging of an important immune checkpoint protein, programmed-death ligand 1 (PD-L1) receptors at the single cell scale in tumors. In vivo two-color near-infrared light sheet sectioning enabled simultaneous volumetric imaging of tumor vasculatures and PD-L1 proteins in live mammals.

bioengineering

A stable, long-term cortical signature underlying consistent behavior

Animals readily execute learned motor behaviors in a consistent manner over long periods of time, yet similarly stable neural correlates remained elusive up to now. How does the cortex achieve this stable control? Using the sensorimotor system as a model of cortical processing, we investigated the hypothesis that the dynamics of neural latent activity, which capture the dominant co-variation patterns within the neural population, are preserved across time. We recorded from populations of neurons in premotor, primary motor, and somatosensory cortices for up to two years as monkeys performed a reaching task. Intriguingly, despite steady turnover in the recorded neurons, the low-dimensional latent dynamics remained stable. Such stability allowed reliable decoding of behavioral features for the entire timespan, while fixed decoders based on the recorded neural activity degraded substantially. We posit that latent cortical dynamics within the manifold are the fundamental and stable building blocks underlying consistent behavioral execution.

neuroscience

Putting species back on the map: devising a robust method for quantifying the biodiversity impacts of land conversion

AimQuantifying connections between the global drivers of habitat loss and biodiversity impact is vital for decision-makers promoting responsible land-use. To that end, biodiversity impact metrics should be able to report linked trends in specific anthropogenic activities and changes in biodiversity state. However, for biodiversity, it is challenging to deliver integrated information on its multiple dimensions (i.e. species richness, endemicity) and keep it practical. Here, we developed a biodiversity footprint indicator that can i) capture the status of different species groups, ii) link biodiversity impact to specific human activities, and iii) be adapted to the most applicable scale for the decision context.\n\nLocationCerrado Biome, Brazil\n\nMethodsWe illustrate this globally-applicable approach for the case of soybean expansion in the Brazilian Cerrado. Using species-specific habitat suitability models, we assessed the impact of soy expansion and other land uses over 2,000 species of amphibians, birds, mammals and plants for three time periods between 2000 and 2014.\n\nResultsOverall, plants suffered the greatest reduction of suitable habitat. However, among endemic and near-endemic species - which face greatest risk of global extinction from habitat conversion in the Cerrado - birds were the most affected group. While planted pastures and cropland expansion were together responsible for most of the absolute biodiversity footprint, soy expansion via direct conversion of natural vegetation had the greatest impact per unit area. The total biodiversity footprint over the period was concentrated in the southern states of Minas Gerais, Goias and Mato Grosso, but the soy footprint was proportionally higher in those northern states (such as Bahia and Piaui) which belong to the new agricultural frontier.\n\nMain conclusionsThe ability and flexibility of our approach to examine linkages between biodiversity loss and specific human activities has substantial potential to better characterise the pathways by which habitat loss drivers operate.

ecology

Epigenome-wide association analysis of daytime sleepiness in the Multi-Ethnic Study of Atherosclerosis reveals African-American specific associations

Study ObjectivesExcessive daytime sleepiness (EDS) is a consequence of inadequate sleep, or of a primary disorder of sleep-wake control. Population variability in prevalence of EDS and susceptibility to EDS are likely due to genetic and biological factors as well as social and environmental influences. Epigenetic modifications (such as DNA methylation-DNAm) are potential influences on a range of health outcomes. Here, we explored the association between DNAm and daytime sleepiness quantified by the Epworth Sleepiness Scale (ESS).\n\nMethodsWe performed multi-ethnic and ethnic-specific epigenome-wide association studies for DNAm and ESS in 619 individuals from the Multi-Ethnic Study of Atherosclerosis. Replication was assessed in the Cardiovascular Health Study (CHS). Genetic variants in genes proximal to ESS-associated DNAm were analyzed to identify methylation quantitative trait loci and followed with replication of genotype-sleepiness associations in the UK Biobank.\n\nResults61 methylation sites were associated with ESS (FDR [≤] 0.1) in African Americans only, including an association in KCTD5, a gene strongly implicated in sleep. One association (cg26130090) replicated in CHS African Americans (p-value 0.0004). We identified a sleepiness-associated methylation site in the gene RAI1, a gene associated with sleep and circadian phenotypes. In a follow-up analysis, a genetic variant within RAI1 associated with both DNAm and sleepiness score. The variants association with sleepiness was replicated in the UK Biobank.\n\nConclusionsOur analysis identified methylation sites in multiple genes that may be implicated in EDS. These sleepiness-methylation associations were specific to African Americans. Future work is needed to identify mechanisms driving ancestry-specific methylation effects.\n\nStatement of SignificanceExcessive daytime sleepiness is associated with negative health outcomes such as reduction in quality of life, increased workplace accidents, and cardiovascular mortality. There are race/ethnic disparities in excessive daytime sleepiness, however, the environmental and biological mechanisms for these differences are not yet understood. We performed an association analysis of DNA methylation, measured in monocytes, and daytime sleepiness within a racially diverse study population. We detected numerous DNA methylation markers associated with daytime sleepiness in African Americans, but not in European and Hispanic Americans. Future work is required to elucidate the pathways between DNA methylation, sleepiness, and related behavioral/environmental exposures.

genomics