Search bioRxivSearch

Biology subjects

Browse preprints

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

At least 1,009 records · Page 56Linked to original sources

CRISPR/Cas9 screens Reveal Dasatinib Targets of Inhibiting T cell Activation and Proliferation

Immune response by T cells is essential for a healthy body against cancer, infection, and pathophysiological alteration. The activation and expansion of T cells can be inhibited by dasatinib, a tyrosine inhibitor, thus improving the outcome of diseases, such as autoimmune disease, graft-versus-host disease, and transplant rejection. The underlying mechanism of inhibition by dasatinib is elusive. Here, we designed and synthesized a CRISPR/Cas9 screening library that includes 6,149 genes. Using the library, we performed dasatinib CRISPR/cas9 screening in Jurkat cell, a T lymphocyte cell. We firstly identified survival essential genes for Jurkat cells. Comparing with other CRISPR/Cas9 screenings, we obtained Jurkat cell specific essential genes. By comparing dasatinib treatment to control, we identified a set of dasatinib targets, which includes known targets: CSK, LCK, ZAP70, and previously unknown targets: ZFP36L2, LRPPRC, CFLAR, PD-1, CD45 et al. Visualizing these target genes on T cell receptor signaling pathway, we found several genes could be inhibited by dasatinib. Furthermore, we introduced a framework, 9-square, to classify genes and found a group of genes that are associated with dasatinib resistance, possibly linking the side effects of dasatinib. These data reveal a set of dasatinib targets and demonstrate the molecular potential functions of dasatinib. Identification of dasatinib targets will broaden our understanding to its molecular mechanism, and thus benefits to clinical outcome.

cancer biology

Nanoscale robots exhibiting quorum sensing

Multi-agent systems demonstrate the ability to collectively perform complex tasks--e.g., construction1-2, search3, and locomotion4,5--with greater speed, efficiency, or effectiveness than could a single agent alone. Direct and indirect coordination methods allow agents to collaborate to share information and adapt their activity to fit dynamic situations. A well-studied example is quorum sensing (QS), a mechanism allowing bacterial communities to coordinate and optimize various phenotypes in response to population density. Here we implement, for the first time, bio-inspired QS in robots fabricated from DNA origami, which communicate by transmitting and receiving diffusing signals. The mechanism we describe includes features such as programmable response thresholds and quorum quenching, and is capable of being triggered by proximity of a specific target cell. Nanoscale robots with swarm intelligence could carry out tasks that have been so far unachievable in diverse fields such as industry, manufacturing and medicine.

synthetic biology

Model selection for biological crystallography

Structural biologists have fit increasingly complex model types to protein X-ray crystallographic data, motivated by higher-resolving crystals, greater computational power, and a growing appreciation for protein dynamics. Once fit, a more complex model will generally fit the experimental data better, but it also provides greater capacity to overfit to experimental noise. While refinement progress is normally monitored for a given model type with a fixed number of parameters, comparatively little attention has been paid to the selection among distinct model types where the number of parameters can vary. Using metrics derived in the statistical field of model comparison, we develop a framework for statistically rigorous inference of model complexity. From analysis of simulated data, we find that the resulting information criteria are less likely to prefer an erroneously complex model type and are less sensitive to noise, compared to the crystallographic cross-validation criterion Rfree. Moreover, these information criteria suggest caution in using complex model types and for inferring protein conformational heterogeneity from experimental scattering data.

biophysics

A piggyBac-based toolkit for inducible genome editing in mammalian cells

We describe the development and application of a novel series of vectors that facilitate CRISPR-Cas9-mediated genome editing in mammalian cells, which we call CRISPR-Bac. CRISPR-Bac leverages the piggyBac transposon to randomly insert CRISPR-Cas9 components into mammalian genomes. In CRISPR-Bac, a single piggyBac cargo vector containing a doxycycline-inducible Cas9 or catalytically-dead Cas9 (dCas9) variant and a gene conferring resistance to Hygromycin B is co-transfected with a plasmid expressing the piggyBac transposase. A second cargo vector, expressing a single-guide RNA (sgRNA) of interest, the reverse-tetracycline TransActivator (rtTA), and a gene conferring resistance to G418, is also cotransfected. Subsequent selection on Hygromycin B and G418 generates polyclonal cell populations that stably express Cas9, rtTA, and the sgRNA(s) of interest. Using Mus musculus-derived embryonic and trophoblast stem cells, we show that CRISPR-Bac can be used to knockdown proteins of interest, to create targeted genetic deletions with high efficiency, and to activate or repress transcription of protein-coding genes and an imprinted long noncoding RNA. The ratio of sgRNA-to-Cas9-to-transposase can be adjusted in transfections to alter the average number of cargo insertions into the genome. sgRNAs targeting multiple genes can be inserted in a single transfection. CRISPR-Bac is a versatile platform for genome editing that simplifies the generation of mammalian cells that stably express the CRISPR-Cas9 machinery.

molecular biology

Cell type-specific role of lamin-B1 and its inflammation-driven reduction in organ building and aging

Cellular architectural proteins often participate in organ development and maintenance. Although functional decay of some of these proteins during aging is known, the cell-type specific developmental role and the cause and consequence of their subsequent decay remain to be established especially in mammals. By studying lamins, the nuclear structural proteins, we demonstrate that lamin-B1 functions specifically in the thymic epithelial cells (TECs) for proper thymus organogenesis. An upregulation of proinflammatory cytokines in the intra-thymic myeloid immune cells during aging accompanies a gradual reduction of adult TEC lamins-B1. These cytokines cause adult TEC senescence and lamin-B1 reduction. We identify 17 adult TEC subsets and show that TEC lamin-B1 maintains the composition of these TECs. Lamin-B1 supports the expression of TEC genes needed for maintaining adult thymic architecture and function. Thus, structural proteins involved in organ building and maintenance can undergo inflammation-driven decay which can in turn contribute to age-associated organ degeneration.

immunology

Microbial predator-prey interactions could favor coincidental selection of diverse virulence factors in marine coastal waters.

Vibrios are ubiquitous in marine environments and opportunistically colonize a broad range of hosts. Strains of Vibrio tasmaniensis present in oyster farms can thrive in oysters during juvenile mortality events. Among them, V. tasmaniensis LGP32 behaves as a facultative intracellular pathogen of oyster hemocytes, a property rather unusual in vibrios. Herein, we asked whether LGP32 resistance to phagocytosis could result from coincidental selection of virulence factors during interactions with heterotrophic protists, such as amoeba, in the environment. To answer that question, we developed an integrative study, from the first description of amoeba diversity in oyster-farming areas to the characterization of LGP32 interactions with amoebae of the Vannella genus that were found abundant in the oyster environment. LGP32 was shown to be resistant to grazing by amoebae and this phenotype was dependent on previously identified virulence factors: the secreted metalloprotease Vsm and the copper efflux p-ATPase CopA. Using dedicated in vitro assays, our results showed that these virulence factors act at different steps during amoeba-vibrio interactions than they do in oysters-vibrio interactions. Hence, the virulence factors of LGP32 are key determinants of biotic interactions with multiple hosts ranging from protozoans to metazoans, suggesting that the selective pressure exerted by amoebae in marine coastal environments favor coincidental selection of virulence factors.

microbiology

Comparative validation of breast cancer risk prediction models and projections for future risk stratification

BackgroundWell-validated risk models are critical for risk stratified breast cancer prevention. We used the Individualized Coherent Absolute Risk Estimation (iCARE) tool for comparative model validation of five-year risk of invasive breast cancer in a prospective cohort, and to make projections for population risk stratification.\n\nMethodsPerformance of two recently developed models, iCARE-BPC3 and iCARE-Lit, were compared with two established models (BCRAT, IBIS) based on classical risk factors in a UK-based cohort of 64,874 women (863 cases) aged 35-74 years. Risk projections in US White non-Hispanic women aged 50-70 years were made to assess potential improvements in risk stratification by adding mammographic breast density (MD) and polygenic risk score (PRS).\n\nResultsThe best calibrated models were iCARE-Lit (expected to observed number of cases (E/O)=0.98 (95% confidence interval [CI]=0.87 to 1.11)) for women younger than 50 years; and iCARE-BPC3 (E/O=1.00 (0.93 to 1.09)) for women 50 years or older. Risk projections using iCARE-BPC3 indicated classical risk factors can identify ~500,000 women at moderate to high risk (>3% five-year risk). Additional information on MD and a PRS based on 172 variants is expected to increase this to ~3.6 million, and among them, ~155,000 invasive breast cancer cases are expected within five years.\n\nConclusionsiCARE models based on classical risk factors perform similarly or better than BCRAT or IBIS. Addition of MD and PRS can lead to substantial improvements in risk stratification. Independent prospective validation of integrated models is needed prior to clinical evaluation risk stratified breast cancer screening and prevention.

epidemiology

SimpactCyan 1.0: An Open-source Simulator for Individual-Based Models in HIV Epidemiology with R and Python Interfaces

SimpactCyan is an open-source simulator for individual-based models in HIV epidemiology. Its core algorithm is written in C++ for computational efficiency, while the R and Python interfaces aim to make the tool accessible to the fast-growing community of R and Python users. Transmission, treatment and prevention of HIV infections in dynamic sexual networks are simulated by discrete events. A generic "intervention" event allows model parameters to be changed over time, and can be used to model medical and behavioural HIV prevention programmes. First, we describe a more efficient variant of the modified Next Reaction Method that drives our continuous-time simulator. Next, we outline key built-in features and assumptions of individual-based models formulated in SimpactCyan, and provide code snippets for how to formulate, execute and analyse models in SimpactCyan through its R and Python interfaces. Lastly, we give two examples of applications in HIV epidemiology: the first demonstrates how the software can be used to estimate the impact of progressive changes to the eligibility criteria for HIV treatment on HIV incidence. The second example illustrates the use of SimpactCyan as a data-generating tool for assessing the performance of a phylodynamic inference framework.

epidemiology

Assessing the response of small RNA populations to allopolyploidy using resynthesized Brassica napus allotetraploids

Allopolyploidy, combining interspecific hybridization with whole genome duplication, has had significant impact on plant evolution. Its evolutionary success is related to the rapid and profound genome reorganizations that allow neo-allopolyploids to form and adapt. Nevertheless, how neo-allopolyploid genomes adapt to regulate their expression remains poorly understood. The hypothesis of a major role for small non-coding RNAs (sRNAs) in mediating the transcriptional response of neo-allopolyploid genomes has progressively emerged. Generally, 21-nt sRNAs mediate post-transcriptional gene silencing (PTGS) by mRNA cleavage whereas 24-nt sRNAs repress transcription (transcriptional gene silencing, TGS) through epigenetic modifications. Here, we characterize the global response of sRNAs to allopolyploidy in Brassica, using three independently resynthesized B. napus allotetraploids surveyed at two different generations in comparison with their diploid progenitors. Our results suggest an immediate but transient response of specific sRNA populations to allopolyploidy. These sRNA populations mainly target non-coding components of the genome but also target the transcriptional regulation of genes involved in response to stresses and in metabolism; this suggests a broad role in adapting to allopolyploidy. We finally identify the early accumulation of both 21- and 24-nt sRNAs involved in regulating the same targets, supporting a PTGS-to-TGS shift at the first stages of the neo-allopolyploid formation. We propose that reorganization of sRNA production is an early response to allopolyploidy in order to control the transcriptional reactivation of various non-coding elements and stress-related genes, thus ensuring genome stability during the first steps of neo-allopolyploid formation.

evolutionary biology

MR-pheWAS with stratification and interaction: Searching for the causal effects of smoking heaviness identified an effect on facial aging

Mendelian randomization (MR) is an established approach for estimating the causal effect of an environmental exposure on a downstream outcome. The gene x environment (GxE) study design can be used within an MR framework to determine whether MR estimates may be biased if the genetic instrument affects the outcome through pathways other than via the exposure of interest (known as horizontal pleiotropy). MR phenome-wide association studies (MR-pheWAS) search for the effects of an exposure, and a recently published tool (PHESANT) means that it is now possible to do this comprehensively, across thousands of traits in UK Biobank. In this study, we introduce the GxE MR-pheWAS approach, and search for the causal effects of smoking heaviness - stratifying on smoking status (ever versus never) - as an exemplar. If a genetic variant is associated with smoking heaviness (but not smoking initiation), and this variant affects an outcome (at least partially) via tobacco intake, we would expect the effect of the variant on the outcome to differ in ever versus never smokers. If this effect is entirely mediated by tobacco intake, we would expect to see an effect in ever smokers but not never smokers. We used PHESANT to search for the causal effects of smoking heaviness, instrumented by genetic variant rs16969968, among never and ever smokers respectively, in UK Biobank. We ranked results by: 1) strength of effect of rs16969968 among ever smokers, and 2) strength of interaction between ever and never smokers. We replicated previously established causal effects of smoking heaviness, including a detrimental effect on lung function and pulse rate. Novel results included a detrimental effect of heavier smoking on facial aging. We have demonstrated how GxE MR-pheWAS can be used to identify causal effects of an exposure, while simultaneously assessing the extent that results may be biased by horizontal pleiotropy.\n\nAuthor summaryMendelian randomization uses genetic variants associated with an exposure to investigate causality. For instance, a genetic variant that relates to how heavily a person smokes has been used to test whether smoking causally affects health outcomes. Mendelian randomization is biased if the genetic variant also affects the outcome via other pathways. We exploit additional information - that the effect of heavy smoking only occurs in people who actually smoke - to overcome this problem. By testing associations in ever and never smokers separately we can assess whether the genetic variant affects an outcome via smoking or another pathway. If the effect is entirely via smoking heaviness, we would expect to see an effect in ever but not never smokers, and this would suggest that smoking causally influences the outcome. Previous Mendelian randomization studies of smoking heaviness focused on specific outcomes - here we searched for the causal effects of smoking heaviness across over 18,000 traits. We identified previously established effects (e.g. a detrimental effect on lung function) and novel results including a detrimental effect of heavier smoking on facial aging. Our approach can be used to search for the causal effects of other exposures, where the exposure only occurs in known subsets of the population.

epidemiology

Evaluation of the NCCN guidelines using the RIGHT Statement and AGREE II instrument: a cross-sectional review.

IntroductionRobust, clearly reported clinical practice guidelines (CPGs) are essential for evidence-based clinical practice. The Reporting Items for practice Guidelines in HealThcare (RIGHT) statement and Appraisal of Guidelines for Research and Evaluation (AGREE) II instrument were published to improve the methodological and reporting quality in healthcare CPGs.\n\nMethodsWe applied the RIGHT statement checklist and AGREE II instrument to 48 National Comprehensive Cancer Network (NCCN) guidelines. Our primary objective was to assess the adherence to RIGHT and AGREE II items. Since neither RIGHT nor AGREE-II can judge the clinical usefulness of a guideline, our study is designed to only focus on the methodological and reporting quality of each guideline.\n\nResultsThe NCCN guidelines demonstrated notable strengths and weaknesses. For example, RIGHT statement items 19 (conflicts of interest), 7b (description of subgroups), and 13a (clear, precise recommendations) were fully reported in all guidelines. However, the guidelines inconsistently incorporated patient values and preferences and cost, nor did they consistently describe the method for assessing the quality and certainty of evidence. Regarding the AGREE II instrument, the NCCN guidelines scored highly on the domains 4 (clear, precise recommendations) and 6 (handling of conflicts of interest), but lowest on domain 2 (inclusion of all relevant stakeholders).\n\nConclusionsIn this investigation we found that NCCN CPGs demonstrate key strengths and weaknesses with respect to the reporting of key items essential to CPGs. We recommend the continued use of NCCN guidelines and adherence to the RIGHT and AGREE II items. Doing so serves to improve the evidence delivered to healthcare providers, thus potentially improving patient care.

epidemiology

Rapid statistical methods for inferring intra- and inter-hospital transmission of nosocomial pathogens from whole genome sequence data

Whole genome sequence (WGS) data for bacterial pathogens can provide evidence as to the source of nosocomial infection, and more specifically the ability to distinguish between intra- and inter-hospital transmission. This is currently achieved either through using SNP thresholds, which can lack statistical robustness, or by constructing phylogenetic trees, which can be computationally expensive and difficult to interpret. Here we compare two alternative statistical approaches using 1022 genomes of methicillin resistant Staphylococcus aureus (MRSA) clone ST22. In 71% of cases both methods predict the same hospital origin, which is also supported by the ML tree. Robust assignments are divided approximately equally between intra-hospital transmission and inter-hospital transmission. Our approaches are rapid and produce intuitive output that could inform on immediate infection control priorities, as well as providing long-term data on inter-hospital transmission networks. We discuss the strengths and weakness of our methods, and the generalisability of this approach.\n\nOne Sentence SummaryWe present rapid statistical methods for distinguishing intra- versus inter-hospital transmission of bacterial pathogens using whole genome sequence data; these methods do not require the use of SNP thresholds or the generation and interpretation of phylogenetic trees.

epidemiology

Sampling for disease absence-deriving informed monitoring from epidemic traits

Monitoring for disease requires subsets of the host population to be sampled and tested for the pathogen. If all the samples return healthy, what are the chances the disease was present but missed? In this paper, we developed a statistical approach to solve this problem considering the fundamental property of infectious diseases: their growing incidence in the host population. The model gives an estimate of the incidence probability density as a function of the sampling effort, and can be reversed to derive adequate monitoring patterns ensuring a given maximum incidence in the population. We then present an approximation of this model, providing a simple rule of thumb for practitioners. The approximation is shown to be accurate for a sample size larger than 20, and we demonstrate its use by applying it to three plant pathogens: citrus canker, bacterial blight and grey mould.

epidemiology

Deep Convolutional modeling of human face selective columns reveals their role in pictorial face representation

Despite the massive accumulation of systems neuroscience findings, their functional meaning remains tentative, largely due to the absence of realistically performing models. The discovery that deep convolutional networks achieve human performance in realistic tasks offers fresh opportunities for such modeling. Here we show that the face-space topography of face-selective columns recorded intra-cranially in 32 patients significantly matches that of a DCNN having human-level face recognition capabilities. Three modeling aspects converge in pointing to a role of human face areas in pictorial rather than person identification: First, the match was confined to intermediate layers of the DCNN. Second, identity preserving image manipulations abolished the brain to DCNN correlation. Third, DCNN neurons matching face-column tuning displayed view-point selective receptive fields. Our results point to a \"convergent evolution\" of pattern similarities in biological and artificial face perception. They demonstrate DCNNs as a powerful modeling approach for deciphering the function of human cortical networks.

neuroscience

LuxRep: a technical replicate-aware method for bisulfite sequencing data analysis

DNA methylation is measured using bisulfite sequencing (BS-seq). Bisulfite conversion can have low efficiency and a DNA sample is then processed multiple times generating DNA libraries with different bisulfite conversion rates. Libraries with low conversion rates are excluded from analysis resulting in reduced coverage and increased costs. We present a method and software, LuxRep, that accounts for technical replicates from different bisulfite-converted DNA libraries. We show that including replicates with low bisulfite conversion rates generates more accurate estimates of methylation levels and differentially methylated sites.\n\nAvailabilityAn implementation of the method is available at https://github.com/tare/LuxGLM/tree/master/LuxRep\n\nContactmaia.malonzo@aalto.fi

bioinformatics

Molecularly distinct models of zebrafish Myc-induced B cell leukemia

Zebrafish models of T cell acute lymphoblastic leukemia (T-ALL) have been studied for over a decade, but curiously, robust zebrafish B cell ALL (B-ALL) models had not been described. Recently, our laboratories reported two seemingly closely-related models of zebrafish B-ALL. In these genetic lines, the primary difference is expression of either murine or human transgenic c-MYC, each controlled by the zebrafish rag2 promoter. Here, we compare ALL gene expression in both models. Surprisingly, we find that B-ALL arise in different B cell lineages, with ighm+ vs. ighz+ B-ALL driven by murine Myc vs. human MYC, respectively. Moreover, these B-ALL types exhibit signatures of distinct molecular pathways, further unexpected dissimilarity. Thus, despite sharing analogous genetic makeup, the ALL types in each model are markedly different, proving subtle genetic changes can profoundly impact model organism phenotypes. Investigating the mechanistic differences between mouse and human c-MYC in these contexts may reveal key functional aspects governing MYC-driven oncogenesis in human malignancies.

cancer biology

Associations between environmental breast cancer risk factors and DNA methylation-based risk-predicting measures

BackgroundGenome-wide average DNA methylation (GWAM) and epigenetic age acceleration have been suggested to predict breast cancer risk. We aimed to investigate the relationships between these putative risk-predicting measures and environmental breast cancer risk factors.\n\nMethodsUsing the Illumina HumanMethylation450K assay methylation data, we calculated GWAM and epigenetic age acceleration for 132 female twin pairs and their 215 sisters. Linear regression was used to estimate associations between these risk-predicting measures and multiple breast cancer risk factors. Within-pair analysis was performed for the 132 twin pairs.\n\nResultsGWAM was negatively associated with number of live births, and positively with age at first live birth (both P<0.05). Epigenetic age acceleration was positively associated with body mass index (BMI), smoking, alcohol drinking and age at menarche, and negatively with age at first live birth (all P<0.05), and the associations with BMI, alcohol drinking and age at first live birth remained in the within-pair analysis.\n\nConclusionsThis exploratory study shows that lifestyle and hormone-related breast cancer risk factors are associated with DNA methylation-based measures that could predict breast cancer risk. The associations of epigenetic age acceleration with BMI, alcohol drinking and age at first live birth are unlikely to be due to familial confounding.

epidemiology

HiAlc Klebsiella pneumonia, one of potential chief culprits of non-alcoholic fatty liver disease: through generation of endogenous ethanol

Non-alcoholic fatty liver disease (NAFLD), a prelude of cirrhosis and hepatocellular carcinoma, is the most common chronic liver disease worldwide. NAFLD has been considerated to be associated with the composition of gut microbiota. However, causal relationship between change of gut microbiome and NAFLD remains unclear. Here we show that Klebsiella pneumoniae was significantly associated with NAFLD through inducing generation of endogenous ethanol. A strain of high alcohol-producing Klebsiella pneumoniae (HiAlc Kpn) was initially isolated from fecal samples of patient with non-alcoholic steatohepatitis (NASH) accompanied with auto-brewery syndrome (ABS). Gavage of HiAlc Kpn was capable of inducing murine model of fatty liver disease (FLD) in which had typical pathological changes of hepatic steatosis and similar liver gene expression profiles to those of alcohol intake in mice. Data derived from germ-free mice by gnotobiotic gavage further demonstrated that the HiAlc Kpn is the major cause of the changes in FLD mice. Furthermore, using proteomic and metabolitic analysis, we found that HiAlc Kpn induced generation of endogenous alcohol through the 2,3-butanediol fermentation pathway. More interestingly, the blood alcohol concentration was elevated in FLD mice induced by HiAlc Kpn after glucose intake. Clinical analysis showed that HiAlc Kpn were observed in up to 60% of patients with NAFLD. Our results suggested that HiAlc Kpn make important contribution to NAFLD, possibly through generation of the endogenous alcohol. Thus, targeting these bacteria might provide a novel therapeutic for clinical treatment of NAFLD.\n\nIn BriefFatty liver disease induced by high alcohol-producing Klebsiella pneumoniae\n\nCompeting Financial Interest StatementThe authors declare no conflicts of interest.

microbiology