Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “systems biology”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,711 records · Page 95Linked to original sources

Identification of a novel therapeutic agent for Inflammatory Bowel Disease guided by systems medicine

ObjectiveInflammatory bowel diseases cause significant morbidity and mortality. Aberrant NF-{kappa}B signalling is strongly associated with these conditions, and several established drugs influence the NF-{kappa}B signalling network to exert their effect. This study aimed to identify drugs which alter NF-{kappa}B signalling and may be repositioned for use in inflammatory bowel disease. DesignThe SysmedIBD consortium established a novel drug-repurposing pipeline based on a combination of in-silico drug discovery and biological assays targeted at demonstrating an impact on NF-kappaB signalling, and a murine model of IBD. ResultsThe drug discovery algorithm identified several drugs already established in IBD, including corticosteroids. The highest-ranked drug was the macrolide antibiotic Clarithromycin, which has previously been reported to have anti-inflammatory effects in aseptic conditions. Clarithromycins effects were validated in several experiments: it influenced NF-{kappa}B mediated transcription in murine peritoneal macrophages and intestinal enteroids; it suppressed NF-{kappa}B protein shuttling in murine reporter enteroids; it suppressed NF-{kappa}B (p65) DNA binding in the small intestine of mice exposed to LPS, and it reduced the severity of dextran sulphate sodium-induced colitis in C57BL/6 mice. Clarithromycin also suppressed NF-{kappa}B (p65) nuclear translocation in human intestinal enteroids. ConclusionsThese findings demonstrate that in-silico drug repositioning algorithms can viably be allied to laboratory validation assays in the context of inflammatory bowel disease; and that further clinical assessment of clarithromycin in the management of inflammatory bowel disease is required.

systems biology

UPS-indel: A Universal Positioning System For Indels

BackgroundIndels, though differing in allele sequence and position, are biologically equivalent when they lead to the same altered sequences. Storing biologically equivalent indels as distinct entries in databases causes data redundancy, and may mislead downstream analysis and interpretations. About 10% of the human indels stored in dbSNP are redundant. It is thus desirable to have a unified system for identifying and representing equivalent indels in publically available databases. Moreover, a unified system is also desirable to compare the indel calling results produced by different tools. This paper describes UPS-indel, a utility tool that creates a universal positioning system for indels so that equivalent indels can be uniquely determined by their coordinates in the new system, which also can be used to compare indel calling results produced by different tools.\n\nResultsUPS-indel identifies nearly 15% indels in dbSNP (version 142) as redundant across all human chromosomes, higher than previously reported. When applied to COSMIC coding and noncoding indel datasets, UPS-indel identifies nearly 29% and 13% indels as redundant, respectively. Comparing the performance of UPS-indel with existing variant normalization tools vt normalize, BCFtools, and GATK LeftAlignAndTrimVariants shows that UPS-indel is able to identify 456,352 more redundant indels in dbSNP; 2,118 more in COSMIC coding, and 553 more in COSMIC noncoding indel dataset in addition to the ones reported jointly by these tools. Moreover, comparing UPS-indel to other state-of-the-art approaches for indel call set comparison demonstrates that UPS-indel is clearly superior to other approaches in finding indels in common among call sets.\n\nConclusionsUPS-indel is theoretically proven to find all equivalent indels, and is thus exhaustive. UPS-indel is written in C++ and the command line version is freely available to download at http://ups-indel.sourceforge.net. The online version of UPS-indel is available at http://bench.cs.vt.edu/ups-indel/.

bioinformatics

GROOLS: reactive graph reasoning for genome annotation through biological processes

BackgroundHigh quality functional annotation is essential for understanding the phenotypic consequences encoded in a genome. Despite improvements in bioinformatics methods, millions of sequences in databanks are not assigned reliable functions. The curation of protein functions in the context of biological processes is a way to evaluate and improve their annotation.\n\nResultsWe developed an expert system using paraconsistent logic, named GROOLS (Genomic Rule Object-Oriented Logic System), that evaluates the completeness and the consistency of predicted functions through biological processes like metabolic pathways. Using a generic and hierarchical representation of knowledge, biological processes are modeled in a graph from which observations (i.e. predictions and expectations) are propagated by rules. At the end of the reasoning, conclusions are assigned to biological process components and highlight uncertainties and inconsistencies. Results on 14 microbial organisms are presented.\n\nConclusionsGROOLS software is designed to evaluate the overall accuracy of functional unit and pathway predictions according to organism experimental data like growth phenotypes. It assists biocurators in the functional annotation of proteins by focusing on missing or contradictory observations.

bioinformatics

Electrical-charge accumulation enables integrative quality control during B. subtilis sporulation

Quality control of offspring is important for the survival of cells. However, the mechanism by which quality of offspring cells may be monitored while running genetic programs of cellular differentiation remains largely unclear. Here we investigated a quality control system during Bacillus subtilis spore formation by combining single-cell time-lapse microscopy, molecular biology and mathematical modelling. Our results revealed that the quality-control system via premature germination is coupled with the accumulation of cations on the surface of developing forespores. Specifically, the forespores accumulating less cations on their surface are more likely to be aborted. This charge accumulation system enables the projection of multidimensional information about the external environment and morphological development of the forespore onto a one-dimensional information of cation accumulation. Based on the insight we gain, we propose a novel use of Nernstian chemicals for reducing the yield and quality of Bacillus endospores.

biophysics

Multiple genetic changes underlie the evolution of long-tailed forest deer mice

Understanding both the role of selection in driving phenotypic change and its underlying genetic basis remain major challenges in evolutionary biology. Here we focus on a classic system of local adaptation in the North American deer mouse, Peromyscus maniculatus, which occupies two main habitat types, prairie and forest. Using historical collections we demonstrate that forest-dwelling mice have longer tails than those from non-forested habitats, even when we account for individual and population relatedness. Based on genome-wide SNP capture data, we find that mice from forested habitats in the eastern and western parts of their range form separate clades, suggesting that increased tail length evolved independently from a short-tailed ancestor. Two major changes in skeletal morphology can give rise to longer tails--increased number and increased length of vertebrae--and we find that forest mice in the east and west have both more and longer caudal vertebrae, but not trunk vertebrae, than nearby prairie forms. Using a second-generation intercross between a prairie and forest pair, we show that the number and length of caudal vertebrae are not correlated in this recombinant population, suggesting that variation in these traits is controlled by separate genetic loci. Together, these results demonstrate convergent evolution of the long-tailed forest phenotype through multiple, distinct genetic mechanisms (controlling vertebral length and vertebral number), thus suggesting that these morphological changes--either independently or together--are adaptive.

Evolutionary Biology

Clinical data specification and coding for cross-analyses with omics data in autoimmune disease trials

ObjectivesAutoimmune and inflammatory diseases (AIDs) form a continuum of autoimmune and inflammatory diseases, yet AIDs nosology is based on syndromic classification. The TRANSIMMUNOM trial (NCT02466217) was designed to re-evaluate AIDs nosology through clinic-biological and multi-omics investigations of patients with one of 19 selected AIDs. To allow cross-analyses of clinic-biological data together with omics data, we needed to integrate clinical data in a harmonized database.\n\nMaterials and MethodsWe assembled a clinical expert consortium (CEC) to select relevant clinic-biological features to be collected for all patients and a cohort management team comprising biologists, clinicians and computer scientists to design an electronic case report form (eCRF). The eCRF design and implementation has been done on OpenClinica, an open-source CFR-part 11 compliant electronic data capture system.\n\nResultsThe CEC selected 865 clinical and biological parameters. The CMT selected coded the items using CDISC standards into 5835 coded values organized in 28 structured eCRFs. Examples of such coding are check boxes for clinical investigation, numerical values with units, disease scores as a result of an automated calculations, and coding of possible treatment formulas, doses and dosage regimens per disease.\n\nDiscussion21 CRFs were designed using OpenClinica v3.14 capturing the 5835 coded values per patients. Technical adjustment have been implemented to allow data entry and extraction of this amount of data, rarely achieved in classical eCRFs designs.\n\nConclusionsA multidisciplinary endeavour offers complete and harmonized CRFs for AID clinical investigations that are used in TRANSIMMUNOM and will benefit translational research team.

clinical trials

A Bayesian Mixture Modelling Approach For Spatial Proteomics

AbstractAnalysis of the spatial sub-cellular distribution of proteins is of vital importance to fully understand context specific protein function. Some proteins can be found with a single location within a cell, but up to half of proteins may reside in multiple locations, can dynamically re-localise, or reside within an unknown functional compartment. These considerations lead to uncertainty in associating a protein to a single location. Currently, mass spectrometry (MS) based spatial proteomics relies on supervised machine learning algorithms to assign proteins to sub-cellular locations based on common gradient profiles. However, such methods fail to quantify uncertainty associated with sub-cellular class assignment. Here we reformulate the framework on which we perform statistical analysis. We propose a Bayesian generative classifier based on Gaussian mixture models to assign proteins probabilistically to sub-cellular niches, thus proteins have a probability distribution over sub-cellular locations, with Bayesian computation performed using the expectation-maximisation (EM) algorithm, as well as Markov-chain Monte-Carlo (MCMC). Our methodology allows proteome-wide uncertainty quantification, thus adding a further layer to the analysis of spatial proteomics. Our framework is flexible, allowing many different systems to be analysed and reveals new modelling opportunities for spatial proteomics. We find our methods perform competitively with current state-of-the art machine learning methods, whilst simultaneously providing more information. We highlight several examples where classification based on the support vector machine is unable to make any conclusions, while uncertainty quantification using our approach provides biologically intriguing results. To our knowledge this is the first Bayesian model of MS-based spatial proteomics data.\n\nAuthor summarySub-cellular localisation of proteins provides insights into sub-cellular biological processes. For a protein to carry out its intended function it must be localised to the correct sub-cellular environment, whether that be organelles, vesicles or any sub-cellular niche. Correct sub-cellular localisation ensures the biochemical conditions for the protein to carry out its molecular function are met, as well as being near its intended interaction partners. Therefore, mis-localisation of proteins alters cell biochemistry and can disrupt, for example, signalling pathways or inhibit the trafficking of material around the cell. The sub-cellular distribution of proteins is complicated by proteins that can reside in multiple micro-environments, or those that move dynamically within the cell. Methods that predict protein sub-cellular localisation often fail to quantify the uncertainty that arises from the complex and dynamic nature of the sub-cellular environment. Here we present a Bayesian methodology to analyse protein sub-cellular localisation. We explicitly model our data and use Bayesian inference to quantify uncertainty in our predictions. We find our method is competitive with state-of-the-art machine learning methods and additionally provides uncertainty quantification. We show that, with this additional information, we can make deeper insights into the fundamental biochemistry of the cell.

systems biology

Personalized characterization of diseases using sample-specific networks

A complex disease generally results not from malfunction of individual molecules but from dysfunction of the relevant system or network, which dynamically changes with time and conditions. Thus, estimating a condition-specific network from a sample is crucial to elucidating the molecular mechanisms of complex diseases at the system level. However, there is currently no effective way to construct such an individual-specific network by expression profiling of a single sample because of the requirement of multiple samples for computing correlations. We developed here with a statistical method, i.e., a sample-specific network method, which allows us to construct individual-specific networks based on molecular expression of a single sample. Using this method, we can characterize various human diseases at a network level. In particular, such sample-specific networks can lead to the identification of individual-specific disease modules as well as driver genes, even without gene sequencing information. Extensive analysis by using the Cancer Genome Atlas data not only demonstrated the effectiveness of the method, but also found new individual-specific driver genes and network patterns for various cancers. Biological experiments on drug resistance further validated one important advantage of our method over the traditional methods, i.e., we even identified those drug resistance genes that actually have no clearly differential expression between samples with and without the resistance, due to the additional network information.

Systems Biology

Versatile approach for functional analysis of human proteins and efficient stable cell line generation using FLP-mediated recombination system

Deciphering a function of a given protein requires investigating various biological aspects. Usually, the protein of interest is expressed with a fusion tag that aids or allows subsequent analyses. Additionally, downregulation or inactivation of the studied gene enables functional studies. Development of the CRISPR/Cas9 methodology opened many possibilities but in many cases it is restricted to non-essential genes. It may also be time-consuming if a homozygote is needed. Recombinase-dependent gene integration methods, like the Flp-In system, are very good alternative. The system is widely used in different research areas, which calls for the existence of compatible vectors and efficient protocols that ensure straightforward DNA cloning and creation of stable cell lines. We have created and validated a robust series of 52 vectors for streamlined generation of stable mammalian cell lines using the FLP recombinase-based methodology. Using the sequence-independent DNA cloning method all constructs for a given coding-sequence can be made with just three universal PCR primers. The collection allows tetracycline-inducible expression of proteins with various tags suitable for protein localization, FRET, bimolecular fluorescence complementation (BiFC), protein dynamics studies (FRAP), co-immunoprecipitation, the RNA tethering assay and cell sorting. Some of the vectors contain a bidirectional promoter for concomitant expression of miRNA and mRNA, so that a gene can be silenced and its product replaced by a mutated miRNA-insensitive version. We demonstrate the efficacy of our vectors by creating stable cell lines with various tagged proteins (numatrin, fibrillarin, coilin, centrin, THOC5, PCNA). We have analysed transgene expression over time to provide a guideline for future experiments and compared the utility of commonly used inducers of tetracycline-responsive promoters. We determined the protein interaction network of the exoribonuclease XRN2 and examined the role of the protein in transcription termination by RNAseq analysis of cells devoid of its ribonucleolytic activity. In total we created more than 500 DNA constructs which proves high efficiency of our strategy.

molecular biology

Development of a multi-locus CRISPR gene drive system in budding yeast

The discovery of CRISPR/Cas gene editing has allowed for major advances in many biomedical disciplines and basic research. One arrangement of this biotechnology, a nuclease-based gene drive, can rapidly deliver a genetic element through a given population and studies in fungi and metazoans have demonstrated the success of such a system. This methodology has the potential to control biological populations and contribute to eradication of insect-borne diseases, agricultural pests, and invasive species. However, there remain challenges in the design, optimization, and implementation of gene drives including concerns regarding biosafety, containment, and control/inhibition. Given the numerous gene drive arrangements possible, there is a growing need for more advanced designs. In this study, we use budding yeast to develop an artificial multi-locus gene drive system. Our minimal setup requires only a single copy of S. pyogenes Cas9 and three guide RNAs to propagate three separate gene drives. We demonstrate how this system could be used for targeted allele replacement of native genes and to suppress NHEJ repair systems by modifying DNA Ligase IV. A multi-locus gene drive configuration provides an expanded suite of options for complex attributes including pathway redundancy, combatting evolved resistance, and safeguards for control, inhibition, or reversal of drive action.

synthetic biology

Tissue-specific tagging of endogenous loci in Drosophila melanogaster

Fluorescent protein tags have revolutionized cell and developmental biology, and in combination with binary expression systems they enable diverse tissue-specific studies of protein function. However these binary expression systems often do not recapitulate endogenous protein expression levels, localization, binding partners, and developmental windows of gene expression. To address these limitations, we have developed a method called T-STEP (Tissue-Specific Tagging of Endogenous Proteins) that allows endogenous loci to be tagged in a tissue specific manner. T-STEP uses a combination of efficient gene targeting and tissue-specific recombinase-mediated tag swapping to temporally and spatially label endogenous proteins. We have employed this method to GFP tag OCRL (a phosphoinositide-5-phosphatase in the endocytic pathway) and Vps35 (a Parkinson's disease-implicated component of the endosomal retromer complex) in diverse Drosophila tissues including neurons, glia, muscles, and hemocytes. Selective tagging of endogenous proteins allows for the first time cell type-specific live imaging and proteomics in complex tissues.

Neuroscience

Dynamic information routing in complex networks

Abstract Flexible information routing fundamentally underlies the function of many biological and artificial networks. Yet, how such systems may specifically communicate and dynamically route information is not well understood. Here we identify a generic mechanism to route information on top of collective dynamical reference states in complex networks. Switching between collective dynamics induces flexible reorganization of information sharing and routing patterns, as quantified by delayed mutual information and transfer entropy measures between activities of a network's units. We demonstrate the power of this generic mechanism specifically for oscillatory dynamics and analyze how individual unit properties, the network topology and external inputs coact to systematically organize information routing. For multi-scale, modular architectures, we resolve routing patterns at all levels. Interestingly, local interventions within one sub-network may remotely determine non-local network-wide communication. These results help understanding and designing information routing patterns across systems where collective dynamics co-occurs with a communication function.

Neuroscience

Error-prone bypass of DNA lesions during lagging strand replication is a common source of germline and cancer mutations

Spontaneously occurring mutations are of great relevance in diverse fields including biochemistry, oncology, evolutionary biology, and human genetics. Studies in experimental systems have identified a multitude of mutational mechanisms including DNA replication infidelity as well as many forms of DNA damage followed by inefficient repair or replicative bypass. However, the relative contributions of these mechanisms to human germline mutations remain completely unknown. Here, based on the mutational asymmetry with respect to the direction of replication and transcription, we suggest that error-prone damage bypass on the lagging strand plays a major role in human mutagenesis. Asymmetry with respect to transcription is believed to be mediated by the action of transcription-coupled DNA repair (TC-NER). TC-NER selectively repairs DNA lesions on the transcribed strand; as a result, lesions on the non-transcribed strand are preferentially converted into mutations. In human polymorphism we detect a striking similarity between transcriptional asymmetry and asymmetry with respect to replication fork direction. This parallels the observation that damage-induced mutations in human cancers accumulate asymmetrically with respect to the direction of replication, suggesting that DNA lesions are asymmetrically resolved during replication. Re-analysis of XR-seq data, Damage-seq data and cancers with defective NER corroborate the preferential error-prone bypass of DNA lesions on the lagging strand. We experimentally demonstrate that replication delay greatly attenuates the mutagenic effect of UV-irradiation, in line with the key role of replication in conversion of DNA damage to mutations. We conservatively estimate that at least 10% of human germline mutations arise due to DNA damage rather than replication infidelity. The number of these damage-induced mutations is expected to scale with the number of replications and, consequently, paternal age.

biochemistry

Genome wide association analysis identifies genetic variants associated with reproductive variation across domestic dog breeds and uncovers links to domestication

The diversity of eutherian reproductive strategies has led to variation in many traits, such as number of offspring, age of reproductive maturity, and gestation length. While reproductive trait variation has been extensively investigated and is well established in mammals, the genetic loci contributing to this variation remain largely unknown. The domestic dog, Canis lupus familiaris is a powerful model for studies of the genetics of inherited disease due to its unique history of domestication. To gain insight into the genetic basis of reproductive traits across domestic dog breeds, we collected phenotypic data for four traits - cesarean section rate (n = 97 breeds), litter size (n = 60), stillbirth rate (n = 57), and gestation length (n = 23) - from primary literature and breeders handbooks. By matching our phenotypic data to genomic data from the Cornell Veterinary Biobank, we performed genome wide association analyses for these four reproductive traits, using body mass and kinship among breeds as co-variates. We identified 14 genome-wide significant associations between these traits and genetic loci, including variants near CACNA2D3 with gestation length, MSRB3 with litter size, SMOC2 with cesarean section rate, MITF with litter size and still birth rate, KRT71 with cesarean section rate, litter size, and stillbirth rate, and HTR2C with stillbirth rate. Some of these loci, such as CACNA2D3 and MSRB3, have been previously implicated in human reproductive pathologies. Many of the variants that we identified have been previously associated with domestication-related traits, including brachycephaly (SMOC2), coat color (MITF), coat curl (KRT71), and tameness (HTR2C). These results raise the hypothesis that the artificial selection that gave rise to dog breeds also shaped the observed variation in their reproductive traits. Overall, our work establishes the domestic dog as a system for studying the genetics of reproductive biology and disease.

evolutionary biology

Multi-Hierarchical Dynamics of Antimicrobial Resistance Simulated in a Nested Membrane Computing Model

Membrane Computing is a bio-inspired computing paradigm, whose devices are the so-called membrane systems or P systems. The P system designed in this work reproduces complex biological landscapes in the computer world. It uses nested \"membrane-surrounded entities\" able to divide, propagate and die, be transferred into other membranes, exchange informative material according to flexible rules, mutate and being selected by external agents. This allows the exploration of hierarchical interactive dynamics resulting from the probabilistic interaction of genes (phenotypes), clones, species, hosts, environments, and antibiotic challenges. Our model facilitates analysis of several aspects of the rules that govern the multi-level evolutionary biology of antibiotic resistance. We examine a number of selected landscapes where we predict the effects of different rates of patient flow from hospital to the community and viceversa, cross-transmission rates between patients with bacterial propagules of different sizes, the proportion of patients treated with antibiotics, antibiotics and dosing in opening spaces in the microbiota where resistant phenotypes multiply. We can also evaluate the selective strength of some drugs and the influence of the time-0 resistance composition of the species and bacterial clones in the evolution of resistance phenotypes. In summary, we provide case studies analyzing the hierarchical dynamics of antibiotic resistance using a novel computing model with reciprocity within and between levels of biological organization, a type of approach that may be expanded in the multi-level analysis of complex microbial landscapes.

microbiology

Pancreatic adenocarcinoma human organoids share structural and genetic features with primary tumors

Patient-derived pancreatic ductal adenocarcinoma (PDAC) organoid systems show great promise for understanding the biological underpinnings of disease and advancing therapeutic precision medicine. Despite the increased use of organoids, the fidelity of molecular features, genetic heterogeneity, and drug response to the tumor of origin remain important unanswered questions limiting their utility. To address this gap in knowledge, we created primary tumor- and PDX-derived organoids, and 2D cultures for in-depth genomic and histopathological comparisons to the primary tumor. Histopathological features and PDAC representative protein markers showed strong concordance. DNA and RNA sequencing of single organoids revealed patient-specific genomic and transcriptomic consistency. Single-cell RNAseq demonstrated that organoids are primarily a clonal population. In drug response assays, organoids displayed patient-specific sensitivities. Additionally, we examined the in vivo PDX response to FOLFIRINOX and Gemcitabine/Abraxane treatments, which was recapitulated in vitro by organoids. The patient-specific molecular and histopathological fidelity of organoids indicate that they can be used to understand the etiology of the patients tumor and the differential response to therapies and suggests utility for predicting drug responses.

cancer biology

Learning Molecule Drug Function from Structure Representations with Deep Neural Networks or Random Forests

Empirical testing of chemicals for desirable properties, such as drug efficacy, toxicity, solubility, and thermal conductivity, costs many billions of dollars every year. Further, the choice of molecules for testing relies on expert knowledge and intuition in the relevant domain of application. The ability to predict the action of a molecule in silico would greatly increase the speed and decrease the cost of prioritizing molecules with desirable function for experimental testing. We explore the use of molecule structures to predict their drug classification without any explicit biological model (e.g. protein structure or cell system). Two approaches were compared: (1) a two-dimensional image representation of molecule structure with transfer learning from a pre-trained convolutional neural network (CNN), and (2) Morgan molecular fingerprints (MFP) with a random forest (RF). Both methods using only molecule information as input achieved significantly better accuracy than previous work that used transcriptomic measurements of molecular effects. Because these data suggest that much of a molecules function is encoded by its chemical structure, we explore which general molecular properties are drug class-specific and how misclassification of structures might provide drug repurposing opportunities.

bioinformatics

Introducing THOR, a model microbiome for genetic dissection of community behavior

The quest to manipulate microbiomes has intensified, but many microbial communities have proven recalcitrant to sustained change. Developing model communities amenable to genetic dissection will underpin successful strategies for shaping microbiomes by advancing understanding of community interactions. We developed a model community with representatives from three dominant rhizosphere taxa: the Firmicutes, Proteobacteria, and Bacteroidetes. We chose Bacillus cereus as a model rhizosphere Firmicute and characterized twenty other candidates, including \"hitchhikers\" that co-isolated with B. cereus from the rhizosphere. Pairwise analysis produced a hierarchical interstrain-competition network. We chose two hitchhikers -- Pseudomonas koreensis from the top tier of the competition network and Flavobacterium johnsoniae from the bottom of the network to represent the Proteobacteria and Bacteroidetes, respectively. The model community has several emergent properties--induction of dendritic expansion of B. cereus colonies by either of the other members and production of more robust biofilms by the three members together than individually. Moreover, P. koreensis produces a novel family of alkaloid antibiotics that inhibit growth of F. johnsoniae, and production is inhibited by B. cereus. We designate this community THOR, because the members are the hitchhikers of the rhizosphere. The genetic, genomic, and biochemical tools available for dissection of THOR provide the means to achieve a new level of understanding of microbial community behavior.\n\nIMPORTANCEThe manipulation and engineering of microbiomes could lead to improved human health, environmental sustainability, and agricultural productivity. However, microbiomes have proven difficult to alter in predictable ways and their emergent properties are poorly understood. The history of biology has demonstrated the power of model systems to understand complex problems such as gene expression or development. Therefore, a defined and genetically tractable model community would be useful to dissect microbiome assembly, maintenance, and processes. We have developed a tractable model rhizosphere microbiome, designated THOR, containing Pseudomonas koreensis, Flavobacterium johnsoniae, and Bacillus cereus, which represent three dominant phyla in the rhizosphere, as well as in soil and the mammalian gut. The model community demonstrates emergent properties and the members are amenable to genetic dissection. We propose that THOR will be a useful model for investigations of community-level interactions.

microbiology