Search bioRxivSearch

Biology subjects

Kim, S.

Publications and source records attributed to Kim, S..

At least 19 recordsLinked to original sources

Cryo-EM structures reveal the mechanism of phosphatidylserine remodeling by membrane-bound glycerophospholipid O-acyltransferase 1

Lands cycle remodeling of glycerophospholipid acyl chains is crucial for cells to maintain appropriate membrane composition. Glycerophospholipids are cleaved at the glycerol sn2-position by phospholipase A. The lysophospholipids are reacylated by enzymes of the membrane-bound O-acyltransferase (MBOAT) family to incorporate specific fatty-acyl chains to adjust membrane properties. How MBOAT enzymes recognize specific acyl-CoA donors, select lysophospholipid acceptors, and release products is unclear. Phosphatidylserine (PS), a critical anionic phospholipid, controls membrane surface charge, signaling-protein recruitment, and cell-death-associated membrane recognition, and PS acyl-chain remodeling is linked to ferroptosis resistance. Here, we showed that MBOAT1 preferentially generates monounsaturated fatty acid-containing PS from lyso-PS. High-resolution cryo-electron microscopy structures of human MBOAT1 captured distinct binding poses of the fatty acyl donor, lyso-PS acceptor, and PS product. With lipidomics, enzymology and molecular dynamics simulations, these structures reveal the mechanism and pathway of MBOAT1-dependent PS remodeling.

biochemistry

The whale shark genome reveals how genomic and physiological properties scale with body size

The endangered whale shark (Rhincodon typus) is the largest fish on Earth and is a long-lived member of the ancient Elasmobranchii clade. To characterize the relationship between genome features and biological traits, we sequenced and assembled the genome of the whale shark and compared its genomic and physiological features to those of 81 animals and yeast. We examined scaling relationships between body size, temperature, metabolic rates, and genomic features and found both general correlations across the animal kingdom and features specific to the whale shark genome. Among animals, increased lifespan is positively correlated to body size and metabolic rate. Several genomic features also significantly correlated with body size, including intron and gene length. Our large-scale comparative genomic analysis uncovered general features of metazoan genome architecture: GC content and codon adaptation index are negatively correlated, and neural connectivity genes are longer than average genes in most genomes. Focusing on the whale shark genome, we identified multiple features that significantly correlate with lifespan. Among these were very long gene length, due to large introns highly enriched in repetitive elements such as CR1-like LINEs, and considerably longer neural genes of several types, including connectivity, activity, and neurodegeneration genes. The whale sharks genome had an expansion of gene families related to fatty acid metabolism and neurogenesis, with the slowest evolutionary rate observed in vertebrates to date. Our comparative genomics approach uncovered multiple genetic features associated with body size, metabolic rate, and lifespan, and showed that the whale shark is a promising model for studies of neural architecture and lifespan.

genomics

Dosage Matters: A Randomized Controlled Trial of Rehabilitation Dose in the Chronic Phase after Stroke

Background and PurposeFor stroke rehabilitation, task-specific training in animal models and human rehabilitation trials is considered important to trigger inherent neuroplasticity, promote motor learning, and functional recovery. Little is known, however, about what constitutes an effective dosage of therapy.\n\nMethodsThis is a parallel group, four arm, single blind, phase I, randomized control trial of four dosages of upper extremity therapy delivered in an outpatient setting during the chronic phase after stroke. Participants were randomized into groups that varied in total dosage of therapy (i.e., 0, 15, 30, or 60 hours). Seven hundred and four participants were assessed for eligibility, 50 were eligible to enroll, 45 were randomized, 44 participated and 41 completed the study. Planned primary analyses used linear mixed effects regression to model baseline to post-intervention changes in the Motor Activity Log-Quality of Movement rating (MALQ) and the Wolf Motor Function Test (WMFT) time score as a function of therapy dosage. A series of hierarchical models were constructed using the MALQ and WMFT.\n\nResultsWe observed a significant dose response curve: the greater the dosage of training, the greater the change in MALQ, with the dose by week slope parameter of 0.0045 ({Delta}MAL/hour/week; p = 0.0011; 95% CI = [0.0019; 0.0071]). Over the 3 weeks of therapy, this corresponds to a gain of 0.81 in MALQ for the 60 hour dose.\n\nConclusionsFor mild-to-moderately impaired stroke survivors, the dosage of a patient-centered, task specific motor therapy was shown to systematically influence the gain in quality of arm use in the natural environment, but not functional capacity as measured in the laboratory. We highlight the importance of recovery outcomes that capture arm use vs. functional capacity.\n\nClinical Trial RegistrationURL: http://www.clinicaltrials.gov. Unique identifier: NCT 01749358

clinical trials

An anti-Gn glycoprotein antibody from a convalescent patient potently inhibits the infection of severe fever with thrombocytopenia syndrome virus

Severe fever with thrombocytopenia syndrome (SFTS) is an emerging infectious disease localized to China, Japan, and Korea that is characterized by severe hemorrhage and a high fatality rate. Currently, no specific vaccine or treatment has been approved for this disease. To develop a therapeutic agent for SFTS, we isolated antibodies from a phage-displayed antibody library that was constructed from a patient who recovered from SFTS virus (SFTSV) infection. One antibody, designated as Ab10, was reactive to the Gn envelope glycoprotein of SFTSV and protected host cells and A129 mice from infection in both in vitro and in vivo experiments. Notably, Ab10 protected 80% of mice, even when injected 5 days after inoculation with a lethal dose of SFTSV. Using cross-linker assisted mass spectrometry and alanine scanning, we located the non-linear epitope of Ab10 on the Gn glycoprotein domain II and an unstructured stem region, suggesting that Ab10 may inhibit a conformational alteration that is critical for cell membrane fusion between the virus and host cell. Ab10 reacted to recombinant Gn glycoprotein in Gangwon/Korea/2012, HB28, and SD4 strains. Additionally, based on its epitope, we predict that Ab10 binds the Gn glycoprotein in 247 of 272 reported SFTSV isolates previously reported. Together, these data suggest that Ab10 has potential to be developed into a therapeutic agent that could protect against more than 90% of reported SFTSV isolates.\n\nAuthor summarySevere fever with thrombocytopenia syndrome (SFTS) is an emerging infectious disease localized to China, Japan, and Korea. This tick-borne virus has infected more than 5,000 humans with a 6.4% to 20.9% fatality rate. Currently, there are no prophylactic or therapeutic measures against this virus. Historically, antibodies from patients who recovered from viral infection have been used to treat new patients. Until now, one recombinant monoclonal antibody was approved for the prophylaxis of respiratory syntial virus infection. We selected 10 antibodies from a patient who recovered from SFTS and found that one antibody potently inhibited SFTS viral infection in both test tube and animal studies. We determined the binding site of this antibody to SFTS virus, which allowed us to predict that this antibody could bind 247 out of 272 SFTS virus isolates reported up to now. We anticipate that this antibody could be developed into a therapeutic measure against SFTS.

biochemistry

RNA polymerases display collaborative and antagonistic group behaviors over long distances through DNA supercoiling

Transcription by RNA polymerases (RNAPs) is essential for cellular life. Genes are often transcribed by multiple RNAPs. While the properties of individual RNAPs are well appreciated, it remains less explored whether group behaviors can emerge from co-transcribing RNAPs under most physiological levels of gene expression. Here, we provide evidence in Escherichia coli that well-separated RNAPs can exhibit collaborative and antagonistic group dynamics. Co-transcribing RNAPs translocate faster than a single RNAP, but the density of RNAPs has no significant effect on their average speed. When a promoter is inactivated, RNAPs that are far downstream from the promoter slow down and experience premature dissociation, but only in the presence of other co-transcribing RNAPs. These group behaviors depend on transcription-induced DNA supercoiling, which can also mediate inhibitory dynamics between RNAPs from neighboring divergent genes. Our findings suggest that transcription on topologically-constrained DNA, a norm across organisms, can provide an intrinsic mechanism for modulating the speed and processivity of RNAPs over long distances according to the promoters on/off state.

microbiology

Learning Gene Networks Underlying Clinical Phenotypes Using SNP Perturbations

Recent technologies are generating an abundance of genome sequence data and molecular and clinical phenotype data, providing an opportunity to understand the genetic architecture and molecular mechanisms underlying diseases. Previous approaches have largely focused on the co-localization of single-nucleotide polymorphisms (SNPs) associated with clinical and expression traits, each identified from genome-wide association studies and expression quantitative trait locus (eQTL) mapping, and thus have provided only limited capabilities for uncovering the molecular mechanisms behind the SNPs influencing clinical phenotypes. Here we aim to extract rich information on the functional role of trait-perturbing SNPs that goes far beyond this simple co-localization. We introduce a computational framework called Perturb-Net for learning the gene network that modulates the influence of SNPs on phenotypes, using SNPs as naturally occurring perturbation of a biological system. Perturb-Net uses a probabilistic graphical model to directly model both the cascade of perturbation from SNPs to the gene network to the phenotype network and the network at each layer of molecular and clinical phenotypes. Perturb-Net learns the entire model by solving a single optimization problem with an extremely fast algorithm that can analyze human genome-wide data within a few hours. In our analysis of asthma data, for a locus that was previously implicated in asthma susceptibility but for which little is known about the molecular mechanism underlying the association, Perturb-Net revealed the gene network modules that mediate the influence of the SNP on asthma phenotypes. Many genes in this network module were well supported in the literature as asthma-related.

bioinformatics

Identifying functional targets from transcription factor binding data using SNP perturbation

Transcription factors (TFs) play a key role in transcriptional regulation by binding to DNA to initiate the transcription of target genes. Techniques such as ChIP-seq and DNase-seq provide a genome-wide map of TF binding sites but do not offer direct evidence that those bindings affect gene expression. Thus, these assays are often followed by TF perturbation experiments to determine functional binding that leads to changes in target gene expression. However, such perturbation experiments are costly and time-consuming, and have a well-known limitation that they cannot distinguish between direct and indirect targets. In this study, we propose to use the naturally occurring perturbation of gene expression by genetic variation captured in population SNP and expression data to determine functional targets from TF binding data. We introduce a computational methodology based on probabilistic graphical models for isolating the perturbation effect of each individual SNP, given a large number of SNPs across genomes perturbing the expression of all genes simultaneously. Our computational approach constructs a gene regulatory network over TFs, their functional targets, and further downstream genes, while at the same time identifying the SNPs perturbing this network. Compared to experimental perturbation, our approach has advantages of identifying direct and indirect targets, and leveraging existing data collected for expression quantitative trait locus mapping, a popular approach for studying the genetic architecture of expression. We apply our approach to determine functional targets from the TF binding data for a lymphoblastoid cell line from the ENCODE Project, using SNP and expression data from the HapMap 3 and 1000 Genomes Project samples. Our results show that from TF binding data, functional target genes can be determined by SNP perturbation of various aspects that impact transcriptional regulation, such as TF concentration and TF-DNA binding affinity.

bioinformatics

Protect TUDCA stimulated CKD-derived hMSCs against the CKD-Ischemic disease via upregulation of PrPC

Although autologous human mesenchymal stem cells (hMSCs) are a promising source for regenerative stem cell therapy, the barriers associated with pathophysiological conditions in this disease limit therapeutic applicability to patients. We proved treatment of CKD-hMSCs with TUDCA enhanced the mitochondrial function of these cells and increased complex I & IV enzymatic activity, increasing PINK1 expression and decreasing mitochondrial O2*- and mitochondrial fusion in a PrPC-dependent pathway. Moreover, TH-1 cells enhanced viability when co-cultured in vitro with TUDCA-treated CKD-hMSC. In vivo, tail vein injection of TUDCA-treated CKD-hMSCs into the mouse model of CKD associated with hindlimb ischemia enhanced kidney recovery, the blood perfusion ratio, vessel formation, and prevented limb loss, and foot necrosis along with restored expression of PrPC in the blood serum of the mice. These data suggest that TUDCA-treated CKD-hMSCs are a promising new autologous stem cell therapeutic intervention that dually treats cardiovascular problems and CKD in patients.

cell biology

A combination of transcription factors mediates inducible interchromosomal pairing

Remodeling of the three-dimensional organization of a genome has been previously described (e.g. condition-specific pairing or looping), but it remains unknown which factors specify and mediate such shifts in chromosome conformation. Here we describe an assay, MAP-C (Mutation Analysis in Pools by Chromosome conformation capture), that enables the simultaneous characterization of hundreds of cis or trans-acting mutations for their effects on a chromosomal contact or loop. As a proof of concept, we applied MAP-C to systematically dissect the molecular mechanism of inducible interchromosomal pairing between HAS1pr-TDA1pr alleles in Saccharomyces yeast. We identified three transcription factors, Leu3, Sdd4 (Ypr022c), and Rgt1, whose collective binding to nearby DNA sequences is necessary and sufficient for inducible pairing between binding site clusters. Rgt1 contributes to the regulation of pairing, both through changes in expression level and through its interactions with the Tup1/Ssn6 repressor complex. HAS1pr-TDA1pr is the only locus with a cluster of binding site motifs for all three factors in both S. cerevisiae and S. uvarum genomes, but the promoter for HXT3, which contains Leu3 and Rgt1 motifs, also exhibits inducible homolog pairing. Altogether, our results demonstrate that specific combinations of transcription factors can mediate condition-specific interchromosomal contacts, and reveal a molecular mechanism for interchromosomal contacts and mitotic homolog pairing.

genomics

Synthetic chromosome fusion: effects on genome structure and function

As part of the Synthetic Yeast 2.0 (Sc2.0) project, we designed and synthesized synthetic chromosome I. The total length of synI is [~]21.4% shorter than wild-type chromosome I, the smallest chromosome in Saccharomyces cerevisiae. SynI was designed for attachment to another synthetic chromosome due to concerns of potential instability and karyotype imbalance. We used a variation of a previously developed, robust CRISPR-Cas9 method to fuse chromosome I to other chromosome arms of varying length: chrIXR (84kb), chrIIIR (202kb) and chrIVR (1Mb). All fusion chromosome strains grew like wild-type so we decided to attach synI to synIII. Through the investigation of three-dimensional structures of fusion chromosome strains, unexpected loops and twisted structures were formed in chrIII-I and chrIX-III-I fusion chromosomes, which depend on silencing protein Sir3. These results suggest a previously unappreciated 3D interaction between HMR and the adjacent telomere. We used these fusion chromosomes to show that axial element Red1 binding in meiosis is not strictly chromosome size dependent even though Red1 binding is enriched on the three smallest chromosomes in wild-type yeast, and we discovered an unexpected role for centromeres in Red1 binding patterns.

synthetic biology

Web-based design and analysis tools for CRISPR base editing

BackgroundAs a result of its simplicity and high efficiency, the CRISPR-Cas system has been widely used as a genome editing tool. Recently, CRISPR base editors, which consist of deactivated Cas9 (dCas9) or Cas9 nickase (nCas9) linked with a cytidine or a guanine deaminase, have been developed. Base editing tools will be very useful for gene correction because they can produce highly specific DNA substitutions without the introduction of any donor DNA, but dedicated web-based tools to facilitate the use of such tools have not yet been developed.\n\nResultsWe present two web tools for base editors, named BE-Designer and BE-Analyzer. BE-Designer provides all possible base editor target sequences in a given input DNA sequence with useful information including potential off-target sites. BE-Analyzer, a tool for assessing base editing outcomes from next generation sequencing (NGS) data, provides information about mutations in a table and interactive graphs. Furthermore, because the tool runs client-side, large amounts of targeted deep sequencing data (>100MB) do not need to be uploaded to a server, substantially reducing running time and increasing data security. BE-Designer and BE-Analyzer can be freely accessed at http://www.rgenome.net/bedesigner/ and http://www.rgenome.net/be-analyzer/respectively\n\nConclusionWe develop two useful web tools to design target sequence (BE-Designer) and to analyze NGS data from experimental results (BE-Analyzer) for CRISPR base editors.

bioinformatics

The structural basis for activation of voltage sensor domains in an ion channel TPC1

Voltage sensing domains (VSDs) couple changes in transmembrane electrical potential to conformational changes that regulate ion conductance through a central channel. Positively charged amino acids inside each sensor cooperatively respond to changes in voltage. Our previous structure of a TPC1 channel captured the first example of a resting-state VSD in an intact ion channel. To generate an activated state VSD in the same channel we removed the luminal inhibitory Ca2+-binding site (Cai2+), that shifts voltage-dependent opening to more negative voltage and activation at 0 mV. Cryo-EM reveals two coexisting structures of the VSD, an intermediate state 1 that partially closes access to the cytoplasmic side, but remains occluded on the luminal side and an intermediate activated state 2 in which the cytoplasmic solvent access to the gating charges closes, while luminal access partially opens. Activation can be thought of as moving a hydrophobic insulating region of the VSD from the external side, to an alternate grouping on the internal side. This effectively moves the gating charges from the inside potential to that of the outside. Activation also requires binding of Ca2+ to a cytoplasmic site (Caa2+). An X-ray structure with Caa2+ removed and a near-atomic resolution cryo-EM structure with Cai2+ removed define how dramatic conformational changes in the cytoplasmic domains may communicate with the VSD during activation. Together four structures provide a basis for understanding the voltage dependent transition from resting to activated state, the tuning of VSD by thermodynamic stability, and this channels requirement of cytoplasmic Ca2+-ions for activation.

biophysics

TGFam-Finder: An optimal solution for target-gene family annotation in eukaryotic genomes

Whole genome annotation errors that omit essential protein-coding genes hinder further research. We developed Target Gene Family Finder (TGFam-Finder), an optimal tool for structural annotation of protein-coding genes containing target domain(s) of interest in eukaryotic genomes. Large-scale re-annotation of 100 publicly available eukaryotic genomes led to the discovery of essential genes that were missed in previous annotations. An average of 117 (346%) and 148 (45%) additional FAR1 and NLR genes were newly identified in 50 plant genomes. Furthermore, 117 (47%) additional C2H2 zinc finger genes were detected in 50 animal genomes including human and mouse. Accuracy of the newly annotated genes was validated by RT-PCR and cDNA sequencing in human, mouse and rice. In the human genome, 26 newly annotated genes were identical with known functional genes. TGFam-Finder along with the new gene models provide an optimized platform for unbiased functional and comparative genomics and comprehensive evolutionary study in eukaryotes.

bioinformatics

High-throughput retrieval of physical DNA for NGS-identifiable clones in phage display library

In antibody discovery, in-depth analysis of an antibody library and high-throughput retrieval of clones in the library are crucial to identifying and exploiting rare clones with different properties. However, existing methods have several technical limitations such as low process throughput from laborious cloning process and waste of the phenotypic screening capacity from unnecessary repetitive tests on the dominant clones. To overcome the limitations, we developed a new high-throughput platform for the identification and retrieval of clones in the library, TrueRepertoire. TrueRepertoire provides highly accurate sequences of the clones with linkage information between heavy and light chains of the antibody fragment. Additionally, the physical DNA of clones can be retrieved in high throughput based on the sequence information. We validated the high accuracy of the sequences and demonstrated that there is no platform-specific bias. Moreover, the applicability of TrueRepertoire was demonstrated by a phage-displayed single-chain variable fragment (scFv) library targeting human hepatocyte growth factor (hHGF) protein.

bioengineering

Addition of Degenerate Bases to DNA-based Data Storage for Increased Information Capacity

Introductory paragraphDNA-based data storage has emerged as a promising method to satisfy the exponentially increasing demand for information storage. However, practical implementation of DNA-based data storage remains a challenge because of the high cost of DNA per unit data. Here, we propose the use of eleven degenerate bases as encoding characters in addition to A, C, G, and T, which increases the information capacity (the amount of data that can be stored per length of DNA sequence designed) and reduce the cost of DNA per unit data. Using the proposed method, we experimentally achieved an information capacity of 3.37 bits/character, which is more than twice when compared to the highest information capacity previously achieved. Finally, the platform was projected to reduce the cost of DNA-based data storage by 50%.

synthetic biology

Joint analysis of matched tumor samples with varying tumor contents improves somatic variant calling in the absence of a germline sample

Archival tumor samples represent a potential rich resource of annotated specimens for translational genomics research. However, standard variant calling approaches require a matched normal sample from the same individual, which is often not available in the retrospective setting, making it difficult to distinguish between true somatic variants and germline variants that are private to the individual. Archival sections often contain adjacent normal tissue, but this normal tissue can include infiltrating tumor cells. Comparative somatic variant callers are designed to exclude variants present in the normal sample, so a novel approach is required to leverage sequencing of adjacent normal tissue for somatic variant calling. Here we present LumosVar 2.0, a software package designed to jointly analyze multiple samples from the same patient. The approach is based on the concept that the allelic fraction of somatic variants, but not germline variants, would be reduced in samples with low tumor content. LumosVar 2.0 estimates allele specific copy number and tumor sample fractions from the data, and uses the model to determine expected allelic fractions for somatic and germline variants and classify variants accordingly. To evaluate using LumosVar 2.0 to jointly call somatic variants with tumor and adjacent normal samples, we used a glioblastoma dataset with matched high tumor content, low tumor content, and germline exome sequencing data (to define true somatic variants) available for each patient. We show that both sensitivity and positive predictive value are improved by analyzing the high tumor and low tumor samples jointly compared to analyzing the samples individually or compared to in-silico pooling of the two samples. Finally, we applied this approach to a set of breast and prostate archival tumor samples for which normal samples were not available for germline sequencing, but tumor blocks containing adjacent normal tissue were available for sequencing. Joint analysis using LumosVar 2.0 detected several variants, including known cancer hotspot mutations that were not detected by standard somatic variant calling tools using the adjacent normal as a reference. Together, these results demonstrate the potential utility of leveraging paired tissue samples to improve somatic variant calling when a constitutional DNA sample is not available.

bioinformatics

PTK2 regulates the UPS impairment via p62 phosphorylation in TDP-43 proteinopathy

TDP-43 proteinopathy is a common feature in a variety of neurodegenerative disorders including Amyotrophic lateral sclerosis (ALS) cases, Frontotemporal lobar degeneration (FTLD), and Alzheimers disease. However, the molecular mechanisms underlying TDP-43-induced neurotoxicity are largely unknown. In this study, we demonstrated that TDP-43 proteinopathy induces impairment in ubiquitin-proteasome system (UPS) evidenced by an accumulation of ubiquitinated proteins and reduction of proteasome activity in neuronal cells. Through kinase inhibitor screening, we identified PTK2 as a suppressor of neurotoxicity induced by UPS impairment. Importantly, PTK2 inhibition significantly reduces ubiquitin aggregates and attenuated TDP-43-induced cytotoxicity in Drosophila model of TDP-43 proteinopathy. We further identified that phosphorylation of p62 at serine 403 (p-p62S403), a key component in the autophagic degradation of poly-ubiquitinated proteins, is increased upon TDP-43 overexpression and dependent on activation of PTK2 in neuronal cells. Moreover, expressing a non-phosphorylated form of p62 (p62S403A) significantly represses accumulation of polyubiquitinated proteins and neurotoxicity induced by TDP-43 overexpression in neuronal cells. In addition, inhibition of TBK1, a kinase which phosphorylates S403 of p62, ameliorates neurotoxicity upon UPS impairment in neuronal cells. Taken together, our data suggest that activation of PTK2-TBK1-p62 axis plays a critical role in the pathogenesis of TDP-43 by regulating neurotoxicity induced by UPS impairment. Therefore, targeting PTK2-TBK1-p62 axis may represent a novel therapeutic intervention for neurodegenerative diseases with TDP-43 proteinopathy.

neuroscience

Modeling large fluctuations of thousands of clones during hematopoiesis: the role of stem cell self-renewal and bursty progenitor dynamics in rhesus macaque

In a recent clone-tracking experiment, millions of uniquely tagged hematopoietic stem cells (HSCs) were autologously transplanted into rhesus macaques and peripheral blood containing thousands of tags were sampled and sequenced over 14 years to quantify the abundance of hundreds to thousands of tags or \"clones.\" Two major puzzles of the data have been observed: consistent differences and massive temporal fluctuations of clone populations. The large sample-to-sample variability can lead clones to occasionally go \"extinct\" but \"resurrect\" themselves in subsequent samples. Although heterogeneity in HSC differentiation rates, potentially due to tagging, and random sampling of the animals blood and cellular demographic stochasticity might be invoked to explain these features, we show that random sampling cannot explain the magnitude of the temporal fluctuations. Moreover, we show through simpler neutral mechanistic and statistical models of hematopoiesis of tagged cells that a broad distribution in clone sizes can arise from stochastic HSC self-renewal instead of tag-induced heterogeneity. The very large clone population fluctuations that often lead to extinctions and resurrections can be naturally explained by a generation-limited proliferation constraint on the progenitor cells. This constraint leads to bursty cell population dynamics underlying the large temporal fluctuations. We analyzed experimental clone abundance data using a new statistic that counts clonal disappearances and provide least-squares estimates of two key model parameters in our model, the total HSC differentiation rate and the maximum number of progenitor-cell divisions.\n\nAuthor summaryHematopoiesis of virally tagged cells in rhesus macaques is analyzed in the context of a mechanistic and statistical model. We find that the clone size distribution and the temporal variability in the abundance of each clone (viral tag) in peripheral blood are consistent with (i) stochastic HSC self-renewal during bone marrow repair, (ii) clonal aging that restricts the number of generations of progenitor cells, and (iii) infrequent and small-size samples. By fitting data, we infer two key parameters that control the level of fluctuations of clone sizes in our model: the total HSC differentiation rate and the maximum proliferation capacity of progenitor cells. Our analysis provides insight into the mechanisms of hematopoiesis and a framework to guide future multiclone barcoding/lineage tracking measurements.

systems biology