Search bioRxivSearch

Biology subjects

Xu, C.

Publications and source records attributed to Xu, C..

At least 19 recordsLinked to original sources

Multi-Center Study of Resectable Lung Lesions by Ultra-Deep Sequencing of Targeted Genes in Plasma Cell-Free DNA to Assess Nodule Malignancy and Detect Lung Cancers

BACKGROUNDEarly detection of lung cancer to allow curative treatment remains challenging. Cell-free circulating tumor DNA (ctDNA) analysis may aid in malignancy assessment and early cancer diagnosis of lung nodules found in screening imagery.\n\nMETHODSThe multi-center clinical study enrolled 192 patients with operable occupying lung diseases. Plasma ctDNA, white blood cell genomic DNA (gDNA) and tumor tissue gDNA of each patient were analyzed by ultra-deep sequencing to an average of 35,000X of the coding regions of 65 lung cancer-related genes.\n\nRESULTSThe cohort consists of a quarter of benign lung diseases and three quarters of cancer patients with all histopathology subtypes. 64% of the cancer patients is at Stage I. Gene mutations detection in tissue gDNA and plasma ctDNA results in a sensitivity of 91% and specificity of 88%. When ctDNA assay was used as the test, the sensitivity was 69% and specificity 96%. As for the lung cancer patients, the assay detected 63%, 83%, 94% and 100%, for Stage I, II, III and IV, respectively. In a linear discriminant analysis, combination of ctDNA, patient age and a panel of serum biomarkers boosted the overall sensitivity to 80% at a specificity of 99%. 29 out of the 65 genes harbored mutations in the lung cancer patients with the largest number found in TP53 (30% plasma and 62% tumor tissue samples) and EGFR (20% and 40%, respectively).\n\nCONCLUSIONPlasma ctDNA was analyzed in lung nodule assessment and early cancer detection while an algorithm combining clinical information enhanced the test performance.

clinical trials

The mitochondrial DNA content can not predict the embryo viability

ObjectiveTo investigate whether the mitochondrial DNA content could predict the embryo viability\n\nDesignRetrospective analysis.\n\nSettingReproductive genetics laboratory\n\nPatient(s)A total of 421 biopsied samples obtained from 129 patients\n\nIntervention(s)Embryo biopsies samples underwent whole genome amplification (WGA) and were tested by next generation sequencing (NGS) and array Comparative Genomic Hybridization (aCGH), 30 samples were selected randomly to undergo quantitative real-time polymerase chain reaction (qPCR).\n\nMain Outcome Measure(s)Those embryos which obtained the consistent chromosome status determined both aCGH and NGS platform were further classified. We investigated the relationship of mtDNA content with several factors including female patient age, embryo morphology, chromosome status, and live birth rate of both blastocysts and blastomeres.\n\nResult(s)A total of 386 (110 blastomeres and 276 blastocysts) out of 399 embryos showed consistent chromosome status outcome. We found no statistically difference was observed in aneuploid and euploid blastocysts (p=0.14), the same phenomenon was observed in aneuploid and euploid blastomeres (p=0.89). Similarly, the mtDNA content was independent of female patient age, embryo morphology and live birth rate.\n\nConclusion(s)The mtDNA content did not provide a reliable prediction of the viability of blastocysts to initiate a pregnancy.

ecology

Genome structure and evolution of Antirrhnum majus L.

Snapdragon (Antirrhinum majus L.), a member of Plantaginaceae, is an important model for plant genetics and molecular studies on plant growth and development, transposon biology and self-incompatibility. Here we report a high-quality genome assembly of A. majus cultivated JI7 (A. majus cv.JI7) of a 510 Mb with 37,714 annotated protein-coding genes. The scaffolds covering 97.12% of the assembled genome were anchored on 8 chromosomes. Comparative and evolutionary analyses revealed that Plantaginaceae and Solanaceae diverged from their most recent ancestor around 62 million years ago (MYA). We also revealed the genetic architectures associated with complex traits such as flower asymmetry and self-incompatibility including a unique TCP duplication around 46-49 MYA and a near complete{psi} S-locus of ca.2 Mb. The genome sequence obtained in this study not only provides the first genome sequenced from Plantaginaceae but also bring the popular plant model system of Antirrhinum into a genomic age.

genomics

Adaptive disinhibitory gating by VIP interneurons permits associative learning

Learning drives behavioral adaptations necessary for survival. While plasticity of excitatory projection neurons during associative learning is studied extensively, little is known about the contributions of local interneurons. Using fear conditioning as a model for associative learning, we find that behaviorally relevant, salient stimuli cause learning by tapping into a local microcircuit consisting of precisely connected subtypes of inhibitory interneurons. By employing calcium imaging and optogenetics, we demonstrate that vasoactive intestinal peptide (VIP)-expressing interneurons in the basolateral amygdala are activated by aversive events and provide an instructive disinhibitory signal for associative learning. Notably, VIP interneuron responses are plastic and shift from the instructive to the predictive cue upon memory formation. We describe a novel form of adaptive disinhibitory gating by VIP interneurons that allows to discriminate unexpected, important from irrelevant information, and might be a general dynamic circuit motif to trigger stimulus-specific learning, thereby ensuring appropriate behavioral adaptations to salient events.

neuroscience

MxB restricts HIV-1 by targeting the tri-hexamer interface of the viral capsid

Myxovirus resistance protein B (MxB) is an interferon-inducible restriction factor of HIV-1 that blocks nuclear import of the viral genome. Evidence suggests that MxB recognizes higher-order interfaces of the HIV capsid lattice, but the mechanistic details of this interaction are not known. Previous studies have mapped the restriction activity of MxB to its N-terminus encompassing a triple arginine motif 11RRR13. Here we demonstrate a direct and specific interaction between the MxB N-terminus and helical assemblies of HIV-1 capsid protein (CA) using highly purified recombinant proteins. We performed thorough mutagenesis to establish the detailed molecular requirements for the CA interaction with MxB. The results map MxB binding to the interface of three CA hexamers, specifically interactions between positively charged MxB N-terminal residues and negatively charged CA residues. Our crystal structures show that the CA mutations affecting MxB interaction and restriction do not alter the conformation of capsid assembly. In addition, 30 microsecond long all-atom molecular dynamics (MD) simulations of the complex between the MxB N-terminus and the HIV CA tri-hexamer interface show persistent MxB binding and identify a MxB-binding pocket surrounded by three CA hexamers. These results establish the molecular details of the binding of a lattice-sensing host factor onto HIV capsid, and provide insight into how MxB recognizes HIV capsid for the restriction of HIV-1 infection.\n\nAuthor summaryThe human antiviral protein MxB is a restriction factor that fights HIV infection. Previous experiments have demonstrated that MxB targets the HIV capsid, a protein shell that protects the viral genome. To make the conical shaped capsid, HIV CA proteins are organized into a lattice composed of hexamer and pentamer building blocks, providing many interfaces for host proteins to recognize. Through extensive biochemical and biophysical studies and molecular dynamics simulations, we show that MxB is targeting the HIV capsid by recognizing the region created at the intersection of three CA hexamers. We are further able to map this interaction to a few CA residues, located in a negatively-charged well at the interface between the three CA hexamers. This work provides detailed residue-level mapping of the targeted capsid interface and how MxB interacts. This information could inspire the development of capsid-targeting therapies for HIV.

biochemistry

Control of synaptic specificity by limiting promiscuous synapse formation

The ability of neurons to distinguish appropriate from inappropriate synaptic partners in their local environment is fundamental to the proper assembly and function of neural circuits. How synaptic partner selection is regulated is a longstanding question in Neurobiology. A prevailing hypothesis is that appropriate partners express complementary molecules that match them together and promote synaptogenesis. Dpr and DIP IgSF proteins bind heterophilically and are expressed in a complementary manner between synaptic partners in the Drosophila visual system. Here, we show that in the lamina, DIP mis-expression is sufficient to promote synapse formation with Dpr-expressing neurons, and that DIP proteins are not necessary for synaptogenesis but rather function to prevent ectopic synapse formation. These findings indicate that Dpr-DIP interactions regulate synaptic specificity by biasing synapse formation towards specific cell-types. We propose that synaptogenesis occurs independent of synaptic partner choice, and that precise synaptic connectivity is established by limiting promiscuous synapse formation.

neuroscience

A cell atlas of the adult Drosophila midgut

Studies of the adult Drosophila midgut have provided a number of insights on cell type diversity, stem cell regeneration, tissue homeostasis and cell fate decision. Advances in single-cell RNA sequencing (scRNA-seq) provide opportunities to identify new cell types and molecular features. We used inDrop to characterize the transcriptome of midgut epithelial cells and identified 12 distinct clusters representing intestinal stem cells (ISCs), enteroblasts (EBs), enteroendocrine cells (EEs), enterocytes (ECs) from different regions, and cardia. This unbiased approach recovered 90% of the known ISCs/EBs markers, highlighting the high quality of the dataset. Gene set enrichment analysis in conjunction with electron micrographs revealed that ISCs are enriched in free ribosomes and possess mitochondria with fewer cristae. We demonstrate that a subset of EEs in the middle region of the midgut expresses the progenitor marker esg and that individual EEs are capable of expressing up to 4 different gut hormone peptides. We also show that the transcription factor klumpfuss (klu) is expressed in EBs and functions to suppress EE formation. Lastly, we provide a web-based resource for visualization of gene expression in single cells. Altogether, our study provides a comprehensive resource for addressing novel functions of genes in the midgut epithelium.

genetics

Diagnosis and Prognosis Using Machine Learning Trained on BrainMorphometry and White Matter Connectomes

Accurate, reliable prediction of risk for Alzheimers disease (AD) is essential for early, disease-modifying therapeutics. Multimodal MRI, such as structural and diffusion MRI, is likely to contain complementary information of neurodegenerative processes in AD. Here we tested the utility of the multimodal MRI (T1-weighted structure and diffusion MRI), combined with high-throughput brain phenotyping--morphometry and structural connectomics--and machine learning, as a diagnostic tool for AD. We used, firstly, a clinical cohort at a dementia clinic (National Health Insurance Service-Ilsan Hospital [NHIS-IH]; N=211; 110 AD, 64 mild cognitive impairment [MCI], and 37 cognitively normal with subjective memory complaints [SMC]) to test the diagnostic models; and, secondly, Alzheimers Disease Neuroimaging Initiative (ADNI)-2 to test the generalizability. Our machine learning models trained on the morphometric and connectome estimates (number of features=34,646) showed optimal classification accuracy (AD/SMC: 97% accuracy, MCI/SMC: 83% accuracy; AD/MCI: 97% accuracy) in NHIS-IH cohort, outperforming a benchmark model (FLAIR-based white matter hyperintensity volumes). In ADNI-2 data, the combined connectome and morphometry model showed similar or superior accuracies (AD/HC: 96%; MCI/HC: 70%; AD/MCI: 75% accuracy) compared with the CSF biomarker model (t-tau, p-tau, and Amyloid {beta}, and ratios). In predicting MCI to AD progression in a smaller cohort of ADNI-2 (n=60), the morphometry model showed similar performance with 69% accuracy compared with CSF biomarker model with 70% accuracy. Our comparison of classifiers trained on structural MRI, diffusion MRI, FLAIR, and CSF biomarkers show the promising utility of the white matter structural connectomes in classifying AD and MCI in addition to the widely used structural MRI-based morphometry, when combined with machine learning.\n\nHighlightsO_LIWe showed the utility of multimodal MRI, combining morphometry and white matter connectomes, to classify the diagnosis of AD and MCI using machine learning.\nC_LIO_LIIn predicting the progression from MCI to AD, the morphometry model showed the best performance.\nC_LIO_LITwo independent clinical datasets were used in this study: one for model building, the other for generalizability testing.\nC_LI

neuroscience

SymSim: simulating multi-faceted variability in single cell RNA sequencing

The abundance of new computational methods for processing and interpreting transcriptomes at a single cell level raises the need for in-silico platforms for evaluation and validation. Simulated datasets which resemble the properties of real datasets can aid in method development and prioritization as well as in questions in experimental design by providing an objective ground truth. Here, we present SymSim, a simulator software that explicitly models the processes that give rise to data observed in single cell RNA-Seq experiments. The components of the SymSim pipeline pertain to the three primary sources of variation in single cell RNA-Seq data: noise intrinsic to the process of transcription, extrinsic variation that is indicative of different cell states (both discrete and continuous), and technical variation due to low sensitivity and measurement noise and bias. Unlike other simulators, the parameters that govern the simulation process directly represent meaningful properties such as mRNA capture rate, the number of PCR cycles, sequencing depth, or the use of unique molecular identifiers. We demonstrate how SymSim can be used for benchmarking methods for clustering and differential expression and for examining the effects of various parameters on their performance. We also show how SymSim can be used to evaluate the number of cells required to detect a rare population and how this number deviates from the theoretical lower bound as the quality of the data decreases. SymSim is publicly available as an R package and allows users to simulate datasets with desired properties or matched with experimental data.

bioinformatics

Inference of Chromosome-length Haplotypes Using Genomic Data of Three to Five Single Gametes

Knowledge of chromosome-length haplotypes will not only advance our understanding of the relationship between DNA and phenotypes, but also promote a variety of genetic applications. Here we present Hapi, an innovative method for chromosomal haplotype inference using only 3 to 5 gametes. Hapi outperformed all existing haploid-based phasing methods in terms of accuracy, reliability, and cost efficiency in both simulated and real gamete datasets. This highly cost-effective phasing method will make large-scale haplotype studies feasible to facilitate human disease studies and plant/animal breeding. In addition, Hapi can detect meiotic crossovers in gametes, which has promise in the diagnosis of abnormal recombination activity in human reproductive cells.

bioinformatics

Synthesis of three major auxins from glucose in Engineered Escherichia coli

Indole-3-acetic acid (IAA) is considered the most common and important naturally occurring auxin in plants and a major regulator of plant growth and development. In addition, phenylacetic acid (PAA) and 4-hydroxyphenylacetic acid (4HPA) can also play a role as auxin in some plants. In recent years, several microbes have been metabolically engineered to produce IAA from L-tryptophan. In this study, we showed that aminotransferase aro8 and decarboxylase kdc from Saccharomyces cerevisiae, and aldehyde dehydrogenase aldH from Escherichia coli have broad substrate ranges and can catalyze the conversion of three kinds of aromatic amino acids (L-tryptophan, L-tyrosine or L-phenylalanine) to the corresponding IAA, 4HPA and PAA. Subsequently, three de novo biosynthetic pathways for the production of IAA, PAA and 4HPA from glucose were constructed in E. coli through strengthening the shikimate pathway. This study described here shows the way for the development of agricultural microorganism for biosynthesis of plant auxin and promoting plant growth in the future.

bioengineering

Single tube bead-based DNA co-barcoding for cost effective and accurate sequencing, haplotyping, and assembly

Obtaining accurate sequences from long DNA molecules is very important for genome assembly and other applications. Here we describe single tube long fragment read (stLFR), a technology that enables this a low cost. It is based on adding the same barcode sequence to sub-fragments of the original long DNA molecule (DNA co-barcoding). To achieve this efficiently, stLFR uses the surface of microbeads to create millions of miniaturized barcoding reactions in a single tube. Using a combinatorial process up to 3.6 billion unique barcode sequences were generated on beads, enabling practically non-redundant co-barcoding with 50 million barcodes per sample. Using stLFR, we demonstrate efficient unique co-barcoding of over 8 million 20-300 kb genomic DNA fragments. Analysis of the genome of the human genome NA12878 with stLFR demonstrated high quality variant calling and phasing into contigs up to N50 34 Mb. We also demonstrate detection of complex structural variants and complete diploid de novo assembly of NA12878. These analyses were all performed using single stLFR libraries and their construction did not significantly add to the time or cost of whole genome sequencing (WGS) library preparation. stLFR represents an easily automatable solution that enables high quality sequencing, phasing, SV detection, scaffolding, cost-effective diploid de novo genome assembly, and other long DNA sequencing applications.

genomics

Marginal Effects of Systemic CCR5 Blockade with Maraviroc on Oral Simian Immunodeficiency Virus Transmission to Infant Macaques

Current approaches do not eliminate all HIV-1 maternal-to-infant transmissions (MTIT); new prevention paradigms might help avert new infections. We administered Maraviroc (MVC) to rhesus macaques (RMs) to block CCR5-mediated entry, followed by repeated oral exposure of a CCR5-dependent clone of simian immunodeficiency virus (SIV)mac251 (SIVmac766). MVC significantly blocked the CCR5 coreceptor in peripheral blood mononuclear cells and tissue cells. All control animals and 60% of MVC-treated infant RMs became infected by the 6th challenge, with no significant difference between the number of exposures (p=0.15). At the time of viral exposures, MVC plasma and tissue (including tonsil) concentrations were within the range seen in humans receiving MVC as a therapeutic. Both treated and control RMs were infected with only a single transmitted/founder variant, consistent with the dose of virus typical of HIV-1 infection. The uninfected RMs expressed the lowest levels of CCR5 on the CD4+ T cells. Ramp-up viremia was significantly delayed (p=0.05) in the MVC-treated RMs, yet peak and postpeak viral loads were similar in treated and control RMs. In conclusion, in spite of apparent effective CCR5 blockade in infant RMs, MVC had marginal impact on acquisition and only a minimal impact on post infection delay of viremia following oral SIV infection. Newly developed, more effective CCR5 blockers may have a more dramatic impact on oral SIV transmission than MVC.\n\nImportanceWe have previously suggested that the very low levels of simian immunodeficiency virus (SIV) maternal-to-infant transmissions (MTIT) in African nonhuman primates that are natural hosts of SIVs are due to a low availability of target cells (CCR5+ CD4+ T cells) in the oral mucosa of the infants, rather than maternal and milk factors. To confirm this new MTIT paradigm, we performed a proof of concept study, in which we therapeutically blocked CCR5 with maraviroc (MVC) and orally exposed MVC treated and naive infant rhesus macaques to SIV. MVC had only a marginal effect on oral SIV transmission. However, the observation that the infant RMs that remained uninfected at the completion of the study, after 6 repeated viral challenges, had the lowest CCR5 expression on the CD4+ T cells prior to the MVC treatment, appear to confirm our hypothesis, also suggesting that the partial effect of MVC is due to a limited efficacy of the drug. Newly, more effective CCR5 inhibitors may have a better effect in preventing SIV and HIV transmission.

immunology

smCounter2: an accurate low-frequency variant caller for targeted sequencing data with unique molecular identifiers

MotivationLow-frequency DNA mutations are often confounded with technical artifacts from sample preparation and sequencing. With unique molecular identifiers (UMIs), most of the sequencing errors can be corrected. However, errors before UMI tagging, such as DNA polymerase errors during end-repair and the first PCR cycle, cannot be corrected with single-strand UMIs and impose fundamental limits to UMI-based variant calling.\n\nResultsWe developed smCounter2, a UMI-based variant caller for targeted sequencing data and an upgrade from the current version of smCounter. Compared to smCounter, smCounter2 features lower detection limit at 0.5%, better overall accuracy (particularly in non-coding regions), a consistent threshold that can be applied to both deep and shallow sequencing runs, and easier use via a Docker image and code for read pre-processing. We benchmarked smCounter2 against several state-of-the-art UMI-based variant calling methods using multiple datasets and demonstrated smCounter2s superior performance in detecting somatic variants. At the core of smCounter2 is a statistical test to determine whether the allele frequency of the putative variant is significantly above the background error rate, which was carefully modeled using an independent dataset. The improved accuracy in non-coding regions was mainly achieved using novel repetitive region filters that were specifically designed for UMI data.\n\nAvailabilityThe entire pipeline is available at https://github.com/qiaseq/qiaseq-dna under MIT license.

bioinformatics

Accurate Prediction of Alzheimer’s Disease Using Multi-Modal MRI and High-Throughput Brain Phenotyping

Accurate, reliable prediction of risk for Alzheimers disease (AD) is essential for early, disease-modifying therapeutics. Multimodal MRI, such as structural and diffusion MRI, is likely to contain complementary information of neurodegenerative processes in AD. Here we tested the utility of commonly available multimodal MRI (T1-weighted structure and diffusion MRI), combined with high-throughput brain phenotyping--morphometry and connectomics--and machine learning, as a diagnostic tool for AD. We used, firstly, a clinical cohort at a dementia clinic (study 1: Ilsan Dementia Cohort; N=211; 110 AD, 64 mild cognitive impairment [MCI], and 37 subjective memory complaints [SMC]) to test and validate the diagnostic models; and, secondly, Alzheimers Disease Neuroimaging Initiative (ADNI)-2 (study 2) to test the generalizability of the approach and the prognostic models with longitudinal follow up data. Our machine learning models trained on the morphometric and connectome estimates (number of features=34,646) showed optimal classification accuracy (AD/SMC: 97% accuracy, MCI/SMC: 83% accuracy; AD/MCI: 97% accuracy) with iterative nested cross-validation in a single-site study, outperforming the benchmark model (FLAIR-based white matter hyperintensity volumes). In a generalizability study using ADNI-2, the combined connectome and morphometry model showed similar or superior accuracies (AD/HC: 96%; MCI/HC: 70%; AD/MCI: 75% accuracy) as CSF biomarker model (t-tau, p-tau, and Amyloid {beta}, and ratios). We also predicted MCI to AD progression with 69% accuracy, compared with the 70% accuracy using CSF biomarker model. The optimal classification accuracy in a single-site dataset and the reproduced results in multi-site dataset show the feasibility of the high-throughput imaging analysis of multimodal MRI and data-driven machine learning for predictive modeling in AD.

neuroscience

Pathway enrichment analysis of -omics data

Pathway enrichment analysis helps gain mechanistic insight into large gene lists typically resulting from genome scale (-omics) experiments. It identifies biological pathways that are enriched in the gene list more than expected by chance. We explain pathway enrichment analysis and present a practical step-by-step guide to help interpret gene lists resulting from RNA-seq and genome sequencing experiments. The protocol comprises three major steps: define a gene list from genome scale data, determine statistically enriched pathways, and visualize and interpret the results. We focus on differentially expressed genes and mutated cancer genes, however the described principles can be applied to diverse -omics data. The protocol is designed for biologists with no prior bioinformatics training and uses freely available software including g:Profiler, GSEA, Cytoscape and Enrichment Map.

bioinformatics

Drosophila Fezf coordinates laminar-specific connectivity through cell-intrinsic and cell-extrinsic mechanisms

Laminar arrangement of neural connections is a fundamental feature of neural circuit organization. Identifying mechanisms that coordinate neural connections within correct layers is thus vital for understanding how neural circuits are assembled. In the medulla of the Drosophila visual system neurons form connections within ten parallel layers. The M3 layer receives input from two neuron types that sequentially innervate M3 during development. Here we show that M3-specific innervation by both neurons is coordinated by Drosophila Fezf (dFezf), a conserved transcription factor that is selectively expressed by the earlier targeting input neuron. In this cell, dFezf instructs layer specificity and activates the expression of a secreted molecule (Netrin) that regulates the layer specificity of the other input neuron. We propose that employment of transcriptional modules that cell-intrinsically target neurons to specific layers, and cell-extrinsically recruit other neurons is a general mechanism for building layered networks of neural connections.

neuroscience

Endodermal differentiation is reconstructed by coordination of two parallel signaling systems derived from the stele in roots

The plant roots represent the exquisitely controlled cell fate map in which different cell types undergo a complete status transition from stem cell division and initial fate specification, to the terminal differentiation. The endodermis is initially specified in meristem but further differentiates to form Casparian strips (CSs), the apoplastic barrier in the mature zone for the selective transport between stele and outer tissues, and thus is regarded as plant inner skin. In the Arabidopsis thaliana root the transcription factors SHORTROOT (SHR) regulate asymmetric cell division in cortical initials to separate endodermal and cortex cell layer. In this paper, we utilized synthetic approach to examine the reconstruction of fully functional Casparian strips in plant roots. Our results revealed that SHR serves as a master regulator of a hierarchical signaling cascade that, combined with stele-derived small peptides, is sufficient to rebuild the functional CS in non-endodermal cells. This is a demonstration of the deployment of two parallel signaling systems, in which both apoplastic and symplastic communication were employed, for coordinately specifying the endodermal cell fate.

plant biology