Search bioRxivSearch

Biology subjects

Xu, H.

Publications and source records attributed to Xu, H..

28 records · Page 2Linked to original sources

Passive Audio Vocal Capture and Measurement in the Evaluation of Selective Mutism

Selective Mutism (SM) is an anxiety disorder often diagnosed in early childhood and characterized by persistent failure to speak in certain social situations but not others. Diagnosing SM and monitoring treatment response can be quite complex, due in part to changing definitions of and scarcity of research about the disorder. Subjective self-reports and parent/teacher interviews can complicate SM diagnosis and therapy, given that similar speech problems of etiologically heterogeneous origin can be attributed to SM. The present perspective discusses the potential for passive audio capture to help overcome psychiatrys current lack of objective and quantifiable assessments in the context of SM. We present evidence from two pilot studies indicating the feasibility of using a digital wearable device to quantify child vocalization features affected by SM. We also highlight limitations in the design and implementation of this preliminary work that can help guide future efforts.

neuroscience

A robust and tunable mitotic oscillator in artificial cells

Single-cell analysis is pivotal to deciphering complex phenomena like cellular heterogeneity, bistable switch, and oscillations, where a population ensemble cannot represent the individual behaviors. Bulk cell-free systems, despite having unique advantages of manipulation and characterization of biochemical networks, lack the essential single-cell information to understand a class of out-of-steady-state dynamics including cell cycles. Here we develop a novel artificial single-cell system by encapsulating Xenopus egg extracts in water-in-oil microemulsions to study mitotic dynamics. These \"cells\", adjustable in sizes and periods, sustain oscillations for over 30 cycles, and function in forms from the simplest cytoplasmic-only to the more complicated ones involving nuclei dynamics, mimicking real mitotic cells. Such innate flexibility and robustness make it key to studying clock properties of tunability and stochasticity. Our result also highlights energy supply as an important regulator of cell cycles. We demonstrate a simple, powerful, and likely generalizable strategy of integrating strengths of single-cell approaches into conventional in vitro systems to study complex clock functions.

systems biology

Assessment of the impact of shared data on the scientific literature

Data sharing is increasingly recommended as a means of accelerating science by facilitating collaboration, transparency, and reproducibility. While few oppose data sharing philosophically, a range of barriers deter most researchers from implementing it in practice (e.g., workforce and infrastructural demands, sociocultural and privacy concerns, lack of standardization). To justify the significant effort required for sharing data (e.g., organization, curation, distribution), funding agencies, institutions, and investigators need clear evidence of benefit. Here, using the International Neuroimaging Data-sharing Initiative, we present a brain imaging case study that provides direct evidence of the impact of open sharing on data use and resulting publications over a seven-year period (2010-2017). We dispel the myth that scientific findings using shared data cannot be published in high-impact journals and demonstrate rapid growth in the publication of such journal articles, scholarly theses, and conference proceedings. In contrast to commonly used pay to play models, we demonstrate that openly shared data can increase the scale (i.e., sample size) of scientific studies conducted by data contributors, and can recruit scientists from a broader range of disciplines. These findings suggest the transformative power of data sharing for accelerating science and underscore the need for the scientific ecosystem to embrace the challenge of implementing data sharing universally.

bioinformatics

Advanced whole genome sequencing and analysis of fetal genomes from amniotic fluid

Amniocentesis is typically performed to identify large chromosomal abnormalities within the fetus. Here we demonstrate that it is feasible to generate an accurate whole genome sequence (WGS) of a fetus from an amniotic sample. DNA from cells and the amniotic fluid were isolated and sequenced from 31 amniocenteses. Concordance of variant calls between the two DNA sources and with parental libraries was high. Two fetal genomes were found to harbor potentially detrimental variants in CHD8 and LRP1, variations in these genes have been associated with Autism Spectrum Disorder (ASD) and Keratosis pilaris atrophicans, respectively. We also discovered drug sensitivities and carrier information of fetuses for a variety of diseases. In this study, we demonstrate for the first time the sequencing of the whole genome of fetuses from amniotic fluid and show that much more information than large chromosomal abnormalities can be gained from an amniocentesis.

genomics

Computational correction of copy-number effect improves specificity of CRISPR-Cas9 essentiality screens in cancer cells

The CRISPR-Cas9 system has revolutionized gene editing both on single genes and in multiplexed loss-of-function screens, enabling precise genome-scale identification of genes essential to proliferation and survival of cancer cells. However, previous studies reported that an anti-proliferative effect of Cas9-mediated DNA cleavage confounds such measurement of genetic dependency, particularly in the setting of copy number gain1-4. We performed genome-scale CRISPR-Cas9 essentiality screens on 342 cancer cell lines and found that this effect is common to all lines, leading to false positive results when targeting genes in copy number amplified regions. We developed CERES, a computational method to estimate gene dependency levels from CRISPR-Cas9 essentiality screens while accounting for the copy-number-specific effect, as well as variable sgRNA activity. We applied CERES to sets of screens performed with different sgRNA libraries and found that it reduces false positive results and provides meaningful estimates of sgRNA activity. As a result, the application of CERES improves confidence in the interpretation of genetic dependency data from CRISPR-Cas9 essentiality screens of cancer cell lines.

genomics

The pomegranate (Punica granatum L.) genome provides insights into fruit quality and ovule developmental biology

Pomegranate (Punica granatum L.) with an uncertain taxonomic status has an ancient cultivation history, and has become an emerging fruit due to its attractive features such as the bright red appearance and the high abundance of medicinally valuable ellagitannin-based compounds in its peel and aril. However, the absence of genomic resources has restricted further elucidating genetics and evolution of these interesting traits. Here we report a 274-Mb high-quality draft pomegranate genome sequence, which covers approximately 81.5% of the estimated 336 Mb genome, consists of 2,177 scaffolds with an N50 size of 1.7 Mb, and contains 30,903 genes. Phylogenomic analysis supported that pomegranate belongs to the Lythraceae family rather than the monogeneric Punicaceae family, and comparative analyses showed that pomegranate and Eucalyptus grandis shares the paleotetraploidy event. Integrated genomic and transcriptomic analyses provided insights into the molecular mechanisms underlying the biosynthesis of ellagitannin-based compounds, the color formation in both peels and arils during pomegranate fruit development, and the unique ovule development processes that are characteristic of pomegranate. This genome sequence represents the first reference in Lythraceae, providing an important resource to expand our understanding of some unique biological processes and to facilitate both comparative biology studies and crop breeding.

genomics

Arabidopsis thaliana Trihelix Transcription factor AST1 mediates abiotic stress tolerance by binding to a novel AGAG-box and some GT motifs

Trihelix transcription factors are characterized by containing a conserved trihelix (helix-loop-helix-loop-helix) domain that bind to GT elements required for light response, play roles in light stress, and also in abiotic stress responses. However, only few of them have been functionally characterised. In the present study, we characterized the function of AST1 (Arabidopsis SIP1 clade Trihelix1) in response to abiotic stress. AST1 shows transcriptional activation activity, and its expression is induced by osmotic and salt stress. The genes regulated by AST1 were identified using qRT-PCR and transcriptome assays. A conserved sequence highly present in the promoters of genes regulated by AST1 was identified, which is bound by AST1, and termed AGAG-box with the sequence [A/G][G/A][A/T]GAGAG. Additionally, AST1 also binds to some GT motifs including GGTAATT, TACAGT, GGTAAAT and GGTAAA, but failed in binding to GTTAC and GGTTAA. Chromatin immunoprecipitation combined with qRT-PCR analysis suggested that AST1 binds to AGAG-box and/or some GT motifs to regulate the expression of stress tolerance genes, resulting in reduced reactive oxygen species, Na+ accumulation, stomatal apertures, lipid peroxidation, cell death and water loss rate, and increased proline content and reactive oxygen species scavenging capability. These physiological changes mediated by AST1 finally improve abiotic stress tolerance.

physiology

Integrative analysis and refined design of CRISPR knockout screens

Genome-wide CRISPR-Cas9 screen has been widely used to interrogate gene functions. However, the analysis remains challenging and rules to design better libraries beg further refinement. Here we present MAGeCK-NEST, which integrates protein-protein interaction (PPI), improves the inference accuracy when fewer guide-RNAs (sgRNAs) are available, and assesses screen qualities using information on PPI. MAGeCK-NEST also adopts a maximum-likelihood approach to remove sgRNA outliers, which are characterized with higher G-nucleotide counts, especially in regions distal from the PAM motif. Using MAGeCK-NEST, we found that choosing non-targeting sgRNAs as negative controls lead to strong bias, which can be mitigated by sgRNAs targeting the \"safe harbor\" regions. Custom-designed screens confirmed our findings, and further revealed that 19nt sgRNAs consistently gave the best signal-to-noise separation. Collectively, our method enabled robust calling of CRISPR screen hits and motivated the design of an improved genome-wide CRISPR screen library.

bioinformatics

DATS: the data tag suite to enable discoverability of datasets

Todays science increasingly requires effective ways to find and access existing datasets that are distributed across a range of repositories. For researchers in the life sciences, discoverability of datasets may soon become as essential as identifying the latest publications via PubMed. Through an international collaborative effort funded by the National Institutes of Health (NIH)s Big Data to Knowledge (BD2K) initiative, we have designed and implemented the DAta Tag Suite (DATS) model to support the DataMed data discovery index. DataMeds goal is to be for data what PubMed has been for the scientific literature. Akin to the Journal Article Tag Suite (JATS) used in PubMed, the DATS model enables submission of metadata on datasets to DataMed. DATS has a core set of elements, which are generic and applicable to any type of datasets, and an extended set that can accommodate more specialized data types. DATS is a platform-independent model also available as a Schema.org annotated serialization to be used beyond DataMed, for example, in projects like DataCite.

bioinformatics

DataMed: Finding useful data across multiple biomedical data repositories

The value of broadening searches for data across multiple repositories has been identified by the biomedical research community. As part of the NIH Big Data to Knowledge initiative, we work with an international community of researchers, service providers and knowledge experts to develop and test a data index and search engine, which are based on metadata extracted from various datasets in a range of repositories. DataMed is designed to be, for data, what PubMed has been for the scientific literature. DataMed supports Findability and Accessibility of datasets. These characteristics - along with Interoperability and Reusability - compose the four FAIR principles to facilitate knowledge discovery in todays big data-intensive science landscape.

bioinformatics