Search bioRxivSearch

Biology subjects

Bai, X.

Publications and source records attributed to Bai, X..

8 recordsLinked to original sources

GLnexus: joint variant calling for large cohort sequencing

As ever-larger cohorts of human genomes are collected in pursuit of genotype/phenotype associations, sequencing informatics must scale up to yield complete and accurate genotypes from vast raw datasets. Joint variant calling, a data processing step entailing simultaneous analysis of all participants sequenced, exhibits this scaling challenge acutely. We present GLnexus (GL, Genotype Likelihood), a system for joint variant calling designed to scale up to the largest foreseeable human cohorts. GLnexus combines scalable joint calling algorithms with a persistent database that grows efficiently as additional participants are sequenced. We validate GLnexus using 50,000 exomes to show it produces comparable or better results than existing methods, at a fraction of the computational cost with better scaling. We provide a standalone open-source version of GLnexus and a DNAnexus cloud-native deployment supporting very large projects, which has been employed for cohorts of >240,000 exomes and >22,000 whole-genomes.

bioinformatics

Programmed Variations of Cytokinesis Contribute to Morphogenesis in the C. elegans embryo

While cytokinesis has been intensely studied, how it is executed during development is not well understood, despite a long-standing appreciation that various aspects of cytokinesis vary across cell and tissue types. To address this, we investigated cytokinesis during the invariant C. elegans embryo lineage and found several reproducibly altered parameters at different stages. During early divisions, furrow ingression asymmetry and midbody inheritance is consistent, suggesting specific regulation of these events. During morphogenesis, we find several unexpected alterations including migration of midbodies to the apical surface during epithelial polarization in different tissues. Aurora B kinase, which is essential for several aspects of cytokinesis, remains localized to the apical membrane after internalization of other midbody components. Inactivation of Aurora B causes cytokinesis failure, which disrupts polarization and tissue formation. Therefore, cytokinesis shows surprising diversity during development and is required during epithelial polarization to establish cellular architecture during morphogenesis.

cell biology

DensityPath: a level-set algorithm to visualize and reconstruct cell developmental trajectories for large-scale single-cell RNAseq data

Cell fates are determined by transition-states which occur during complex biological pro-cesses such as proliferation and differentiation. The advance in single-cell RNA sequencing (scRNAseq) provides the snapshots of single cell transcriptomes, thus offering an essential opportunity to study such complex biological processes. Here, we introduce a novel algorithm, DensityPath, which visualizes and reconstructs the underlying cell developmental trajectories for large-scale scRNAseq data. DensityPath has three merits. Firstly, by adopting the nonlinear dimension reduction algorithm elastic embedding, DensityPath reveals the intrinsic structures of the data. Secondly, by applying the powerful level set clustering method, DensityPath extracts the separate high density clusters of representative cell states (RCSs) from the single cell multimodal density landscape of gene expression space, enabling it to handle the heterogeneous scRNAseq data elegantly and accurately. Thirdly, DensityPath constructs cell state-transition path by finding the geodesic minimum spanning tree of the RCSs on the surface of the density landscape, making it more computationally efficient and accurate for large-scale dataset. The cell state-transition path constructed by DensityPath has the physical interpretation as the minimum-transition-energy (least-cost) path. We demonstrate that DensityPath is capable of identifying complex cell development trajectories with bifurcating and trifurcating branches on the human preimplantation embryos. We demonstrate that DensityPath is robust and has high accuracy of pseudotime calculation and branch assignment on the real scRNAseq as well as simulated datasets.

systems biology

Cryo-EM structure of the adenosine A2A receptor coupled to an engineered heterotrimeric G protein

The adenosine A2A receptor (A2AR) is a prototypical G protein-coupled receptor (GPCR) that couples to the heterotrimeric G protein GS. Here we determine the structure by electron cryo-microscopy (cryo-EM) of A2AR at pH 7.5 bound to the small molecule agonist NECA and coupled to an engineered heterotrimeric G protein, which contains mini-GS, the {beta}{gamma} subunits and nanobody Nb35. Most regions of the complex have a resolution of ~3.8 [A] or better. Comparison with the 3.4 [A] resolution crystal structure shows that the receptor and mini-GS are virtually identical and that the density of the side chains and ligand are of comparable quality. However, the cryo-EM density map also indicates regions that are flexible in comparison to the crystal structures, which unexpectedly includes regions in the ligand binding pocket. In addition, an interaction between intracellular loop 1 of the receptor and the {beta} subunit of the G protein was observed.

biophysics

KaryoScan: abnormal karyotype detection from whole-exome sequence

MotivationDetection of abnormal karyotypes from whole-exome sequencing has significant clinical potential, enabling a primary screen for chromosomal anomalies among samples undergoing short-read sequencing for nucleotide resolution genomic characterization.\n\nResultsWe present KaryoScan, a high-throughput method for detecting chromosomal anomalies within large cohort exome sequencing studies. We detect and validate autosomal and sex chromosomal aneuploidies in a large exome sequencing cohort, and demonstrate detection of smaller and complex events (partial chromosome, mosaic, copy neutral, and complex rearrangements), representing the range of anomalies that can be uncovered from the exome.\n\nAvailabilityhttps://github.com/rgcgithub/karyoscan

bioinformatics

Profiling and leveraging relatedness in a precision medicine cohort of 92,455 exomes

Large-scale human genetics studies are ascertaining increasing proportions of populations as they continue growing in both number and scale. As a result, the amount of cryptic relatedness within these study cohorts is growing rapidly and has significant implications on downstream analyses. We demonstrate this growth empirically among the first 92,455 exomes from the DiscovEHR cohort and, via a custom simulation framework we developed called SimProgeny, show that these measures are in-line with expectations given the underlying population and ascertainment approach. For example, we identified [~]66,000 close (first- and second-degree) relationships within DiscovEHR involving 55.6% of study participants. Our simulation results project that >70% of the cohort will be involved in these close relationships as DiscovEHR scales to 250,000 recruited individuals. We reconstructed 12,574 pedigrees using these relationships (including 2,192 nuclear families) and leveraged them for multiple applications. The pedigrees substantially improved the phasing accuracy of 20,947 rare, deleterious compound heterozygous mutations. Reconstructed nuclear families were critical for identifying 3,415 de novo mutations in [~]1,783 genes. Finally, we demonstrate the segregation of known and suspected disease-causing mutations through reconstructed pedigrees, including a tandem duplication in LDLR causing familial hypercholesterolemia. In summary, this work highlights the prevalence of cryptic relatedness expected among large healthcare population genomic studies and demonstrates several analyses that are uniquely enabled by large amounts of cryptic relatedness.

genomics

Genetic Identification of Novel Separase regulators in Caenorhabditis elegans

Separase is a highly conserved protease required for chromosome segregation. Although observations that separase also regulates membrane trafficking events have been made, it is still not clear how separase achieves this function. Here we present an extensive ENU mutagenesis suppressor screen aimed at identifying suppressors of sep-1(e2406), a temperature sensitive maternal effect embryonic lethal separase mutant. We screened nearly a million haploid genomes, and isolated sixty-eight suppressed lines. We identified fourteen independent intragenic sep-1(e2406) suppressed lines. These intragenic alleles map to seven SEP-1 residues within the N-terminus, compensating for the original mutation within the poorly conserved N-terminal domain. Interestingly, 47 of the suppressed lines have novel mutations throughout the entire coding region of the pph-5 phosphatase, indicating that this is an important regulator of separase. We also found that a mutation near the MEEVD motif of HSP-90, which binds and activates PPH-5, also rescues sep-1(e2406) mutants. Finally, we identified six potentially novel suppressor lines that fall into five complementation groups. These new alleles provide the opportunity to more exhaustively investigate the regulation and function of separase.

genetics

Profiling copy number variation and disease associations from 50,726 DiscovEHR Study exomes

Copy number variants (CNVs) are a substantial source of genomic variation and contribute to a wide range of human disorders. Gene-disrupting exonic CNVs have important clinical implications as they can underlie variability in disease presentation and susceptibility. The relationship between exonic CNVs and clinical traits has not been broadly explored at the population level, primarily due to technical challenges. We surveyed common and rare CNVs in the exome sequences of 50,726 adult DiscovEHR study participants with linked electronic health records (EHRs). We evaluated the diagnostic yield and clinical expressivity of known pathogenic CNVs, and performed tests of association with EHR-derived serum lipids, thereby evaluating the relationship between CNVs and complex traits and phenotypes in an unbiased, real-world clinical context. We identified CNVs from megabase to exon-level resolution, demonstrating reliable, high-throughput detection of clinically relevant exonic CNVs. In doing so, we created a catalog of high-confidence common and rare CNVs and refined population frequency estimates of known and novel gene-disrupting CNVs. Our survey among an unselected clinical population provides further evidence that neuropathy-associated duplications and deletions in 17p12 have similar population prevalence but are clinically under-diagnosed. Similarly, adults who harbor 22q11.2 deletions frequently had EHR documentation of neurodevelopmental/neuropsychiatric disorders and congenital anomalies, but not a formal genetic diagnosis (i.e., deletion). In an exome-wide association study of lipid levels, we identified a novel five-exon duplication within LDLR segregating in a large kindred with features of familial hypercholesterolemia. Exonic CNVs provide new opportunities to understand and diagnose human disease.

genomics