Search bioRxivSearch

Biology subjects

Wang, W.

Publications and source records attributed to Wang, W..

At least 73 records · Page 4Linked to original sources

Gene markers for exon capture and phylogenomics in ray-finned fishes

Gene capture coupled with the next generation sequencing has become one of the favorable methods in subsampling genomes for phylogenomic studies. Many target gene markers have been developed in plants, sharks, frogs, reptiles and others, but few have been reported in the ray-finned fishes. Here, we identified a suite of \"single-copy\" protein coding sequence (CDS) markers through comparing eight fish genomes, and tested them empirically in 83 species (33 families and 11 orders) of ray-finned fishes. Sorting through the markers according to their completeness and phylogenetic decisiveness in taxa tested resulted in a selection of 4,434 markers, which were proven to be useful in reconstructing phylogenies of the ray-finned fishes at different taxonomic level. We also proposed a strategy of refining baits (probes) design a posteriori based on empirical data. The markers that we have developed may fill a gap in the tool kit of phylogenomic study in vertebrates.

evolutionary biology

Expanding primary cells from mucoepidermoid and other salivary gland neoplasms for genetic and chemosensitivity testing

Restricted availability of cell and animal models is a rate-limiting step for investigation of salivary gland neoplasm pathophysiology and therapeutic response. Conditionally reprogrammed cell (CRC) technology enables establishment of primary epithelial cell cultures from patient material. This study tested a translational workflow for acquisition, expansion and testing of CRC-derived primary cultures of salivary gland neoplasms from patients presenting to an academic surgical practice. Results showed cultured cells were sufficient for epithelial cell-specific transcriptome characterization to detect candidate therapeutic pathways and fusion genes in addition to screening for cancer-risk-associated single nucleotide polymorphisms (SNPs) and driver gene mutations through exome sequencing. Focused study of primary cultures of a low-grade mucoepidermoid carcinoma demonstrated Amphiregulin-Mechanistic Target of Rapamycin-AKT/Protein kinase B (AKT) pathway activation, identified through bioinformatics and subsequently confirmed as present in primary tissue and preserved through different secondary 2D and 3D culture media and xenografts. Candidate therapeutic testing showed that the allosteric AKT inhibitor MK2206 reproducibly inhibited cell survival across different culture formats. In contrast, the cells appeared resistant to the adenosine triphosphate competitive AKT inhibitor GSK690693. Procedures employed here illustrate an approach for reproducibly obtaining material for pathophysiological studies of salivary gland neoplasms, and other less common epithelial cancer types, that can be executed without compromising pathological examination of patient specimens. The approach permits combined genetic and cell-based physiological and therapeutic investigations in addition to more traditional pathologic studies and can be used to build sustainable bio-banks for future inquiries.

cancer biology

Uncovering Medical Insights from Vast Amounts of Biomedical Data in Clinical Case Reports

Clinical case reports (CCRs) have a time-honored tradition in serving as an important means of sharing clinical experiences on patients presenting with atypical disease phenotypes or receiving new therapies. However, the huge amount of accumulated case reports are isolated, unstructured, and heterogeneous clinical data, posing a great challenge to clinicians and researchers in mining relevant information through existing indexing tools. In this investigation, in order to render CCRs more findable, accessible, interoperable, and reusable (FAIR) by the biomedical community, we created a resource platform, including the construction of a test dataset consisting of 1000 CCRs spanning 14 disease phenotypes, a standardized metadata template and metrics, and a set of computational tools to automatically retrieve relevant medical information and to analyze all published PubMed clinical case reports with respect to trends in publication journals, citations impact, MeSH Terms, drug use, distributions of patient demographics, and relationships with other case reports and databases. Our standardized metadata template and CCR test dataset may be valuable resources to advance medical science and improve patient care for researchers who are using machine learning approaches with a high-quality dataset to train and validate their algorithms. In the future, our analytical tools may be applied towards other large clinical data sources as well.

bioinformatics

Integrative DNA copy number detection and genotyping from sequencing and array-based platforms

MotivationCopy number variations (CNVs) are gains and losses of DNA segments and have been associated with disease. Many large-scale genetic association studies are performing CNV analysis using whole exome sequencing (WES) and whole genome sequencing (WGS). In many of these studies, previous SNP-array data are available. An integrated cross-platform analysis is expected to improve resolution and accuracy, yet there is no tool for effectively combining data from sequencing and array platforms. The detection of CNVs using sequencing data alone can also be further improved by the utilization of allele-specific reads.\n\nResultsWe propose a statistical framework, integrated Copy Number Variation detection algorithm (iCNV), which can be applied to multiple study designs: WES only, WGS only, SNP array only, or any combination of SNP and sequencing data. iCNV applies platform specific normalization, utilizes allele specific reads from sequencing and integrates matched NGS and SNP-array data by a Hidden Markov Model (HMM). We compare integrated two-platform CNV detection using iCNV to naive intersection or union of platforms and show that iCNV increases sensitivity and robustness. We also assess the accuracy of iCNV on WGS data only, and show that the utilization of allele-specific reads improve CNV detection accuracy compared to existing methods.\n\nAvailabilityhttps://github.com/zhouzilu/iCNV\n\nContactnzh@wharton.upenn.edu, zhouzilu@mail.med.upenn.edu\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

bioinformatics

Time-gated detection of protein-protein interactions with transcriptional readout

Transcriptional assays such as yeast two hybrid, split ubiquitin, and Tango that convert transient protein-protein interactions (PPIs) in cells into stable expression of transgenes are powerful tools for PPI discovery, high-throughput screens, and analysis of large cell populations. However, these assays frequently suffer from high background and they lose all information about PPI dynamics. To address these limitations, we developed a light-gated transcriptional assay for PPI detection called PPI-FLARE (PPI-Fast Light- and Activity-Regulated Expression). PPI-FLARE requires both a PPI to deliver TEV protease proximal to its cleavage peptide, and externally-applied blue light to uncage the cleavage peptide, in order to release a membrane-tethered transcription factor (TF) for translocation to the nucleus. We used PPI-FLARE to detect the ligand-induced association of 12 different PPIs in living mammalian cells, with a temporal resolution of 5 minutes and a {+/-}ligand signal ratio up to 37. By systematically shifting the light irradiation window, we could reconstruct PPI time-courses, distinguishing between GPCRs that engage in transient versus sustained interactions with the cytosolic effector arrestin. When combined with FACS, PPI-FLARE enabled >100-fold enrichment of cells experiencing a specific GPCR-arrestin PPI during a short 10-minute light window over cells missing that PPI during the same time window. Due to its high specificity, sensitivity, and generality, PPI-FLARE should be a broadly useful tool for PPI analysis and discovery.

bioengineering

Systematic mapping of chromatin state landscapes during mouse development

Embryogenesis requires epigenetic information that allows each cell to respond appropriately to developmental cues. Histone modifications are core components of a cells epigenome, giving rise to chromatin states that modulate genome function. Here, we systematically profile histone modifications in a diverse panel of mouse tissues at 8 developmental stages from 10.5 days post conception until birth, performing a total of 1,128 ChIP-seq assays across 72 distinct tissue-stages. We combine these histone modification profiles into a unified set of chromatin state annotations, and track their activity across developmental time and space. Through integrative analysis we identify dynamic enhancers, reveal key transcriptional regulators, and characterize the role of chromatin-based repression in developmental gene regulation. We also leverage these data to link enhancers to putative target genes, revealing connections between coding and non-coding sequence variation in disease etiology. Our study provides a compendium of resources for biomedical researchers, and achieves the most comprehensive view of embryonic chromatin states to date.

genomics

Systems-level identification of transcription factors critical for mouse embryonic development

Dynamic changes in the transcriptional regulatory circuit can influence the specification of distinct cell types. Numerous transcription factors (TFs) have been shown to function through dynamic rewiring during embryonic development but a comprehensive survey on the global regulatory network is still lacking. Here, we performed an integrated analysis of epigenomic and transcriptomic data to reveal key regulators from 2 cells to postnatal day 0 in mouse embryogenesis. We predicted 3D chromatin interactions including enhancer-promoter interactions in 12 tissues across 8 developmental stages, which facilitates linking TFs to their target genes for constructing genetic networks. To identify driver TFs particularly those not necessarily differentially expressed ones, we developed a new algorithm, dubbed as Taiji, to assess the global importance of TFs in development. Through comparative analysis across tissues and developmental stages, we systematically uncovered TFs that are critical for lineage-specific and stage-dependent tissue specification. Most interestingly, we have identified TF combinations that function in spatiotemporal order to form transcriptional waves regulating developmental progress and differentiation. Not only does our analysis provide the first comprehensive map of transcriptional regulatory circuits during mouse embryonic development, the identified novel regulators and the predicted 3D chromatin interactions also provide a valuable resource to guide further mechanistic studies.

bioinformatics

The evolutionary history of 2,658 cancers

Cancer develops through a process of somatic evolution. Here, we use whole-genome sequencing of 2,778 tumour samples from 2,658 donors to reconstruct the life history, evolution of mutational processes, and driver mutation sequences of 39 cancer types. The early phases of oncogenesis are driven by point mutations in a small set of driver genes, often including biallelic inactivation of tumour suppressors. Early oncogenesis is also characterised by specific copy number gains, such as trisomy 7 in glioblastoma or isochromosome 17q in medulloblastoma. By contrast, increased genomic instability, a nearly four-fold diversification of driver genes, and an acceleration of point mutation processes are features of later stages. Copy-number alterations often occur in mitotic crises leading to simultaneous gains of multiple chromosomal segments. Timing analysis suggests that driver mutations often precede diagnosis by many years, and in some cases decades, providing a window of opportunity for early cancer detection.

cancer biology

Draft genome of the Reindeer (Rangifer tarandus)

AbstractO_ST_ABSBackgroundC_ST_ABSReindeer (Rangifer tarandus) is the only fully domesticated species in the Cervidae family, and is the only cervid with a circumpolar distribution. Unlike all other cervids, female reindeer regularly grow cranial appendages (antlers, the defining characteristics of cervids), as well as males. Moreover, reindeer milk contains more protein and less lactose than bovids milk. A high quality reference genome of this specie will assist efforts to elucidate these and other important features in the reindeer.\n\nFindingsWe obtained 723.2 Gb (Gigabase) of raw reads by an Illumina Hiseq 4000 platform, and a 2.64 Gb final assembly, representing 95.7% of the estimated genome (2.76 Gb according to k-mer analysis), including 92.6% of expected genes according to BUSCO analysis. The contig N50 and scaffold N50 sizes were 89.7 kilo base (kb) and 0.94 mega base (Mb), respectively. We annotated 21,555 protein-coding genes and 1.07 Gb of repetitive sequences by de novo and homology-based prediction. Homology-based searches detected 159 rRNA, 547 miRNA, 1,339 snRNA and 863 tRNA sequences in the genome of R. tarandus. The divergence time between R. tarandus, and ancestors of Bos taurus and Capra hircus, is estimated to be 29.55 million years ago (Mya).\n\nConclusionsOur results provide the first high-quality reference genome for the reindeer, and a valuable resource for studying evolution, domestication and other unusual characteristics of the reindeer.

genomics

Dual functions of Discoidin domain receptor coordinate cell-matrix adhesion and collective polarity in migratory cardiopharyngeal progenitors

Integrated analyses of regulated effector genes, cellular processes, and extrinsic signals are required to understand how transcriptional networks coordinate fate specification and cell behavior during embryogenesis. Migratory pairs of cardiac progenitors in the tunicate Ciona provide the simplest model of collective migration in chordate embryos. Ciona cardiopharyngeal progenitors (aka trunk ventral cells, TVCs) polarize as leader and trailer cells, and migrate between the ventral epidermis and trunk endoderm, which influences collective polarity. Using functional perturbations and quantitative analyses, we show that the TVC-specific and collagen-binding Discoidin-domain receptor (Ddr) cooperates with Integrin-{beta}1 to promote cell-matrix adhesion to the epidermis. We found that endoderm cells secrete a collagen, Col9-a1, that is deposited in the basal epidermal matrix and activates Ddr at the ventral membrane of migrating TVCs. A functional antagonism between Ddr/Int{beta}1-mediated cell-matrix adhesion and Vegfr signaling appears to modulate the position of cardiopharyngeal progenitors between the endoderm and epidermis. Finally, we show that Ddr activity promotes leader-trailer-polarized BMP-Smad signaling independently of its role in cell-matrix adhesion. We propose that dual functions of Ddr act downstream of cardiopharyngeal-specific transcriptional inputs to coordinate subcellular processes underlying collective polarity and directed migration.

developmental biology

A single cell transcriptional roadmap for cardiopharyngeal fate diversification

In vertebrates, multipotent progenitors located in the pharyngeal mesoderm form cardiomyocytes and branchiomeric head muscles, but the dynamic gene expression programs and mechanisms underlying cardiopharyngeal multipotency and heart vs. head muscle fate choices remain elusive. Here, we used single cell genomics in the simple chordate model Ciona, to reconstruct developmental trajectories forming first and second heart lineages, and pharyngeal muscle precursors, and characterize the molecular underpinnings of cardiopharyngeal fate choices. We show that FGF-MAPK signaling maintains multipotency and promotes the pharyngeal muscle fate, whereas signal termination permits the deployment of a pan-cardiac program, shared by the first and second lineages, to define heart identity. In the second heart lineage, a Tbx1/10-Dach pathway actively suppresses the first heart lineage program, conditioning later cell diversity in the beating heart. Finally, cross-species comparisons between Ciona and the mouse evoke the deep evolutionary origins of cardiopharyngeal networks in chordates.

developmental biology

Transcriptome Deconvolution of Heterogeneous Tumor Samples with Immune Infiltration

Transcriptomic deconvolution in cancer and other heterogeneous tissues remains challenging. Available methods lack the ability to estimate both component-specific proportions and expression profiles for individual samples. We present DeMixT, a new tool to deconvolve high dimensional data from mixtures of more than two components. DeMixT implements an iterated conditional mode algorithm and a novel gene-set-based component merging approach to improve accuracy. In a series of experimental validation studies and application to TCGA data, DeMixT showed high accuracy. Improved deconvolution is an important step towards linking tumor transcriptomic data with clinical outcomes. An R package, scripts and data are available: https://github.com/wwylab/DeMixT/.

bioinformatics

Protein Profiling In Cancer Cell Lines And Tumor Tissue Using Reverse Phase Protein Arrays

Reverse phase protein array (RPPA) technology is an antibody-based high-throughput assay for protein profiling of biological specimens that allows for many measurements with very small amounts of cell lysate. Here, we report the sensitivity, reproducibility, and accuracy of a particular RPPA platform called Zeptosens. We customized the RPPA protocol for our in-house setup, and measured more than 80 total protein and phospho-protein levels in various cancer samples, including cell lines, organoids, tumor chunks, core needle biopsies, and laser-capture microdissected tissue samples. We discuss pros and cons of the RPPA platform, and describe results from profiling 15 cancer cell line cells using RPPA.

systems biology

Cells Interpret Temporal Information From TGF-β Through A Nested Relay Mechanism

The detection and transmission of the temporal quality of intracellular and extracellular signals is an essential cellular mechanism. It remains largely unexplored how cells interpret the duration information of a stimulus. In this paper, through an integrated quantitative and computational approach we demonstrate that crosstalk among multiple TGF-{beta} activated pathways forms a relay from SMAD to GLI1 that initializes and maintains SNAILl expression, respectively. This transaction is smoothed and accelerated by another temporal switch from elevated cytosolic GSK3 enzymatic activity to reduced nuclear GSK3 enzymatic activity. The intertwined network places SNAIL1 as a key integrator of information from TGF-{beta} signaling subsequently distributed through upstream divergent pathways; essentially cells generate a transient or sustained expression of SNAIL1 depending on TGF-{beta} duration. Other signaling pathways may use similar network structure to encode duration information.

systems biology

Establishment In Culture Of Expanded Potential Stem Cells

Mouse embryonic stem cells are derived from in vitro explantation of blastocyst epiblasts1,2 and contribute to both the somatic lineage and germline when returned to the blastocyst3 but are normally excluded from the trophoblast lineage and primitive endoderm4-6. Here, we report that cultures of expanded potential stem cells (EPSCs) can be established from individual blastomeres, by direct conversion of mouse embryonic stem cells (ESCs) and by genetically reprogramming somatic cells. Remarkably, a single EPSC contributes to the embryo proper and placenta trophoblasts in chimeras. Critically, culturing EPSCs in a trophoblast stem cell (TSC) culture condition permits direct establishment of TSC lines without genetic modification. Molecular analyses including single cell RNA-seq reveal that EPSCs share cardinal pluripotency features with ESCs but have an enriched blastomere transcriptomic signature and a dynamic DNA methylome. These proof-of-concept results open up the possibility of establishing cultures of similar stem cells in other mammalian species.

developmental biology

High fidelity hypothermic preservation of primary tissues in organ transplant preservative for single cell transcriptome analysis

BackgroundHigh-fidelity preservation strategies for primary tissues are in great demand in the single cell RNAseq community. A reliable method will greatly expand the scope of feasible collaborations and maximize the utilization of technical expertise. When choosing a method, standardizability is as important a factor to consider as fidelity due to the susceptibility of single-cell RNAseq analysis to technical noises. Existing approaches such as cryopreservation and chemical fixation are less than ideal for failing to satisfy either or both of these standards.\n\nResultsHere we propose a new strategy that leverages preservation schemes developed for organ transplantation. We evaluated the strategy by storing intact mouse kidneys in organ transplant preservative solution at hypothermic temperature for up to 4 days (6 hrs, 1, 2, 3, and 4 days), and comparing the quality of preserved and fresh samples using FACS and single cell RNAseq. We demonstrate that the strategy effectively maintained cell viability, transcriptome integrity, cell population heterogeneity, and transcriptome landscape stability for samples after up to 3 days of preservation. The strategy also facilitated the definition of the diverse spectrum of kidney resident immune cells, to our knowledge the first time at single cell resolution.\n\nConclusionsHypothermic storage of intact primary tissues in organ transplant preservative maintains the quality and stability of the transcriptome of cells for single cell RNAseq analysis. The strategy is readily generalizable to primary specimens from other tissue types for single cell RNAseq analysis.

genomics

Using eDNA to Detect the Distribution and Density of Invasive Crayfish in the Honghe-Hani Rice Terrace World Heritage Site

The Honghe-Hani landscape in China is a UNESCO World Natural Heritage site due to the beauty of its thousands of rice terraces, but these structures are in danger from the invasive crayfish Procambarus clarkii. Crayfish dig nest holes, which collapse terrace walls and destroy rice production. Under the current control strategy, farmers self-report crayfish and are issued pesticide, but this strategy is not expected to eradicate the crayfish nor to prevent their spread since farmers are not able to detect small numbers of crayfish. Thus, we tested whether environmental DNA (eDNA) from paddy-water samples could provide a sensitive detection method. In an aquarium experiment, Real-time Quantitative polymerase chain reaction (qPCR) successfully detected crayfish, even at a simulated density of one crayfish per average-sized paddy (with one false negative). In a field test, we tested eDNA and bottle traps against direct counts of crayfish. eDNA successfully detected crayfish in all 25 paddies where crayfish were observed and in none of the 7 paddies where crayfish were absent. Bottle-trapping was successful in only 68% of the crayfish-present paddies. eDNA concentrations also correlated positively with crayfish counts. In sum, these results suggest that single samples of eDNA are able to detect small crayfish populations, but not perfectly. Thus, we conclude that a program of repeated eDNA sampling is now feasible and likely reliable for measuring crayfish geographic range and for detecting new invasion fronts in the Honghe Hani landscape, which would inform regional control efforts and help to prevent the further spread of this invasive crayfish.

ecology

Novel insights into the molecular heterogeneity of hepatocellularcarcinoma

Hepatocellular carcinoma (HCC) is influenced by numerous factors, which results in diverse genetic, epigenetic and transcriptional scenarios, thus posing obvious challenges for disease management. We scrutinized the molecular heterogeneity of HCC with a multi-omics approach in two small cohorts of resected and explanted livers. Whole-genome transcriptomics was conducted, including polyadenylated transcripts and micro (mi)-RNAs. Copy number variants (CNV) were inferred from whole genome low-pass sequencing data. Fifty-six cancer-related genes were screened using an oncology panel assay. HCC was associated with a dramatic transcriptional deregulation of hundreds of protein-coding genes suggesting downregulation of drugs catabolism, induction of inflammatory responses, and increased cell proliferation in resected livers. Moreover, several long non-coding RNAs and miRNAs not reported previously in the context of HCC were found deregulated. In explanted livers, downregulation of genes involved in energy-producing processes and upregulation of genes aiding in glycolysis were detected. Numerous CNV events were observed, with conspicuous hotspots on chromosomes 1 and 17. Amplifications were more common than deletions, and spanned regions containing genes potentially involved in tumorigenesis. CSF1R, FGFR3, FLT3, NPM1, PDGFRA, PTEN, SMO and TP53 were mutated in all tumors, while other 26 cancer-related genes were mutated with variable penetrance. Our results highlight a remarkable molecular heterogeneity between HCC tumors and reinforce the notion that precision medicine approaches are urgently needed for cancer treatment. We expect that our results will serve as a valuable dataset that will generate hypotheses for us or other researchers to evaluate to ultimately improve our understanding of HCC biology.

cancer biology