Search bioRxivSearch

Biology subjects

Yang, C.

Publications and source records attributed to Yang, C..

At least 19 recordsLinked to original sources

LPM: a latent probit model to characterize the relationship among complex traits using summary statistics from multiple GWASs and functional annotations

Much effort has been made toward understanding the genetic architecture of complex traits and diseases. Recent results from genome-wide association studies (GWASs) suggest the importance of regulatory genetic effects and pervasive pleiotropy among complex traits. In this study, we propose a unified statistical approach, aiming to characterize relationship among complex traits, and prioritize risk variants by leveraging regulatory information collected in functional annotations. Specifically, we consider a latent probit model (LPM) to integrate summary-level GWAS data and functional annotations. The developed computational framework not only makes LPM scalable to hundreds of annotations and phenotypes, but also ensures its statistically guaranteed accuracy. Through comprehensive simulation studies, we evaluated LPMs performance and compared it with related methods. Then we applied it to analyze 44 GWASs with nine genic category annotations and 127 cell-type specific functional annotations. The results demonstrate the benefits of LPM and gain insights of genetic architecture of complex traits. The LPM package is available at https://github.com/mingjingsi/LPM.

genetics

Structural and biochemical characterization on the cognate and heterologous interactions of the MazEF-mt9 TA system

The toxin-antitoxin (TA) modules widely exist in bacteria, and their activities are associated with the persister phenotype of the pathogen Mycobacterium tuberculosis (M. tb). M. tb causes Tuberculosis, a contagious and severe airborne disease. There are ten MazEF TA systems in M. tb, which play important roles in stress adaptation. How the antitoxins antagonize toxins in M. tb or how the ten TA systems crosstalk to each other are of interests, but the detailed molecular mechanisms are largely unclear. MazEF-mt9 is a unique member among the MazEF families due to its tRNase activity, which is usually carried out by the VapC family toxins. Here we present the cocrystal structure of the MazEF-mt9 complex at 2.7 [A]. By characterizing the association mode between the TA pairs through various characterization techniques, we found that MazF-mt9 not only bound its cognate antitoxin, but also the non-cognate antitoxin MazE-mt1, a phenomenon that could be also observed in vivo. Based on our structural and biochemical work, we proposed that the cognate and heterologous interactions among different TA systems work together to relieve MazF-mt9s toxicity to M. tb cells, which may facilitate their adaptation to the stressful conditions encountered during host infection.\n\nIMPORTANCETuberculosis (TB) is one of the most severe contagious diseases. Caused by Mycobacterium tuberculosis (M. tb), it poses a serious threat to human health. Additionally, TB is difficult to cure because of the multipledrug-resistant (MDR) and extensively drug-resistant (XDR) M. tb strains. Toxin-antitoxin (TA) systems have been discovered to widely exist in prokaryotic organisms with diverse roles, normally composed of a pair of molecules that antagonize each other. M. tb has ten MazEF systems, and some of them have been proved to be directly associated with the genesis of persisters and drug-resistance of M. tb. We here report the MazEF-mt9 complex structure, and thoroughly characterized the interactions between MazF-mt9 with MazEs within or outside the MazEF-mt9 family. Our study not only revealed the crosstalks between TA families and its significance to M. tb survival but also offers insights into potential anti-TB drug design.

biochemistry

Boosting subdominant neutralizing antibody responses with a computationally designed epitope-focused immunogen

Throughout the last decades, vaccination has been key to prevent and eradicate infectious diseases. However, many pathogens (e.g. respiratory syncytial virus (RSV), influenza, dengue and others) have resisted vaccine development efforts, largely due to the failure to induce potent antibody responses targeting conserved epitopes. Deep profiling of human B-cells often reveals potent neutralizing antibodies that emerge from natural infection, but these specificities are generally subdominant (i.e., are present in low titers). A major challenge for next-generation vaccines is to overcome established immunodominance hierarchies and focus antibody responses on crucial neutralization epitopes. Here, we show that a computationally designed epitope-focused immunogen presenting a single RSV neutralization epitope elicits superior epitope-specific responses compared to the viral fusion protein. In addition, the epitope-focused immunogen efficiently boosts antibodies targeting the Palivizumab epitope, resulting in enhanced neutralization. Overall, we show that epitope-focused immunogens can boost subdominant neutralizing antibody responses in vivo and reshape established antibody hierarchies.

immunology

SuperCT: A supervised-learning-framework to enhance the characterization of single-cell transcriptomic profiles

Characterization of individual cell types is fundamental to the study of multicellular samples such as tumor tissues. Single-cell RNAseq techniques, which allow high-throughput expression profiling of individual cells, have significantly advanced our ability of this task. Currently, most of the scRNA-seq data analyses are commenced with unsupervised clustering of cells followed by visualization of clusters in a low-dimensional space. Clusters are often assigned to different cell types based on canonical markers. However, the efficiency of characterizing the known cell types in this way is low and limited by the investigator[s] knowledge. In this study, we present a technical framework of training the expandable supervised-classifier in order to reveal the single-cell identities based on their RNA expression profiles. Using multiple scRNA-seq datasets we demonstrate the superior accuracy, robustness, compatibility and expandability of this new solution compared to the traditional methods. We use two examples of model upgrade to demonstrate how the projected evolution of the cell-type classifier is realized.

bioinformatics

Detection of cell-type-specific risk-CpG sites in epigenome-wide association studies

In epigenome-wide association studies, the measured signals for each sample are a mixture of methylation profiles from different cell types. The current approaches to the association detection only claim whether a cytosine-phosphate-guanine (CpG) site is associated with the phenotype or not, but they cannot determine the cell type in which the risk-CpG site is affected by the phenotype. Here, we propose a solid statistical method, HIgh REsolution (HIRE), which not only substantially improves the power of association detection at the aggregated level as compared to the existing methods but also enables the detection of risk-CpG sites for individual cell types.

bioinformatics

Environmental DNA for the enumeration and management of Pacific salmon

Pacific salmon are a keystone resource in Alaska, generating annual revenues of well over [~]US$500 million/yr. Due to their anadromous life history, adult spawners distribute amongst thousands of streams, posing a huge management challenge. Currently, spawners are enumerated at just a few streams because of reliance on human counters and, rarely, sonar. The ability to detect organisms by shed tissue (environmental DNA, eDNA) promises a more efficient counting method. However, although eDNA correlates generally with local fish abundances, we do not know if eDNA can accurately enumerate salmon. Here we show that daily, and near-daily, flow-corrected eDNA rate closely tracks daily numbers of returning sockeye and coho spawners and outmigrating sockeye smolts. eDNA thus promises accurate and efficient enumeration, but to deliver the most robust numbers will need higher-resolution stream-flow data, at-least-daily sampling, and a focus on species with simple life histories, since shedding rate varies amongst jacks, juveniles, and adults.

ecology

TP promotes malignant progression in hepatocellular carcinoma through pentose Warburg effect

Tumor progression is dependent on metabolic reprogramming. Metastasis and vasculogenic mimicry (VM) are typical tumor progression. The relationship of metastasis, VM and metabolic reprogramming is not clear. In this study, we identified the novel role of Twist1, a VM regulator, in the transcriptional regulation of the expression of thymidine phosphorylase (TP). We demonstrated that TP promoted extracellular thymidine metabolization into ATP and amino acids through pentose Warburg effect by coupling the pentose phosphate pathway and glycolysis. Moreover, Twist1 relied on TP-induced metabolic reprogramming to promote hepatocellular carcinoma (HCC) metastasis and VM formation mediated by VE-Cad, VEGFR1, and VEGFR2 in vitro and in vivo. TP inhibitor tipiracil reduced promotion effect of TP enzyme activity on HCC VM formation and metastasis. Our findings demonstrate that TP, transcriptionally activated by Twist1, promotes HCC VM formation and metastasis through pentose Warburg effect, contributing to tumor progression.

cell biology

Inducible formation of leading cells driven by CD44 switching gives rise to collective invasion

Collective invasion into adjacent tissue is a hallmark of luminal breast cancer, with about 20% of cases that eventually undergo metastasis. It remained unclear how less aggressive luminal-like breast cancer transit to invasive cancer. Our study revealed that CD44hi cancer cells are the leading subpopulation in collective invading cancer cells, which could efficiently lead the collective invasion of CD44lo/follower cells. CD44hi/leading subpopulation showed specific gene signature of a cohort of hybrid epithelial/mesenchymal state genes and key functional co-regulators of collective invasion, which was distinct from CD44lo/follower cells. However, the CD44hi/leading cells, in partial-EMT state, were readily switching to CD44lo phenotype along with collective movements and vice versa, which is spontaneous and sensitive to tumor microenvironment. The CD44lo-to-CD44hi conversion is accompanied with a shift of CD44s-to-CD44v, but not corresponding to the conversion of non-CSC-to-CSC. Therefore, the CD44hi leader cells are not a stable subpopulation in breast tumors. This plasticity and ability to generate CD44hi carcinoma cells with enhanced invasion-initiating powers might be responsible for the transition from in situ to invasive behavior of luminal-type breast cancer.\n\nSignificanceNow, the mechanisms involved in local invasion and distant metastasis are still unclear. We identified a switch of CD44 that drives leader cell formation during collective invasion in luminal breast cancer. We provided evidence that interconversions between low and high CD44 states occur frequently during collective invasion. Furthermore, these findings demonstrated that the CD44hi/leader cells featuring partial EMT are inducible and attainable in response to tumor microenvironment. The CD44lo cancer cells are plastic that readily shift to CD44hi state, accompanied with shifts of CD44s-to-CD44v, thereby increasing tumorigenic and malignant potential. There are many \"non-invasiveness\" epithelial/follower cells with reversible invasive potential within an individual tumor, that casting some challenges on molecular targeting therapy.

cancer biology

Why panmictic bacteria are rare

BackgroundBacteria typically have more structured populations than higher eukaryotes, but this difference is surprising given high recombination rates, enormous population sizes and effective geographical dispersal in many bacterial species.\n\nResultsWe estimated the recombination scaled effective population size Ner in 21 bacterial species and find that it does not correlate with synonymous nucleotide diversity as would be expected under neutral models of evolution. Only two species have estimates substantially over 100, consistent with approximate panmixia, namely Helicobacter pylori and Vibrio parahaemolyticus. Both species are far from demographic equilibrium, with diversity predicted to increase more than 30 fold in V. parahaemolyticus if the current value of Ner were maintained, to values much higher than found in any species. We propose that panmixia is unstable in bacteria, and that persistent environmental species are likely to evolve barriers to genetic exchange, which act to prevent a continuous increase in diversity by enhancing genetic drift.\n\nConclusionsOur results highlight the dynamic nature of bacterial population structures and imply that overall diversity levels found within a species are poor indicators of its size.

microbiology

Rosetta FunFolDes - a general framework for the computational design of functional proteins

The robust computational design of functional proteins has the potential to deeply impact translational research and broaden our understanding of the determinants of protein function, nevertheless, it remains a challenge for state-of-the-art methodologies. Here, we present a computational design approach that couples conformational folding with sequence design to embed functional motifs into heterologous proteins. We performed extensive benchmarks, where the most unexpected finding was that the design of function into proteins may not necessarily reside in the global minimum of the energetic landscape, which could have important implications in the field. We have computationally designed and experimentally characterized a distant structural template and a de novo \"functionless\" fold, two prototypical design challenges, to present important viral epitopes. Overall, we present an accessible strategy to repurpose old protein folds for new functions, which may lead to important improvements on the computational design of functional proteins.

bioinformatics

Introduce a New Approach to Detect Genes Associated to Oral Squamous Cell Carcinoma

Oral squamous cell carcinoma (OSCC) represents the most frequent of all oral neoplasms in the world. Genetics plays an important role in the etiopathogenesis of OSCC. However, the investigation of the molecular mechanism of OSCC is still incomplete. In this article, we introduced a new approach to detect OSCC-associated genes, in which we not only compare mean difference, but also variance difference between cases and controls. Based on two OSCC datasets from Gene Expression Omnibus, we identified 456 differentially variable (DV) gene probes, in addition to 2,375 differentially expressed (DE) gene probes. There are 2,193 DE-only probes, 274 DV-only probes, and 182 DE-and-DV probes. DAVID functional analysis showed that genes corresponding to DE-only, DV-only, and DE-and-DV probes were enriched in different KEGG pathways, indicating they play different roles in OSCC. This new approach can be used to investigate the genetic risk factors for other complex human diseases.

genetics

The landscape of coadaptation in Vibrio parahaemolyticus

Investigating fitness interactions in natural populations remains a considerable challenge. We take advantage of the unique population structure of Vibrio parahaemolyticus, a bacterial pathogen of humans and shrimp, to perform a genome-wide screen for coadapted genetic elements. We identified 90 interaction groups involving 1,560 coding genes. 82 of these interaction groups are between accessory genes, many of which have functions related to carbohydrate transport and metabolism. Only 8 interaction groups involve both core and accessory genomes. The largest includes 1,540 SNPs in 82 genes and 338 accessory genome elements, many involved in lateral flagella and cell wall biogenesis. The interactions have a complex hierarchical structure encoding at least four distinct ecological strategies. Preliminary experiments imply that the strategies influence biofilm formation and bacterial growth rate in vitro. One strategy involves a divergent profile in multiple genome regions, implying that strains have irreversibly specialized, while the others involve fewer genes and are more plastic. Our results imply that most genetic alliances are ephemeral but that increasingly complex strategies can evolve and eventually cause speciation.

microbiology

Coordinative Metabolism of Glutamine Carbon and Nitrogen in proliferating Cancer Cells Under Hypoxia

Under hypoxia, most of glucose is converted to secretory lactate, which leads to the lack of carbon source from glucose and thus the overuse of glutamine-carbon. However, under such a condition how glutamine nitrogen is disposed to avoid releasing potentially toxic ammonia remains to be determined. Here we identify a metabolic flux of glutamine to secretory dihydroorotate under hypoxia. We found that glutamine nitrogen is indispensable to nucleotide biosynthesis, but enriched in dihyroorotate and orotate rather than processing to its downstream uridine monophosphate under hypoxia. Dihyroorotate, not orotate, is then secreted out of cells. The specific metabolic pathway occurs in vivo and is required for tumor growth. Such a metabolic pathway renders glutamine mainly to acetyl coenzyme A for lipogenesis, with the rest carbon and nitrogen being safely removed. Our results reveal how glutamine carbon and nitrogen are coordinatively metabolized under hypoxia, and provide a comprehensive understanding on glutamine metabolism.\n\nSignificanceTumor cells often addict to glutamine, and particularly utilize its carbon for lipogenesis under hypoxia. We reveal that tumor cells package the excessive glutamine-nitrogen into secretory dihydroorotate, instead of toxic ammonia. This specifically reprogrammed pathway supports in vivo tumor growth, and could offer diagnostic markers and therapeutic targets for cancers.

biochemistry

Recent mixing of Vibrio parahaemolyticus populations

BackgroundHumans have profoundly affected the ocean environment but little is known about anthropogenic effects on the distribution of microbes. Vibrio parahaemolyticus is found in warm coastal waters and causes gastroenteritis in humans and economically significant disease in shrimps.\n\nResultsBased on data from 1,103 genomes, we show that V. parahaemolyticus is divided into four diverse populations, VppUS1, VppUS2, VppX and VppAsia. The first two are largely restricted to the US and Northern Europe, while the others are found worldwide, with VppAsia making up the great majority of isolates in the seas around Asia. Patterns of diversity within and between the populations are consistent with them having arisen by progressive divergence via genetic drift during geographical isolation. However, we find that there is substantial overlap in their current distribution. These observations can be reconciled without requiring genetic barriers to exchange between populations if dispersal between oceans has increased dramatically in the recent past. We found that VppAsia isolates from the US have an average of 1.01% more shared ancestry with VppUS1 and VppUS2 isolates than VppAsia isolates from Asia itself. Based on time calibrated trees of divergence within epidemic lineages, we estimate that recombination affects about 0.017% of the genome per year, implying that the genetic mixture has taken place within the last few decades.\n\nConclusionsThese results suggest that human activity, such as shipping and aquatic products trade, are responsible for the change of distribution pattern of this marine species.

microbiology

Cell type specific profiling of alternative translation identifies novel protein isoforms in the mouse brain

Translation canonically begins at a single AUG and terminates at the stop codon, generating one protein species per transcript. However, some transcripts may use alternative initiation sites or sustain translation past their stop codon, generating multiple protein isoforms. Through other mechanisms such as alternative splicing, both neurons and glia exhibit remarkable transcriptional diversity, and these other forms of post-transcriptional regulation are impacted by neural activity and disease. Here, using ribosome footprinting, we demonstrate that alternative translation is likewise abundant in the central nervous system and modulated by stimulation and disease. First, in neuron/glia mixed cultures we identify hundreds of transcripts with alternative initiation sites and confirm the protein isoforms corresponding to a subset of these sites by mass spectrometry. Many of them modulate their alternative initiation in response to KCl stimulation, indicating activity-dependent regulation of this phenomenon. Next, we detect several transcripts undergoing stop codon readthrough thus generating novel C-terminally-extended protein isoforms in vitro. Further, by coupling Translating Ribosome Affinity Purification to ribosome footprinting to enable cell-type specific analysis in vivo, we find that several of both neuronal and astrocytic transcripts undergo readthrough in the mouse brain. Functional analyses of one of these transcripts, Aqp4, reveals readthrough confers perivascular localization, indicating readthrough can be a conserved mechanism to modulate protein function. Finally, we show that AQP4 readthrough is disrupted in multiple gliotic disease models. Our study demonstrates the extensive and regulated use of alternative translational events in the brain and indicates that some of these events alter key protein properties.

neuroscience

Germline genetics encode the resistance, risk, and lymphatic metastasis of triple-negative breast cancer in the southern Chinese population

Early identification of the risk for triple-negative breast cancer (TNBC) at the asymptomatic phase could lead to better prognosis. Here we developed a machine learning method to quantify systematic impact of all rare germline mutations on each pathway. We collected 106 TNBC patients and 287 elder healthy women controls. The spectra of activity profiles in multiple pathways were mapped and most pathway activities exhibited globally suppressed by the portfolio of individual germline mutations in TNBC patients. Accordingly, all individuals were delineated into two types: A and B. Type A patients could be differentiated from controls (AUC = 0.89) and sensitive to BRCA1/2 damages; Type B patients can be also differentiated from controls (AUC = 0.69) but probably being protected from BRCA1/2 damages. Further we found that Individuals with the lowest activity of selected pathways had extreme high relative risk (up to 21.67 in type A) and increased lymph node metastasis in these patients. Our study showed that genomic DNA contains information of unimaginable pathogenic factors. And this information is in a distributed form that could be applied to risk assessment for more cancer types. SignificanceWe identified individuals who are more susceptible to triple negative breast cancer. Our method performs much better than previous assessments based on BRCA1/2 damages, even polygenic risk scores. We disclosed previously unimaginable pathogens in a distributed form on genome and extended risk prediction to scenarios for other cancers.

cancer biology

Recombinant expression of Proteorhodopsin and biofilm regulators in Escherichia coli for nanoparticle binding and removal in a wastewater treatment model

The small size of nanoparticles is both an advantage and a problem. Their high surface-area-to-volume ratio enables novel medical, industrial, and commercial applications. However, their small size also allows them to evade conventional filtration during water treatment, posing health risks to humans, plants, and aquatic life. This project aims to remove nanoparticles during wastewater treatment using genetically modified Escherichia coli in two ways: 1) binding citrate-capped nanoparticles with the membrane protein Proteorhodopsin, and 2) trapping nanoparticles using Escherichia coli biofilm produced by overexpressing two regulators: OmpR234 and CsgD. We demonstrate experimentally that Escherichia coli expressing Proteorhodopsin binds to 60 nm citrate-capped silver nanoparticles. We also successfully upregulate biofilm production and show that Escherichia coli biofilms are able to trap 30 nm gold particles. Finally, both Proteorhodopsin and biofilm approaches are able to bind and remove nanoparticles in simulated wastewater treatment tanks. We envision integrating our trapping system in both rural and urban wastewater treatment plants to efficiently capture all nanoparticles before treated water is released into the environment.\n\nFinancial DisclosureThis work was funded by the Taipei American School. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.\n\nCompeting InterestsThe authors have declared that no competing interests exist.\n\nEthics StatementN/A\n\nData AvailabilityYes - all data are fully available without restriction. Sequences for the plasmids used in this study are available through the Registry of Standard Biological Parts. Links to raw data are included in Supplementary Information.

synthetic biology

PKD2 influence uric acid levels and gout risk by interacting with ABCG2

BackgroundUric acid is the final product of purine metabolism and elevated serum urate levels can cause gout. Conflicting results were reported for the effect of PKD2 on serum urate levels and gout risk. Therefore, our study attempted to state the important role of PKD2 in influencing the pathogenesis of gout.\n\nMethodSNPs in PKD2 (rs2725215 and rs2728121) and ABCG2 (rs2231137 and rs1481012) were tested in approximately 5,000 Chinese individuals.\n\nResultsTwo epistatic interactions between loci in PKD2 (rs2728121) and ABCG2 (rs1481012 and rs2231137) showed distinct contributions to uric acid levels with P int values of 0.018 and 0.004, respectively, and the associations varies by gender and BMI. The SNP pair of rs2728121 and rs1481012 justly played roles in uric acid in females (P int = 0.006), while the other pair did in males (P int = 0.017). Regarding BMI, the former SNP pair merely contributed in overweigh subjects (P int = 0.022) and the latter one did in both normal and overweigh individuals (P int = 0.013 and 0.047, respectively). Furthermore, the latter SNP pair was also associated with gout pathology (P int = 0.001), especially in males (P int = 0.001). Finally, functional analysis showed potential epistatic interactions in those genes region and PKD2 mRNA expression had a positive correlation with ABCG2s (r = 0.743, P = 5.83e-06).\n\nConclusionOur study for the first time identified that epistatic interactions between PKD2 and ABCG2 influenced serum urate concentrations and gout risk, and PKD2 might affect the pathogenesis from elevated serum urate to hyperuricemia to gout by modifying ABCG2.

genetics