Search bioRxivSearch

Biology subjects

Jun Z Li

Publications and source records attributed to Jun Z Li.

4 recordsLinked to original sources

High-density SNP array and genome sequencing reveal signatures of selection in a divergent selection rat model for aerobic running capacity

We have previously established two lines of rat for studying the functional basis of aerobic exercise capacity (AEC) and its impact on metabolic health. The two lines, high capacity runners (HCR) and low capacity runners (LCR), have been selectively bred for high and low intrinsic AEC, respectively. They were started from the same genetically heterogeneous population and have now diverged in both AEC and many other physiological measures, including weight, body composition, blood pressure, body mass index, lung capacity, lipid and glucose metabolism, and natural life span. In order to exploit this rat model to understand the genomic regions under differential selection within the two lines, we used SNP genotype and whole genome pooled sequencing data to identify signatures of selection using three different statistics: runs of homozygosity, fixation index, and aberrant allele frequency spectrum, and developed a composite score that combined the three signals. We found that several pathways (ATP transport and fatty acid metabolism) are enriched in regions under differential selection. The candidate genes and pathways under selection will be integrated with the previous mRNA expression data and future F2 QTL results for a multi-omics approach to understanding the biological basis of AEC and metabolic traits.

Genetics

Cancer classification in the genomic era: five contemporary problems

Classification is an everyday instinct as well as a full-fledged scientific discipline. Throughout the history of medicine, disease classification is central to how we develop knowledge, make diagnosis, and assign treatment. Here we discuss the classification of cancer, the process of categorizing cancer subtypes based on their observed clinical and biological features. Traditionally, cancer nomenclature is primarily based on organ location, e.g., \"lung cancer\" designates a tumor originating in lung structures. Within each organ-specific major type, finer subgroups can be defined based on patient age, cell type, histological grades, and sometimes molecular markers, e.g., hormonal receptor status in breast cancer, or microsatellite instability in colorectal cancer. In the past 15+ years, high-throughput technologies have generated rich new data regarding somatic variations in DNA, RNA, protein, or epigenomic features for many cancers. These data, collected for increasingly large tumor collections, have provided not only new insights into the biological diversity of human cancers, but also exciting opportunities to discover previously unrecognized cancer subtypes. Meanwhile, the unprecedented volume and complexity of these data pose significant challenges for biostatisticians, cancer biologists, and clinicians alike. Here we review five related issues that represent contemporary problems in cancer taxonomy and interpretation. 1. How many cancer subtypes are there? 2. How can we evaluate the robustness of a new classification system? 3. How are classification systems affected by intratumor heterogeneity and tumor evolution? 4. How should we interpret cancer subtypes? 5. Can multiple classification systems coexist? While related issues have existed for a long time, we will focus on those aspects that have been magnified by the recent influx of complex multi-omics data. Ongoing exploration of these problems is essential for data-driven refinement of cancer classification and the successful application of these concepts in precision medicine.

Genomics

Selection-, age-, and exercise-dependence of skeletal muscle gene expression patterns in a rat model of metabolic fitness

Aerobic exercise capacity can influence many complex traits including obesity and type 2 diabetes. We established two rat lines by divergent selection of intrinsic aerobic capacity. The high capacity runners (HCR) and low capacity runners (LCR) differed by ~9-fold in aerobic capacity after 32 generations, and diverged in body fat, blood glucose, and other health indicators. To study the interplay among genetic differentiation, age, and strenuous exercise, we performed microarray-based gene expression analyses in skeletal muscle with a 2x2x2 design to compare HCR and LCR, old and young animals, and between rest and exhaustion, for a total of eight groups (n=6 each). Transcripts for mitochondrial function are expressed higher in HCR than LCR at both rest and exhaustion, for both age groups. Extracellular matrix components decrease with age in both lines and both rest and exhaustion. Interestingly, age-effects in many pathways are more pronounced in LCR, suggesting that HCRs higher innate aerobic capacity underlies both increased lifespan and heathspan.

Genomics

A reassessment of consensus clustering for class discovery

Consensus clustering (CC) is an unsupervised class discovery method widely used to study sample heterogeneity in high-dimensional datasets. It calculates \"consensus rate\" between any two samples as how frequently they are grouped together in repeated clustering runs under a certain degree of random perturbation. The pairwise consensus rates form a between-sample similarity matrix, which has been used (1) as a visual proof that clusters exist, (2) for comparing stability among clusters, and (3) for estimating the optimal number (K) of clusters. However, the sensitivity and specificity of CC have not been systemically studied. To assess its performance, we investigated the most common implementations of CC; and compared CC with other popular methods that also focus on cluster stability and estimation of K. We evaluated these methods using simulated datasets with either known structure or known absence of structure. Our results showed that (1) CC was able to divide randomly generated unimodal data into pre-specified numbers of clusters, and was able to show apparent stability of these chance partitions of known cluster-less data; (2) for data with known structure, the proportion of ambiguously clustered (PAC) pairs infers the known number of clusters more reliably than several commonly used K estimating methods; and (3) validation of the optimal K by choosing the most discriminant genes from the discovery cohort and applying them in an independent cohort often exaggerates the confidence in K due to inherent gene-gene correlations among the selected genes. While these results do not yet prove that any of the published studies using CC has generated false positive findings, they show that datasets with subtle or no structure are fully capable of producing strong evidence of consensus clustering. We therefore recommend caution is using CC in class discovery and validation.

Bioinformatics