Search bioRxiv⌕ Search

Biology subjects

Lyu, A.

Publications and source records attributed to Lyu, A..

5 recordsLinked to original sources

Accurate and interpretable gene expression imputation on scRNA-seq data using IGSimpute

Single-cell RNA-sequencing (scRNA-seq) enables the quantification of gene expression at the transcriptomic level with single-cell resolution, enhancing our understanding of cellular heterogeneity. However, the excessive missing values present in scRNA-seq data (termed dropout events) hinder downstream analysis. While numerous imputation methods have been proposed to recover scRNA-seq data, high imputation performance often comes with low or no interpretability. Here, we present IGSimpute, an accurate and interpretable imputation method for recovering missing values in scRNA-seq data with an interpretable instance-wise gene selection layer. IGSimpute outperforms ten other state-of-the-art imputation methods on nine tissues of the Tabula Muris atlas with the lowest mean squared error as the chosen benchmark metric. We demonstrate that IGSimpute can give unbiased estimates of the missing values compared to other methods, regardless of whether the average gene expression values are small or large. Clustering results of imputed profiles show that IGSimpute offers statistically significant improvement over other imputation methods. By taking the heart-and-aorta and the limb muscle tissues as examples, we show that IGSimpute can also denoise gene expression profiles by removing outlier entries with unexpected high expression values via the instance-wise gene selection layer. We also show that genes selected by the instance-wise gene selection layer could indicate the age of B cells from bladder fat tissue of the Tabula Muris Senis atlas. IGSimpute has linear time-complexity with respect to cell number, and thus applicable to large datasets.

bioinformatics↗

A machine learning model for disease risk prediction by integrating genetic and non-genetic factors

Polygenic risk score (PRS) has been widely used to identify the high-risk individuals from the general population, which would be helpful for disease prevention and early treatment. Many methods have been developed to calculate PRS by weighted aggregating the phenotype-associated risk alleles from genome-wide association studies. However, only considering genetic effects may not be sufficient for risk prediction because the disease risk is not only related to genetic factors but also non-genetic factors, e.g., diet, physical exercise et al. But it is still a challenge to integrate these genetic and non-genetic factors into a unified machine learning framework for disease risk prediction. In this paper, we proposed PRSIMD (PRS Integrating Multi-source Data), a machine learning model that applies posterior regularization to integrate genetic and non-genetic factors to improve disease risk prediction. Also, we applied Mendelian Randomization analysis to identify the causal non-genetic risk factors for the selected diseases. We applied PRSIMD to predict type 2 diabetes and coronary artery disease from UK Biobank and observed that PRSIMD was significantly better than the methods to calculate PRS including p-value threshold (P+T), PRSice2, SBLUP, DMSLMM, and LDpred2. In addition, we observed that PRSIMD achieved the better predictive power than the composite risk score.

bioinformatics↗

Effect of habitual reading direction on saccadic eye movements

Cognitive processes can influence the characteristics of saccadic eye movements. Reading habits, including habitual reading direction, also affects cognitive and visuospatial processes, favouring attention to the side where reading begins. Few studies have investigated the effect of habitual reading direction on saccade directionality of low-cognitive-demand stimuli (such as dots). The current study examined horizontal prosaccade, antisaccade and self-paced saccade in subjects with two primary habitual reading directions. We hypothesised that saccades responding to the target in subjects habitual reading direction would show a longer prosaccade latency and lower antisaccade error rate (errors being a reflexive glance to a sudden-appearing target, rather than a saccade away from it). Sixteen young Chinese participants with primary habitual reading direction from left to right and sixteen young Arabic and Persian participants with primary habitual reading direction from right to left were recruited. Subjects needed to look towards a 5{degrees}/ 10{degrees}target in the prosaccade task or look towards the mirror image location of the target in the antisaccade task and look between two 10-degree targets in the self-paced saccade task. Only Arabic and Persian participants showed a shorter and directional prosaccade latency towards 5{degrees}target against their habitual reading direction. No significant effect of primary reading direction on antisaccade latency towards the correct directions was found. However, we found that Chinese readers generated significantly shorter prosaccade latencies and higher antisaccade directional errors compared with Arabic and Persian readers. The present study provides an insight into the effect of reading habits on saccadic eye movements in response to low-cognitive-demand stimuli and offers a platform for future studies to investigate the relationship between reading habits and neural mechanisms of eye movement behaviours.

neuroscience↗

Integrin signaling is critical for myeloid-mediated support of T-cell acute lymphoblastic leukemia

We previously found that T-cell acute lymphoblastic leukemia (T-ALL) requires support from tumor-associated myeloid cells, which activate IGF1R signaling in the leukemic blasts. However, IGF1 is not sufficient to sustain T-ALL survival in vitro, implicating additional myeloid-mediated signals in T-ALL progression. Here, we find that T-ALL cells require close contact with myeloid cells to survive. Transcriptional profiling and in vitro assays demonstrate that integrin-mediated cell adhesion and activation of the downstream FAK/PYK2 kinases are required for myeloid-mediated support of T-ALL cells and promote IGF1R activation. Consistent with these findings, inhibition of integrins or FAK/PYK2 signaling diminishes leukemia burden in multiple organs and confers a survival advantage in a mouse model of T-ALL. Inhibiting integrin-mediated cell adhesion or FAK/PYK2 also diminishes survival of primary patient T-ALL cells co-cultured with myeloid cells. Furthermore, elevated integrin pathway gene signatures correlate significantly with myeloid enrichment and an inferior prognosis in pediatric T-ALL patients. Statement of significanceAlthough tumor-associated myeloid cells provide critical support for T-ALL, our understanding of the underlying mechanisms remains limited. This study reveals that integrin-mediated adhesion and signaling are key mechanisms by which myeloid cells promote survival and progression of T-ALL blasts in the leukemic microenvironment.

cancer biology↗

dynDeepDRIM: a dynamic deep learning model to infer direct regulatory interactions using single cell time-course gene expression data

Time-course single-cell RNA sequencing (scRNA-seq) data have been widely applied to reconstruct the cell-type-specific gene regulatory networks by exploring the dynamic changes of gene expression between transcription factors (TFs) and their target genes. The existing algorithms were commonly designed to analyze bulk gene expression data and could not deal with the dropouts and cell heterogeneity in scRNA-seq data. In this paper, we developed dynDeepDRIM that represents gene pair joint expression as images and considers the neighborhood context to eliminate the transitive interactions. dynDeepDRIM integrated the primary image, neighbor images with time-course into a four-dimensional tensor and trained a convolutional neural network to predict the direct regulatory interactions between TFs and genes. We evaluated the performance of dynDeepDRIM on five time-course gene expression datasets. dynDeepDRIM outperformed the state-of-the-art methods for predicting TF-gene direct interactions and gene functions. We also observed gene functions could be better performed if more neighbor images were involved.

bioinformatics↗