Search bioRxiv⌕ Search

Biology subjects

Maiti, T.

Publications and source records attributed to Maiti, T..

8 recordsLinked to original sources

Multiview Graph Learning for single-cell RNA sequencing data

Characterizing the underlying topology of gene regulatory networks is one of the fundamental problems of systems biology. Ongoing developments in high throughput sequencing technologies has made it possible to capture the expression of thousands of genes at the single cell resolution. However, inherent cellular heterogeneity and high sparsity of the single cell datasets render void the application of regular Gaussian assumptions for constructing gene regulatory networks. Additionally, most algorithms aimed at single cell gene regulatory network reconstruction, estimate a single network ignoring group-level (cell-type) information present within the datasets. To better characterize single cell gene regulatory networks under different but related conditions we propose the joint estimation of multiple networks using multiview graph learning (mvGL). The proposed method is developed based on recent works in graph signal processing (GSP) for graph learning, where graph signals are assumed to be smooth over the unknown graph structure. Graphs corresponding to the different datasets are regularized to be similar to each other through a learned consensus graph. We further kernelize mvGL with the kernel selected to suit the structure of single cell data. An efficient algorithm based on prox-linear block coordinate descent is used to optimize mvGL. We study the performance of mvGL using synthetic data generated with a diverse set of parameters. We further show that mvGL successfully identifies well-established regulators in a mouse embryonic stem cell differentiation study and a cancer clinical study of medulloblastoma.

genomics↗

In silico analysis for potential proteins and microRNAs in Glioblastoma and Parkinsonism.

In todays world, neurodegenerative diseases such as Alzheimers disease, Parkinsons Disease, Huntingtons Disease as well as brain cancers such as astrocytomas, ependymomas, glioblastomas have become a great threat to us. In this study, we are trying to find a probable molecular connection associated with two very much different diseases, Glioblastoma, also known as Glioblastoma Multiforme (cancers of microglial cells of our brain) and Parkinsons disease. We at first downloaded the microarray datasets of these two diseases from Gene Expression Omnibus (GEO) and then analyzed them by the GEO2R tool. After analysis, we found 249 common upregulated differential expressed genes and 135 common downregulated differential expressed genes of these two diseases. Therefore the common differentially expressed genes, both upregulated and downregulated, were imported into STRING online tool to find out the protein-protein interactions. Now, this whole network was subjected to Cytoscape and the top ten hub genes were found by Cyto-Hubba plug-in. The top then hub genes are EGFR, CCNB1, CDK1, CCNA2, CHEK1, RAD51, MAD2L1, KIF20A, BUB1, and CCNB2. These all genes are upregulated in both diseases. To find out the biological processes, molecular functions, cellular components, and pathways associated with these hub genes Enrichr online software was used. We used miRNet software to determine the interactions of hub genes with microRNAs. This study will be useful in the future for drug targets discovery for these diseases.

neuroscience↗

Identification of Common genes and proteins in Alzheimer's Disease, Multiple Sclerosis and Duchenne Muscular Dystrophy using in-silico methods.

Alzheimers is a type of dementia symptom that slowly worsens over some time. In its early stages, memory loss is mild, but with late-stage Alzheimers, the patient loses the ability to carry on a conversation and respond to their environment. Multiple sclerosis (MS) is an autoimmune disease of the central nervous system (CNS) characterized by chronic inflammation, demyelination, gliosis, and neuronal loss. Duchenne muscular dystrophy (DMD) is one of the most severe forms of inherited muscular dystrophies. It is the most common hereditary neuromuscular disease and does not exhibit a predilection for any race or ethnic group. In this study, several gene expression study data were analyzed and there were 557 Differentially Expressed Genes (DEGs) in all three chosen Datasets. A protein-protein interaction network was created using STRING and CytoHubba plug-in was used to identify the top ten genes which are POLR2A, SETD2, EFTUD2, RBM25, PRPF40A, CDK13, BPTF, THOC2, SNRNP70, and SCAF11. Online software Enrichr was used for Gene Ontology and KEGG pathway enrichment analysis to find out the biological process, molecular function, cellular component, and the pathways that are commonly affected in these diseases.

neuroscience↗

In-silico analysis: common biomarkers of NDs

Neurodegenerative disorders (NDs) are a class of rapidly rising devastating diseases and the reason behind are might be an improper function of related genes or a mutation in a particular gene or even could be autoimmune also. Parkinsons disease (PD), Multiple sclerosis (MS), Huntingtons disease (HD) are some of the NDs, and still, incurable fully. Apart from the similarities in symptoms, there are common genes that express somehow a differential manner in patients of PDs, MSs, and HDs. A total of 1197 differentially expressed genes (DEGs) are obtained by analyzing the chosen datasets. The protein interactions by STRING online tool and degree sorted hubs obtained through a plug-in in Cytoscape; Cyto-Hubba. Among the sorted hubs KRAS, CREB1, PIK3CA, JAK2 are the ones that are not only common to all the studied datasets of NDs but also in other neurological disorders like Alzheimers. The enriched pathways with biological process, molecular function, cellular component, and KEGG pathway details are obtained and analyzed using Enricher. This paper frames that the obtained hub genes could be potential biomarkers also and a need for further drug design for finding a possible cure.

neuroscience↗

Identification of potential proteins and microRNAs in Multiple Sclerosis and Huntington's diseases using in silico methods.

Neurogenerative diseases like multiple sclerosis, Huntingtons disease are the major roadblocks in the way towards a healthy brain. Neurodegenerative diseases like multiple sclerosis and Huntingtons disease are affected by several factors such as environmental, immunological, genetics, and the worse scenario we can think of is that they are on the rise worldwide. Degenerative diseases specifically target a limited group of neurons at first resulting in the loss of specific functions associated with the specific part of the brain. The early diagnosis of these neurodegenerative diseases is important so that treatments can start from the early stages of these diseases. In this study, we have established a link between Multiple sclerosis and Huntingtons disease, and also we were able to establish the possible microRNAs that were connected to the expression of genes associated with these two diseases. In this present study, we analyzed the microarray datasets obtained from Gene Expression Omnibus and we identified 266 differentially expressed genes tried to identify using in silico methods the Hub genes involved in Multiple sclerosis and Huntingtons disease. After identifying the genes and proteins we tried to identify the microRNAs that are interacting with the Hub genes. In our study, we identified that the protein network has PTPRC, CXCL8, RBM25 proteins that have maximum connectivity. The top Hub genes are then subjected to a database that contains information concerning the microRNAs that are interacting with the Hub proteins as well as with each other. According to our study, the hsa-mir-155-5p has one of the highest degrees in the microRNA network. Our study will be useful in the future for the development of new drug targets for these neurodegenerative diseases.

neuroscience↗

Benchmarking of a Bayesian single cell RNAseq differential gene expression test for dose-response study designs.

The application of single-cell RNA sequencing (scRNAseq) for the evaluation of chemicals, drugs, and food contaminants presents the opportunity to consider cellular heterogeneity in pharmacological and toxicological responses. Current differential gene expression analysis (DGEA) methods focus primarily on two group comparisons, not multi-group dose-response study designs used in safety assessments. To benchmark DGEA methods for dose-response scRNAseq experiments, we proposed a multiplicity corrected Bayesian testing approach and compare it against 8 other methods including two frequentist fit-for-purpose tests using simulated and experimental data. Our Bayesian test method outperformed all other tests for a broad range of accuracy metrics including control of false positive error rates. Most notable, the fit-for-purpose and standard multiple group DGEA methods were superior to the two group scRNAseq methods for dose-response study designs. Collectively, our benchmarking of DGEA methods demonstrates the importance in considering study design when determining the most appropriate test methods.

genomics↗

Light Potentials of Photosynthetic Energy Storage in the Field: What limits the ability to use or dissipate rapidly increased light energy?

The responses of plant photosynthesis to rapid fluctuations in environmental conditions are thought to be critical for efficient capture of light energy. Such responses are not well represented under laboratory conditions, but have also been difficult to probe in complex field environments. We demonstrate an open science approach to this problem that combines multifaceted measurements of photosynthesis and environmental conditions, and an unsupervised statistical clustering approach. In a selected set of data on mint (Mentha sp.), we show that the "light potential" for increasing linear electron flow (LEF) and nonphotochemical quenching (NPQ) upon rapid light increases are strongly suppressed in leaves previously exposed to low ambient PAR or low leaf temperatures, factors that can act both independently and cooperatively. Further analyses allowed us to test specific mechanisms. With decreasing leaf temperature or PAR, limitations to photosynthesis during high light fluctuations shifted from rapidly-induced NPQ to photosynthetic control (PCON) of electron flow at the cytochrome b6f complex. At low temperatures, high light induced lumen acidification, but did not induce NPQ, leading to accumulation of reduced electron transfer intermediates, a situation likely to induce photodamage, and represents a potential target for improving the efficiency and robustness of photosynthesis. Finally, we discuss the implications of the approach for open science efforts to understand and improve crop productivity.

plant biology↗

scSGL: Signed Graph Learning for Single-Cell Gene Regulatory Network Inference

MotivationElucidating the topology of gene regulatory networks (GRNs) from large single-cell RNA sequencing (scRNAseq) datasets, while effectively capturing its inherent cell-cycle heterogeneity and dropouts, is currently one of the most pressing problems in computational systems biology. Recently, graph learning (GL) approaches based on graph signal processing (GSP) have been developed to infer graph topology from signals defined on graphs. However, existing GL methods are not suitable for learning signed graphs, which represent a characteristic feature of GRNs, as they account for both activating and inhibitory relationships between genes. They are also incapable of handling high proportion of zero values, which represent dropouts in single cell experiments. To this end, we propose a novel signed GL approach, scSGL, that learns GRNs based on the assumption of the smoothness and non-smoothness of gene expressions over activating and inhibitory edges, respectively. scSGL is then extended with kernels to take the nonlinearity of co-expressions into account and handle high proportion of dropouts. From GSP perspective, this extension corresponds to assuming smoothness/non-smoothness of graph signals in a higher dimensional space defined by the kernel. The proposed approach is formulated as a non-convex optimization problem and solved using an efficient ADMM framework. ResultsIn our experiments on simulated and real single cell datasets, scSGL compares favorably with other single cell gene regulatory network reconstruction algorithms. AvailabilityThe scSGL code and analysis scripts are available at (https://github.com/Single-Cell-Graph-Learning/scSGL).

bioinformatics↗