Search bioRxiv⌕ Search

Biology subjects

Kazmi, A.

Publications and source records attributed to Kazmi, A..

2 recordsLinked to original sources

Beyond the Hype: The Complexity of Automated Cell Type Annotations with GPT-4

Recent research has shown the impressive capability of large language models like GPT-4 in various downstream tasks in single-cell data analysis. Among these tasks, cell type annotation remains particularly challenging, with researchers exploring various methods to improve accuracy and efficiency. While recent studies on GPT-like models have demonstrated annotation performance comparable to manual annotations, a significant gap remains in understanding their limitations and generalizability. In this work, we compare and evaluate the annotation performance of the GPT-4 model against traditional methods on nine randomly selected public single-cell RNA seq datasets from cellxgene, covering diverse tissue types. Our evaluation highlights the complexity of annotating cell types in single-cell data, revealing key differences between automated and manual approaches. We found specific cases where GPT-4 underperforms, demonstrating its limitations in certain contexts. We further introduce an automated approach to incorporate literature search using a RAG approach which enhances and outperforms GPT-4 cell type annotation when compared to traditional methods. We also introduce metrics based on taxonomic distance in the ontology tree to evaluate the granularity of the cell type annotations. To support future research, we also release an open-source Python package 1 that enables fully automated cell-type annotation of single-cell data using GPT-4 alongside other methods. The pipeline can take paper as an input and do cell type annotations on its own.

bioinformatics↗

Bioinformatics techniques for efficient structure prediction of SARS-CoV-2 protein ORF7a via structure prediction approaches

Protein is the building block for all organisms. Protein structure prediction is always a complicated task in the field of proteomics. DNA and protein databases can find the primary sequence of the peptide chain and even similar sequences in different proteins. Mainly, there are two methodologies based on the presence or absence of a template for Protein structure prediction. Template-based structure prediction (threading and homology modeling) and Template-free structure prediction (ab initio). Numerous web-based servers that either use templates or do not can help us forecast the structure of proteins. In this current study, ORF7a, a transmembrane protein of the SARS-coronavirus, is predicted using Phyre2, IntFOLD, and Robetta. The protein sequence is straightforwardly entered into the sequence bar on all three web servers. Their findings provided information on the domain, the region with the disorder, the global and local quality score, the predicted structure, and the estimated error plot. Our study presents the structural details of the SARS-CoV protein ORF7a. This immunomodulatory component binds to immune cells and induces severe inflammatory reactions.

microbiology↗