Search bioRxiv⌕ Search

Biology subjects

Behera, A. K.

Publications and source records attributed to Behera, A. K..

4 recordsLinked to original sources

RNA-coupled CRISPR Screens Reveal ZNF207 as a Regulator of LMNA Aberrant Splicing in Progeria

Despite progress in understanding pre-mRNA splicing, the regulatory mechanisms controlling most alternative splicing events remain unclear. We developed CRASP-Seq, a method that integrates pooled CRISPR-based genetic perturbations with deep sequencing of splicing reporters, to quantitively assess the impact of all human genes on alternative splicing from a single RNA sample. CRASP-Seq identifies both known and novel regulators, enriched for proteins involved in RNA splicing and metabolism. As proof-of-concept, CRASP-Seq analysis of an LMNA cryptic splicing event linked to progeria uncovered ZNF207, primarily known for mitotic spindle assembly, as a regulator of progerin splicing. ZNF207 depletion enhances canonical LMNA splicing and decreases progerin levels in patient-derived cells. High-throughput mutagenesis further showed that ZNF207s zinc finger domain broadly impacts alternative splicing through interactions with U1 snRNP factors. These findings position ZNF207 as a U1 snRNP auxiliary factor and demonstrate the power of CRASP-Seq to uncover key regulators and domains of alternative splicing. Main PointsO_LICRASP-Seq: RNA-coupled CRISPR screen quantifying gene and domain impact on splicing C_LIO_LIProfiling of five events identified 370 genes influencing alternative splicing C_LIO_LIZNF207 regulates splicing by interacting with U1 snRNP via its zinc-finger domains C_LIO_LIZNF207 depletion corrects LMNA aberrant splicing causing progeria C_LI

genetics↗

Copy number variation analysis of 9,482 Mycobacterium tuberculosis isolates identifies lineage-specific molecular determinants.

BackgroundClinical manifestations of tuberculosis (TB) caused by Mycobacterium tuberculosis (Mtb) show lineage-specific differences contributed by genetic polymorphism such as phylo-single nucleotide variations (PhyloSNPs) and insertion or deletions (INDELs). Intragenomic rearrangement events, such as gene duplications and deletions, may cause gene copy number differences in Mtb, contributing to lineage-specific phenotypic variations, if any, which need better understanding. ResultsThe relative gene copy number differences in high-quality publicly available whole genome sequencing datasets of 9,482 clinical Mtb isolates were determined by repurposing and modifying an RNA-seq data analysis pipeline. The pipeline included various steps, viz., alignment of reads, sorting by coordinate, GC bias correction, and variant stabilising transformation. The strategy showed maximum separation of lineage-specific clusters in two principal components, capturing [~]54% variability. Unsupervised hierarchical clustering of the top 100 genes and pairwise comparisons between Mtb lineages revealed an overlapping subset of genes (n=42) having significantly perturbed copy numbers (Benjamin Hochberg adjusted P-value < 0.05 and log2(drug-resistant/sensitive) > {+/-} 1). These 42 genes formed multiple tandem gene clusters and are known to be involved in virulence, pathogenicity and defence response to invading phages. A separate comparison showed a significantly high copy number of phage genes and a recently reported druggable target Rv1525 in pre- and extensively drug-resistant (Pre-XDR, XDR) compared to drug-sensitive clinical Mtb isolates. ConclusionThe identified gene sets in Mtb clinical isolates may be useful targets for lineage-specific therapeutics and diagnostics development.

genomics↗

Systematic assessment of long-read RNA-seq methods for transcript identification and quantification

AbstractThe Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.

genomics↗

Structural, functional and evolutionary analysis of wheat WRKY45 protein: A combined bioinformatics and MD simulation approach

Bread wheat (Triticum aestivum L.) is the worlds second-most important cereal crop, as well as Indias. It is an allohexaploid composed of three homeologous sub-genomes (AA, BB, and DD), which is a constraint in determining the complete genome sequence. Several transcription factors have been implicated in both abiotic and biotic stress. WRKY transcription factors are among the best characterised in the context of pathogen defence mechanisms. Different members of the WRKY transcription factors have been shown to confer resistance to stress. But very little is known about the wheat WRKY transcription factors. In silico analysis of the TaWRKY45 protein was performed in the present study using several bioinformatics tools like motif scan, CD search, Netphos, NGlycos, GRAVY, and the SWISS MODEL. The study revealed that TaWRKY45 belongs to the group III family and contains hydrophilic proteins with 19 potential phosphorylation sites. TaWRKY45 protein was found to be orthologous to rice OsWRKY45 by phylogenetic analysis. The catalytic domain was analysed by motif scan which showed that TaWRKY45 has one WRKY domain and a C2-HC zinc finger motif. TaWRKY45s structure was determined to be more stable, more constrained, more compact, and have greater potential to interact with other molecules than OsWRKY45, according to MD simulation analysis. Thus, in silico analysis of transcription factors helps study protein function, interaction, and regulatory pathways.

bioinformatics↗