Search bioRxiv⌕ Search

Biology subjects

Reinosa, R.

Publications and source records attributed to Reinosa, R..

5 recordsLinked to original sources

NSeqVerify: An Easy-to-Use Desktop Suite for Integrated NGS Data Analysis, from Raw Reads to Taxonomic Assignment

MotivationThe proliferation of next-generation sequencing (NGS) data has created a computational bottleneck, especially for researchers lacking specialized bioinformatics training. Standard analysis workflows require mastering multiple command-line tools, hindering exploratory data analysis and delaying scientific discovery. ResultsThis work presents NSeqVerify, a new cross-platform, open-source desktop software developed in Java, designed to overcome these barriers. NSeqVerify implements a fully integrated genomic workflow within a single, intuitive graphical user interface (GUI). The suite includes: (1) a preprocessing module for quality control and filtering of FASTQ files; (2) a de novo assembler employing a sophisticated De Bruijn graph algorithm with an iterative multi-k-mer strategy to maximize contiguity; and (3) a taxonomic assignment module that automates BLAST searches against NCBI databases and displays the results in an easily interpretable tabular format. The tool was validated through controlled use cases, demonstrating its ability to accurately reconstruct reference viral genomes (HIV-1) and to deconvolute metagenomic mixtures (HIV-1 and SARS-CoV-2). The final test consisted of analyzing a real elephant fecal virome (SRA: SRR35776009), where NSeqVerify successfully assembled contigs -- two of which overlapped and appeared to form a partial 1555 bp genome of a putative Smacovirus, enabling the identification of its capsid protein and the prediction of its 3D structure using AlphaFold. Conclusion and AvailabilityNSeqVerify democratizes NGS data analysis, providing a robust "all-in-one" solution that empowers molecular biologists, students, and clinicians to perform end-to-end genomic analyses. The software is freely available under the GNU GPLv3 license at (https://github.com/roberto117343/NSeqVerify). Contactroberto117343@gmail.com

bioinformatics↗

OrfViralScan 3.0: An intuitive tool for the identification and tracking of open reading frames in viral genomes

The identification and analysis of open reading frames (ORFs) are fundamental steps in genome annotation, which requires accurate bioinformatics tools. OrfViralScan 3.0 is presented, a desktop application developed in Java (requires Java 11 to run), featuring an intuitive graphical user interface (GUI) designed to facilitate the annotation and tracking of ORFs. The program offers functionalities to search for ORFs (initiated by ATG) in individual sequences (ORF Search), track specific ORFs based on length and location across multiple genomes or sequence sets (Track Specific ORF), preprocess FASTA files to standardize formatting (Preprocess File), and split large genomes into manageable fragments (Divide Into Fragments). The utility of OrfViralScan 3.0 is demonstrated through the analysis of the SARS-CoV-2 reference genome (NC_045512.2), the successful tracking of the Spike protein in 983 out of 1000 complete viral genomes, and the preparation of the Escherichia coli genome (NC_000913.3) for fragmented analysis. The softwares capabilities, limitations, and potential future applications are discussed. An example is also included featuring the SARS-CoV-2 Spike protein, showing the folded ORF obtained with OrfViralScan 3.0 using AlphaFold 3. The programs source code is available on GitHub under the GNU GPLv3 license. Contactroberto117343@gmail.com

bioinformatics↗

Bioinformatic analysis of the variability of Tax1 and Tax2 proteins of the human T-cell lymphotropic virus (HTLV)

MotivationThe pathological differentiation in humans between HTLV-1 and HTLV-2 viruses is due, among other parameters, to the variability in a viral protein, Tax (Tax1 and Tax2 respectively), with the second virus being a candidate to interfere with HIV-1 infection, which has allowed us to assess the importance of further investigating Tax. ResultsBoth the degree of conservation and the genomic and protein variability between Tax1 and Tax2 have been expressed in this work through bioinformatics tools. The knowledge about Tax2 related to the interfering capacity on HIV-1 compared to the absence of such a function in Tax1 is corroborated with our analyses, highlighting certain variations between both sequences that could explain the functional differences in pathogenicity that we show in this work. Contactroberto117343@gmail.com; hwfcojavier32@gmail.com; hwfcojavier32@hotmail.com

bioinformatics↗

Bioinformatics Study of Genetic Variability in Obelisks

MotivationThe importance of in silico genetic variability analyses of new biological entities, such as obelisks, lies in their potential to serve as a starting point for experimental studies. The study of obelisks could open new horizons in the field of biological sciences. ResultsSeveral findings have been determined, including high conservation among the sequences within each cluster, despite the presence of significant nucleotide mutations; high variability among the Oblin-1 proteins found in the consensus sequences of each cluster, despite all sharing a known conserved region; structural consistency in the 3D model of the Oblin-2 consensus with previous studies; a series of nucleotide and protein- level consensus sequences for each cluster, where nucleotide-level patterns show color- coded conservation percentages that could be used in further research, such as primer or probe design, and protein-level sequences are presented in FASTA format for Oblin- 1, Oblin-2, and unknown Open Reading Frames (ORFs).

bioinformatics↗

EpiMolBio: A Novel User-Friendly Bioinformatic Program for Genetic Variability Analysis

MotivationGenetic sequence analysis has become essential in many medicine, biology, and epidemiology fields. However, the currently available tools can pose a challenge for users without advanced computational skills. ResultsWe present EpiMolBio, a free-to-use software designed with an intuitive, user-friendly interface that enables a broad spectrum of users to explore genetic variability. Its diverse toolkit encompasses sequence processing, conservation and variability analysis, consensus sequence generation, and genome mutation or amino acid changes identification, including specialized tools for HIV and SARS-CoV-2 analysis. AvailabilityFreely available on the web at https://www.epimolbio.com and https://github.com/EpiMolBio/EpiMolBio. Contactafrica.holguin@salud.madrid.org; roberto117343@gmail.com Supplementary informationSupplementary Table 1, 2, and Supplementary Text 1.

bioinformatics↗