Search bioRxivSearch

Biology subjects

Randhawa, G.

Publications and source records attributed to Randhawa, G..

2 recordsLinked to original sources

ML-DSP: Machine Learning with Digital Signal Processing for ultrafast, accurate, and scalable genome classification at all taxonomic levels

BackgroundAlthough methods and software tools abound for the comparison, analysis, identification, and taxonomic classification of the enormous amount of genomic sequences that are continuously being produced, taxonomic classification remains challenging. The difficulty lies within both the magnitude of the dataset and the intrinsic problems associated with classification. The need exists for an approach and software tool that addresses the limitations of existing alignment-based methods, as well as the challenges of recently proposed alignment-free methods.\n\nResultsWe combine supervised Machine Learning with Digital Signal Processing to design ML-DSP, an alignment-free software tool for ultrafast, accurate, and scalable genome classification at all taxonomic levels.\n\nWe test ML-DSP by classifying 7,396 full mitochondrial genomes from the kingdom to genus levels, with 98% classification accuracy. Compared with the alignment-based classification tool MEGA7 (with sequences aligned with either MUSCLE, or CLUSTALW), ML-DSP has similar accuracy scores while being significantly faster on two small benchmark datasets (2,250 to 67,600 times faster for 41 mammalian mitochondrial genomes). ML-DSP also successfully scales to accurately classify a large dataset of 4,322 complete vertebrate mtDNA genomes, a task which MEGA7 with MUSCLE or CLUSTALW did not complete after several hours, and had to be terminated. ML-DSP also outperforms the alignment-free tool FFP (Feature Frequency Profiles) in terms of both accuracy and time, being three times faster for the vertebrate mtDNA genomes dataset.\n\nConclusionsWe provide empirical evidence that ML-DSP distinguishes complete genome sequences at all taxonomic levels. Ultrafast and accurate taxonomic classification of genomic sequences is predicted to be highly relevant in the classification of newly discovered organisms, in distinguishing genomic signatures, in identifying mechanistic determinants of genomic signatures, and in evaluating genome integrity.

genomics

Identification of reference genes for real-time PCR gene expression studies during seed development and under abiotic stresses in Cyamopsis tetragonoloba (L.) Taub.

Guar (Cyamopsis tetragonoloba) is an important industrial crop. The knowledge about genes of guar involved in various processes can help in developing improved varieties of this crop. qRT-PCR is a preferred technique for accurate quantification of gene expression data. This technique requires the use of appropriate reference genes from the crop to be studied. Such genes have not been yet identified in guar. The expression stabilities of the 10 candidate reference genes, viz., CYP, ACT 11, EF-1, TUA, TUB, ACT 7, UBQ 10, UBC 2, GAPDH and 18S rRNA were evaluated in various tissues of guar under normal and abiotic stress conditions. Four different algorithms, geNorm, NormFinder, BestKeeper and {triangleup}Ct approach, were used to assess the expression stabilities and the results obtained were integrated into comprehensive stability rankings. The most stable reference genes were found to be CYP and ACT 11(tissues), ACT 11, UBC 2 and ACT 7 (seed development), ACT 7 and TUB (drought stress), TUA, UBC 2 and CYP (nitrogen stress), TUA and UBC 2 (cold stress), GAPDH and ACT 7 (heat stress) and GAPDH and EF-1a (salt stress). These results indicated the necessity of identifying a suitable reference gene for each experimental condition. Four selected reference genes were validated by normalizing the expression of CtMT1 gene. To the best of our knowledge this is the first report on the identification of reference genes in guar. These findings are likely to provide a boost to the gene expression studies in this important crop.

genetics