Search bioRxiv⌕ Search

Biology subjects

Kashyap, N.

Publications and source records attributed to Kashyap, N..

3 recordsLinked to original sources

NanoLabel: A fast and accurate real-time nanopore signal classifier

Oxford Nanopore Technologies adaptive sampling capability promises to reduce sequencing cost and turnaround time. At its core, adaptive sampling is a real-time classification problem that distinguishes reads originating from regions of interest. Direct signal-based classification approaches bypass the computational bottleneck of basecalling and can eliminate the need for powerful GPUs. However, operating directly on noisy raw signals remains challenging in real-time settings, where classification decisions must be made quickly. In this work, we propose NanoLabel, a new method for real-time classification of nanopore signals. We build NanoLabel on top of signal-based read mapping tool, RawHash2. We accelerate the classification workflow by mapping reads using only the target regions as the reference. To further improve accuracy, we train a lightweight classifier on mapping-derived features and introduce a data augmentation strategy to construct sufficiently large and class-balanced training datasets. We evaluate NanoLabel using publicly available real sequencing datasets from three human genomes (HG001, HG002, and HG005), while assuming a cancer gene panel as the target. Compared to directly mapping reads with RawHash2, we demonstrate 80 x improvement in the classification time and 0.10 - 0.25 units improvement in the F1 score.

genomics↗

EEG data quality in large scale field studies in India and Tanzania

There is a growing imperative to understand the neurophysiological impact of our rapidly changing and diverse technological, social, chemical, and physical environments. To untangle the multidimensional and interacting effects requires data at scale across diverse populations, taking measurement out of a controlled lab environment and into the field. Electroencephalography (EEG), which has correlates with various environmental factors as well as cognitive and mental health outcomes, has the advantage of both portability and cost-effectiveness for this purpose. However, with numerous field researchers spread across diverse locations, data quality issues and researcher idle time due to insufficient participants can quickly become unmanageable and expensive problems. In programs we have established in India and Tanzania, we demonstrate that with appropriate training, structured teams, and daily automated analysis and feedback on data quality, non-specialists can reliably collect EEG data alongside various survey and assessments with consistently high throughput and quality. Over a 30-week period, research teams were able to maintain an average of 25.6 subjects per week, collecting data from a diverse sample of 7,933 participants ranging from Hadzabe hunter-gatherers to office workers. Furthermore, data quality, computed on the first 2,400 records using two common methods, PREP and FASTER, was comparable to benchmark datasets from controlled lab conditions. Altogether this resulted in a cost per subject of under $50, a fraction of the cost typical of such data collection, opening up the possibility for large-scale programs particularly in low- and middle-income countries. Significance StatementWith wide human diversity, a rapidly changing environment and growing rates of neurological and mental health disorders, there is an imperative for large scale neuroimaging studies across diverse populations that can deliver high quality data and be affordably sustained. Here we demonstrate, across two large-scale field data acquisition programs operating in India and Tanzania, that with appropriate systems it is possible to generate high throughput EEG data of quality comparable to controlled lab settings. With effective costs of under $50 per subject, this opens new possibilities for low- and middle-income countries to implement large-scale programs, and to do so at scales that previously could not be considered.

neuroscience↗

Tracing India's Canine Heritage through SNP-Based Haplotype Identification

Dog breeds/germplasm in India is mostly unexplored and the population structure of owned dogs has not been studied at all. The current study was designed to determine the haplotypes and explore the population structure among divergent breeds of dogs using genome-wide distributed SNPs, followed by validation of selected haplotypes through PCR-sequencing. The research employed custom ddRAD-GBS sequencing carried out using Illumina 150bp paired-end sequencing of 50 dog samples generated 2,18,433 high-quality SNPs meeting the screening criteria. The data was further analyzed for population structure assessment and haplotype identification using bash, and R-environment. Subsequently, three haplotypes (on AFAP1, CELSR1, and GBGT1 genes) were selected (based on SNP density and haplotype length) for validation via PCR followed by paired-end Sanger sequencing in seven different dog breeds (n=21). The results revealed notable connections between dog breeds from Punjab and Haryana, while the affiliation with Karnataka was found to be less pronounced. The sequencing results indicated that CELSR1 and GBGT1 genes contained SNPs, with the AFAP1 gene lacking SNPs. The results also provided insights into the molecular-level population structure, SNPs, and haplotypes of diverse dog breeds reared in India. This SNP variation could be used for molecular characterization of indigenous dogs.

evolutionary biology↗