Search bioRxiv⌕ Search

Biology subjects

Hari, A.

Publications and source records attributed to Hari, A..

3 recordsLinked to original sources

Aldy 4: An efficient genotyper and star-allele caller for pharmacogenomics

High-throughput sequencing provides sufficient means for determining genotypes of clinically important pharmacogenes that can be used to tailor medical decisions to individual patients. However, pharmacogene genotyping, also known as star-allele calling, is a challenging problem that requires accurate copy number calling, structural variation discovery, variant calling and phasing within each pharmacogene copy present in the sample. Here we introduce Aldy 4, a fast and efficient tool for genotyping pharmacogenes that utilizes combinatorial optimization for accurate star-allele calling across different sequencing technologies. Aldy 4 adds support for long reads and ships with a novel phasing model and improved copy number and variant calling models. We compare Aldy 4 against the current state-of-the-art star-allele callers on a large and diverse set of samples and genes sequenced by various sequencing technologies, such as whole-genome and targeted Illumina sequencing, barcoded 10X Genomics and PacBio HiFi. We show that Aldy 4 is the most accurate star-allele caller with near-perfect accuracy in all evaluated contexts. We hope that Aldy remains an invaluable tool in the clinical toolbox even with the advent of long-read sequencing technologies. AvailabilityAldy 4 is available at https://github.com/0xTCG/aldy.

bioinformatics↗

mergem: merging and comparing genome-scale metabolic models using universal identifiers

Numerous methods exist to produce and refine genome-scale metabolic models. However, due to the use of incompatible identifier systems for metabolites and reactions, computing and visualizing the metabolic differences and similarities of such models is a current challenge. Furthermore, there is a lack of automated tools that can combine the strengths of multiple reconstruction pipelines into a curated single comprehensive model by merging different drafts, which possibly use incompatible namespaces. Here we present mergem, a novel method to compare and merge two or more metabolic models. Using a universal metabolic identifier mapping system constructed from multiple metabolic databases, mergem robustly can compare models from different pipelines and merge their common elements. mergem is implemented as a command line tool, a Python package, and on the web-application Fluxer, which allows simulating and visually comparing multiple models with different interactive flux graphs. The ability to merge and compare diverse genome scale metabolic models can facilitate the curation of comprehensive reconstructions and the discovery of unique and common metabolic features among different organisms.

systems biology↗

ImmunoTyper-SR: A Novel Computational Approach for Genotyping Immunoglobulin Heavy Chain Variable Genes using Short Read Data

Human immunoglobulin heavy chain (IGH) locus on chromosome 14 includes more than 40 functional copies of the variable gene (IGHV), which, together with the joining genes (IGHJ), diversity genes (IGHD), constant genes (IGHC) and immunoglobulin light chains, code for antibodies that identify and neutralize pathogenic invaders as a part of the adaptive immune system. Because of its highly repetitive sequence composition, the IGH locus has been particularly difficult to assemble or genotype through the use of standard short read sequencing technologies. Here we introduce ImmunoTyper-SR, an algorithmic method for genotype and CNV analysis of the germline IGHV genes using Illumina whole genome sequencing (WGS) data. ImmunoTyper-SR is based on a novel combinatorial optimization formulation that aims to minimize the total edit distance between reads and their assigned IGHV alleles from a given database, with constraints on the number and distribution of reads across each called allele. We have validated ImmunoTyper-SR on 12 individuals with Illumina WGS data from the 1000 Genomes Project, whose IGHV allele composition have been studied extensively through the use of long read and targeted sequencing platforms, as well as nine individuals from the NIAID COVID Consortium who have been subjected to WGS twice. We have then applied ImmunoTyper-SR on 585 samples from the NIAID COVID Consortium to investigate associations between distinct IGHV alleles and anti-type I IFN autoantibodies which have been linked to COVID-19 severity.

bioinformatics↗