Search bioRxiv⌕ Search

Biology subjects

Lai, H.-S.

Publications and source records attributed to Lai, H.-S..

3 recordsLinked to original sources

deCYPher: Star Allele-Resolution Computational Framework of Pharmacogenes for Haplotype-Resolved Long-Read Assemblies

Although existing next-generation sequencing (NGS) tools, such as Aldy and Cyrius, have been applied for allele typing, they cannot achieve complete accuracy due to various genomic challenges including pseudogenes, structural variations, hybrid genes, copy number variations, and gene deletions. These complexities make accurate pharmacogene interpretation more challenging, despite the crucial role pharmacogenomics plays in precision medicine. We developed deCYPher, a tool that generates personalized pharmacogenomic reports from haplotype-resolved assemblies. The tool enables analysis of all PharmVar 1A level genes, such as CYP2B6, CYP2C9, CYP2C19, CYP2D6, CYP3A5, CYP4F2, DPYD, NUDT15, and SLCO1B1. Applied to all HPRC haplotypes (including both release 1 and release 2 data), deCYPher demonstrated high accuracy in resolving complex gene structures. In the case of CYP2D6, release 1 identified 6% gene multiplications, 6% full gene deletions, and 4% CYP2D6/CYP2D7 hybrids. By contrast, release 2 demonstrated an increased prevalence of multiplications (14%) and hybrids (11%), while the frequency of full gene deletions remained comparable at 5%. Comparison with pb-StarPhase revealed discrepancies in 12 of 94 assemblies in the release 1 dataset. For instance, in sample HG02257, Aldy, Cyrius, and deCYPher consistently identified the genotype as *2/*35, whereas pb-StarPhase reported *2/*2. Notably, the *35-defining variants were present in the BAM and VCF files in the pb-StarPhase pipeline, but the local read depth over the *35-specific region was only 5x in HG02257-p, suggesting that the misclassification likely resulted from insufficient coverage - a known limitation of pb-StarPhase under low-depth conditions.

bioinformatics↗

TypeAssembly: Copy number estimation and allele typing for haplotype assemblies

Accurately annotating complex genes in the human genome, particularly from haplotype assemblies, remains a significant challenge. To overcome this, we developed TypeAssembly, a local alignment-based framework for copy number estimation and allele typing. Operating in two modes, mode-FASTA and mode-VCF, TypeAssembly can define alleles by either sequence or variant information. We successfully applied it to annotate 41 genes in the MHC locus, 17 KIR genes, and, for the first time, 15 pharmacogenes across 466 haplotype assemblies. This study establishes TypeAssembly as a robust method for accurately annotating complex genomic regions and provides an evaluation of existing gene annotations and callers.

bioinformatics↗

Evaluating the performance of protein structure prediction in detecting structural changes of pathogenic nonsynonymous single nucleotide variants

AO_SCPLOWBSTRACTC_SCPLOWProtein structure prediction serves as an efficient tool, saving time and circumventing the need for laborious experimental endeavors. Distinguished methodologies, including AlphaFold, RoseTTAFold, and ESMFold, have proven their precision through rigorous evaluation based on the last Critical Assessment of Protein Structure Prediction (CASP14). The success of protein structure prediction raises the following question: can the prediction tools discern structural alterations resulting from single amino acid changes? In this regard, the objective of this study is to assess the performance of existing structure prediction tools on mutated sequences. In this study, we posited that a specific fraction of the pathogenic nonsynonymous single nucleotide variants (nsSNVs) would experience structural alterations following amino acid mutations. We meticulously assembled an extensive dataset by initially sourcing data from ClinVar and subsequently applying multiple filters, resulting in 964 alternative sequences and their corresponding reference sequences. Utilizing UniProt, we acquired reference sequences and generated the corresponding alternative sequences based on variant information. This study performed three tools of structure prediction on both the reference and alternative sequences and expected some structural changes upon mutations. Our findings affirm AlphaFold as the foremost prediction tool presently; nonetheless, our experimental results underscore persistent challenges in accurately predicting structural alterations induced by nsSNVs. Discrepancies between the predicted structures of reference and alternative sequences, when observed, often stem from a lack of confidence in the predictions or the spatial separation between compact domains interrupted by disordered regions, posing challenges to successful alignment.

bioinformatics↗