Search bioRxiv⌕ Search

Biology subjects

Vo, N. S.

Publications and source records attributed to Vo, N. S..

6 recordsLinked to original sources

VN1K: a genome graph-based and function-driven multi-omics and phenomics resource for the Vietnamese population

Vietnam, the 16th most populated nation, remains profoundly underrepresented in global genomic databases. Here, we present VN1K, a first-ever comprehensive and well-curated resource of multi-omics data with a wide-range of phenotypic information of 1,011 unrelated Vietnamese individuals. High-depth short-read whole-genome sequencing data were generated for all samples along with various - omic data, including microarray, long-read whole-genome sequencing, and RNA sequencing. Using a high-sensitivity variant detection pipeline, which included a pangenome graph reference and a deep-learning framework, we identified nearly 40 million variants of which 8.5 million are novel with nearly 900 thousand short insertions/deletions and 39 thousand structural variants. Specifically, VN1K featured a first-ever whole-genome methylation profile based on long read sequencing. A genotype imputation panel was also created with the highest accuracy on the Vietnamese population. Variants with significantly different allele frequencies in the Vietnamese population compared to others were found to be functionally significant, especially in genes associated with immune diseases (HLA-B, KIR3DL3, KIR2DL1, KIR2DL4) or drug responses (CYP2C19, CYP2D6, VKORC1, CYP2B6). We were also able to map various loci related to hepatitis B virus infection as well as six disease traits, including triglyceride levels, LDL-C, serum glucose levels, HbA1c, and levels of two liver enzymes (ALT and AST). VN1K dataset is accessible via genome.vinbigdata.org, an integrated platform with both linear and graph-based genome browser for facilitating data exploration, research, and applications in precision medicine.

genomics↗

AMRomics: a scalable workflow to analyze large microbial genome collection

Whole genome analysis for microbial genomics is critical to studying and monitoring antimicrobial resistance strains. The exponential growth of microbial sequencing data necessitates a fast and scalable computational pipeline to generate the desired outputs in a timely and cost-effective manner. Recent methods have been implemented to integrate individual genomes into large collections of specific bacterial populations and are widely employed for systematic genomic surveillance. However, they do not scale well when the population expands and turnaround time remains the main issue for this type of analysis. Here, we introduce AMRomics, a minimalized microbial genomics pipeline that can work efficiently with big datasets. We use different bacterial data collections to compare AMRomics against competitive tools and show that our pipeline can generate similar results of interest but with better performance. The software is open source and is publicly available at https://github.com/amromics/amromics under an MIT license.

bioinformatics↗

A study of genetic variants associated with skin traits in the Vietnamese population

BackgroundMost skin-related traits have been studied from Caucasian genetic background. A comprehensive study on skin-associated genetic effects on under-represented populations like Vietnam is needed to fill the gaps in the field. ObjectivesTo develop a computational pipeline to predict the effect of genetic factors on skin traits using public data (GWAS catalogs and whole genome sequencing (WGS) data of 1000 genomes project-1KGP) and in-house Vietnamese data (WGS and genotyping by SNP array). By using this information we may have a better understanding of the susceptibility of Vietnamese people. MethodsVietnamese cohorts of whole genome sequencing (WGS) of 1008 healthy individuals for the reference and 96 genotyping samples (which do not have any skin cutaneous issues) by Infinium Asian Screening Array-24 v1.0 BeadChip were employed to predict skin-associated genetic variants of 25 skin-related and micronutrients requirement traits in population analysis and correlation analysis. Simultaneously, we compared the landscape of cutaneous issues of Vietnamese people with other populations by assessing their genetic profiles. ResultsThe skin-related genetic profile of Vietnamese cohorts is similar at most with East Asian (JPT: Fst=0.036, CHB: Fst=0.031, CHS: Fst=0.027, CDX: Fst=0.025) in the population study. In addition, we identified pairs of skin traits being at high risk of frequent co-occurrence (such as skin aging and wrinkles (r = 0.45, p =1.50e-5) or collagen degradation and moisturizing (r = 0.35, p = 1.1e-3). ConclusionThis is the first investigation in Vietnam to explore genetic variants of facial skin. These findings could improve inadequate skin-related genetic diversity in the currently published database.

bioinformatics↗

PanTA: An ultra-fast method for constructing large and growing microbial pangenomes

Pangenome analysis is an indispensable step in bacterial genomics to address the high variability of bacteria genomes. However, speed and scalability remain a challenge for pangenome inference software tools to cope with the fast-growing genomic collections. We present PanTA, a software package for constructing the pangenomes of large bacterial collections. We show that PanTA exhibits an unprecedented multiple times more efficient than the current state-of-the-arts while maintaining a similar pangenome accuracy. In addition, PanTA introduces a novel mechanism to construct the pangenome progressively where new samples are added into an existing pangenome without rebuilding the accumulated collection from scratch. In the progressive mode, PanTA is demonstrated to consume orders of magnitude less computational resource than existing solutions in managing the pangenomes of growing microbial datasets. We further show that PanTA can build the pangenome of the entire collection of >28000 Escherichia coli genomes from the RefSeq database on a laptop computer in 32 hours, highlighting the scalability and practicality of PanTA.The software is open source and is publicly available at https://github.com/amromics/panta under an MIT license.

bioinformatics↗

Imputation and polygenic score performances of human genotyping arrays in diverse populations

Regardless of the overwhelming use of next-generation sequencing technologies, microarray-based genotyping combined with the imputation of untyped variants remains a cost-effective means to interrogate genetic variations across the human genome. This technology is widely used in genome-wide association studies (GWAS) at bio-bank scales, and more recently, in polygenic score (PGS) analysis to predict and to stratify disease risk. Over the last decade, human genotyping arrays have undergone a tremendous growth in both number, and content making a comprehensive evaluation of their performances became more important. Here, we performed a comprehensive performance assessment for 23 available human genotyping arrays in 6 ancestry groups using diverse public, and in-house datasets. The analyses focus on performance estimation of derived imputation (in terms of accuracy and coverage) and PGS (in term of concordance to PGS estimated from whole genome sequencing data) in three different traits and diseases. We found that the arrays with a higher number of SNPs are not necessarily the ones with higher imputation performance, but the arrays that are well-optimized for the targeted population could provide very good imputation performance. In addition, PGS estimated by imputed SNP array data is highly correlated to PGS estimated by whole genome sequencing data in most of cases. When optimal arrays are used, the correlations of key PGS metrics between two types of data can be higher than 0.97, but interestingly, arrays with high density can result in lower PGS performance. Our results suggest the importance of properly selecting a suitable genotyping array for PGS applications. Finally, we developed a web tool that provide interactive analyses of tag SNP contents and imputation performance based on population and genomic regions of interest. This study would act as a practical guide for researchers to design their genotyping arrays-based studies. The tool is available at: https://genome.vinbigdata.org/tools/saa/

genomics↗

LmTag: functional-enrichment and imputation-aware tag SNP selection for population-specific genotyping arrays

Despite the rapid development of sequencing technology, single-nucleotide polymorphism (SNP) array is still the most cost-effective genotyping solutions for large-scale genomic research and applications. Recent years have witnessed the rapid development of numerous genotyping platforms of different sizes and designs, but population-specific platforms are still lacking, especially for those in developing countries. We aim to develop methods to design SNP arrays for thse countries, so the arrays should be cost-effective (small size), yet can still generate key information needed to associate genotypes with traits. A key design principle for most current platforms is to improve genome-wide imputation so that more SNPs (imputed tag SNPs) not included in the array can be predicted. However, current tag SNP selection methods mostly focus on imputation accuracy and coverage, but not the functional content of the measured and imputed SNPs. It is those functional SNPs that are most likely associated to traits. Here, we propose LmTag, a novel method for tag SNP selection that not only improves imputation performance but also prioritizes highly functional SNP markers. We apply LmTag on a wide range of populations using both public and in-house whole genome sequencing databases. Our results showed that LmTag improved both functional marker prioritization and genome-wide imputation accuracy compared to existing methods. This novel approach could contribute to the next generation genotyping arrays that provide excellent imputation capability as well as facilitate array-based functional genetic studies. Such arrays are particularly suitable for under-represented populations in developing countries or non-model species, where little genomics data are available while investment in genome sequencing or high-density SNP arrays is limited.

bioinformatics↗