Search bioRxivSearch

Biology subjects

Foo, J. N.

Publications and source records attributed to Foo, J. N..

2 recordsLinked to original sources

Large-scale whole-genome sequencing of three diverse Asian populations in Singapore

Asian populations are currently underrepresented in human genetics research. Here we present whole-genome sequencing data of 4,810 Singaporeans from three diverse ethnic groups: 2,780 Chinese, 903 Malays, and 1,127 Indians. Despite a medium depth of 13.7x, we achieved essentially perfect (>99.8%) sensitivity and accuracy for detecting common variants and good sensitivity (>89%) for detecting extremely rare variants with <0.1% allele frequency. We found 89.2 million single-nucleotide polymorphisms (SNPs) and 9.1 million small insertions and deletions (INDELs), more than half of which have not been cataloged in dbSNP. In particular, we found 126 common deleterious mutations (MAF>0.01) that were absent in the existing public databases, highlighting the importance of local population reference for genetic diagnosis. We describe fine-scale genetic structure of Singapore populations and their relationship to worldwide populations from the 1000 Genomes Project. In addition to revealing noticeable amounts of admixture among three Singapore populations and a Malay-related novel ancestry component that has not been captured by the 1000 Genomes Project, our analysis also identified some fine-scale features of genetic structure consistent with two waves of prehistoric migration from south China to Southeast Asia. Finally, we demonstrate that our data can substantially improve genotype imputation not only for Singapore populations, but also for populations across Asia and Oceania. These results highlight the genetic diversity in Singapore and the potential impacts of our data as a resource to empower human genetics discovery in a broad geographic region.

genetics

Array-based sequencing of filaggrin gene for comprehensive detection of disease-associated variants

The filaggrin gene (FLG) is essential for skin differentiation and epidermal barrier formation with links to skin diseases, however it has a highly repetitive nucleotide sequence containing very limited stretches of unique nucleotides for precise mapping to reference genomes. Sequencing strategies using polymerase chain reaction (PCR) and conventional Sanger sequencing have been successful for complete FLG coding DNA sequence amplification to identify pathogenic mutations but this time-consuming, labour intensive method has restricted utility. Next-generation sequencing (NGS) offers obvious benefits to accelerate FLG analysis but standard re-sequencing techniques such as oligoprobe-based exome or customized targeted-capture can be expensive, especially for a single target gene of interest. We therefore designed a protocol to improve FLG sequencing throughput using a set of FLG-specific PCR primer assays compatible with microfluidic amplification, multiplexing and current NGS protocols. Using DNA reference samples with known FLG genotypes for benchmarking, this protocol is shown to be concordant for variant detection across different sequencing methodologies. We applied this methodology to analyze cohorts from ethnicities previously not studied for FLG variants and demonstrate usefulness for discovery projects. This comprehensive coverage sequencing protocol is labour-efficient and offers an affordable solution to scale up FLG sequencing for larger cohorts. Robust and rapid FLG sequencing can improve patient stratification for research projects and provide a framework for gene specific diagnosis in the future.

genetics