Search bioRxivSearch

Biology subjects

Wilm, A.

Publications and source records attributed to Wilm, A..

3 recordsLinked to original sources

Large-scale whole-genome sequencing of three diverse Asian populations in Singapore

Asian populations are currently underrepresented in human genetics research. Here we present whole-genome sequencing data of 4,810 Singaporeans from three diverse ethnic groups: 2,780 Chinese, 903 Malays, and 1,127 Indians. Despite a medium depth of 13.7x, we achieved essentially perfect (>99.8%) sensitivity and accuracy for detecting common variants and good sensitivity (>89%) for detecting extremely rare variants with <0.1% allele frequency. We found 89.2 million single-nucleotide polymorphisms (SNPs) and 9.1 million small insertions and deletions (INDELs), more than half of which have not been cataloged in dbSNP. In particular, we found 126 common deleterious mutations (MAF>0.01) that were absent in the existing public databases, highlighting the importance of local population reference for genetic diagnosis. We describe fine-scale genetic structure of Singapore populations and their relationship to worldwide populations from the 1000 Genomes Project. In addition to revealing noticeable amounts of admixture among three Singapore populations and a Malay-related novel ancestry component that has not been captured by the 1000 Genomes Project, our analysis also identified some fine-scale features of genetic structure consistent with two waves of prehistoric migration from south China to Southeast Asia. Finally, we demonstrate that our data can substantially improve genotype imputation not only for Singapore populations, but also for populations across Asia and Oceania. These results highlight the genetic diversity in Singapore and the potential impacts of our data as a resource to empower human genetics discovery in a broad geographic region.

genetics

Structure mapping of dengue and Zika viruses reveals new functional long-range interactions

Dengue and Zika are clinically important members of the Flaviviridae family that utilizes an 11kb positive strand RNA for genome regulation. While structures have been mapped primarily in the UTRs, much remains to be learnt about how the rest of the genome folds to enable function. Here, we performed secondary structure and pair-wise interaction mapping on four dengue serotypes and four Zika strains in their native virus particles and infected cells. Comparative analysis of SHAPE reactivities across serotypes nominated potentially functional regions that are highly structured, show structure conservation, and low synonymous mutation rates, including a structure associated with ribosome pausing. Pair-wise interaction mapping by SPLASH further reveals new pair-wise interactions, in addition to the known circularization sequence. 40% of pair-wise interactions form alternative structures, suggesting extensive structural heterogeneity. Analysis of shared pair-wise interactions between serotypes revealed macro-organization whereby interactions are preserved at their physical locations, beyond their sequence identities. In addition, structure mapping of virus genomes released in solution-as well as inside host cells-showed that other helicases, in addition to the ribosome, play a role in unwinding viral structures inside cells. Mutational experiments that disrupt in cell and in virion pair-wise interactions result in virus attenuation, demonstrating their importance during the virus life-cycle.

genomics

Single-Virion Sequencing Of Lamivudine Treated HBV Populations Reveal Population Evolution Dynamics And Demographic History

Viral populations are complex, dynamic, and fast evolving. The evolution of groups of closely related viruses in a competitive environment is termed quasispecies. To fully understand the role that quasispecies play in viral evolution, characterizing the trajectories of viral genotypes in an evolving population is the key. In particular, long-range haplotype information for thousands of individual viruses is critical; yet generating this information is non-trivial. Popular deep sequencing methods generate relatively short reads that do not preserve linkage information, while third generation sequencing methods have higher error rates that make detection of low frequency mutations a bioinformatics challenge. Here we applied BAsE-Seq, an Illumina-based single-virion sequencing technology, to eight samples from four chronic hepatitis B (CHB) patients - once before antiviral treatment and once after viral rebound due to resistance. We obtained 248-8,796 single-virion sequences per sample, which allowed us to find evidence for both hard and soft selective sweeps. We were also able to reconstruct population demographic history that was independently verified by clinically collected data. We further verified four of the samples independently on PacBio and Illumina sequencers. Overall, we showed that single-virion sequencing yields insight into viral evolution and population dynamics in an efficient and high throughput manner. We believe that single-virion sequencing is widely applicable to the study of viral evolution in the context of drug resistance, differentiating between soft or hard selective sweeps, and the reconstruction of intra-host viral population demographic history.

evolutionary biology