Search bioRxivSearch

Biology subjects

Lucia Carbone

Publications and source records attributed to Lucia Carbone.

3 recordsLinked to original sources

Construction of thousands of single cell genome sequencing libraries using combinatorial indexing

Single cell genome sequencing has proven to be a valuable tool for the detection of somatic variation, particularly in the context of tumor evolution and neuronal heterogeneity. Current technologies suffer from high per-cell library construction costs which restrict the number of cells that can be assessed, thus imposing limitations on the ability to quantitatively measure genomic heterogeneity within a tissue. Here, we present Single cell Combinatorial Indexed Sequencing (SCI-seq) as a means of simultaneously generating thousands of low-pass single cell libraries for the purpose of somatic copy number variant detection. In total, we constructed libraries for 16,698 single cells from a combination of cultured cell lines, frontal cortex tissue from Macaca mulatta, and two human adenocarcinomas. This novel technology provides the opportunity for low-cost, deep characterization of somatic copy number variation in single cells, providing a foundational knowledge across both healthy and diseased tissues.

Genomics

Whole-genome characterization in pedigreed non-human primates using Genotyping-By-Sequencing and imputation.

BackgroundRhesus macaques are widely used in biomedical research, but the application of genomic information in this species to better understand human disease is still undeveloped. Whole-genome sequence (WGS) data in pedigreed macaque colonies could provide substantial experimental power, but the collection of WGS data in large cohorts remains a formidable expense. Here, we describe a cost-effective approach that selects the most informative macaques in a pedigree for whole-genome sequencing, and imputes these dense marker data into all remaining individuals having sparse marker data, obtained using Genotyping-By-Sequencing (GBS).\n\nResultsWe developed GBS for the macaque genome using a single digest with PstI, followed by sequencing to 30X coverage. From GBS sequence data collected on all individuals in a 16-member pedigree, we characterized an optimal 22,455 sparse markers spaced ~125 kb apart. To characterize dense markers for imputation, we performed WGS at 30X coverage on 9 of the 16 individuals, yielding ~10.2 million high-confidence variants. Using the approach of \"Genotype Imputation Given Inheritance\" (GIGI), we imputed alleles at an optimized dense set of 4,920 variants on chromosome 19, using 490 sparse markers from GBS. We assessed changes in accuracy of imputed alleles, 1) across 3 different strategies for selecting individuals for WGS, i.e., a) using \"GIGI-Pick\" to select informative individuals, b) sequencing the most recent generation, or c) sequencing founders only; and 2) when using from 1-9 WGS individuals for imputation. We found that accuracy of imputed alleles was highest using the GIGI-Pick selection strategy (median 92%), and improved very little when using >4 individuals with WGS for imputation. We used this ratio of 4 WGS to 12 GBS individuals to impute an expanded set of ~14.4 million variants across all 20 macaque autosomes, achieving ~85-88% accuracy per chromosome.\n\nConclusionsWe conclude that an optimal tradeoff exists at the ratio of 1 individual selected for WGS using the GIGI-Pick algorithm, per 3-5 relatives selected for GBS, a cost savings of ~67-83% over WGS of all individuals. This approach makes feasible the collection of accurate, dense genome-wide sequence data in large pedigreed macaque cohorts without the need for expensive WGS data on all individuals.

Genomics

An Approximate Bayesian Computation Approach to Examining the Phylogenetic Relationships among the Four Gibbon Genera using Whole Genome Sequence Data

Gibbons are believed to have diverged from the larger great apes [~]16.8 Mya and today reside in the rainforests of Southeast Asia. Based on their diploid chromosome number, the family Hylobatidae is divided into four genera, Nomascus, Symphalangus, Hoolock and Hylobates. Genetic studies attempting to elucidate the phylogenetic relationships among gibbons using karyotypes, mtDNA, the Y chromosome, and short autosomal sequences have been inconclusive. To examine the relationships among gibbon genera in more depth, we performed 2nd generation whole genome sequencing to a mean of [~]15X coverage in two individuals from each genus. We developed a coalescent-based Approximate Bayesian Computation method incorporating a model of sequencing error generated by high coverage exome validation to infer the branching order, divergence times, and effective population sizes of gibbon taxa. Although Hoolock and Symphalangus are likely sister taxa, we could not confidently resolve a single bifurcating tree despite the large amount of data analyzed. Our combined results support the hypothesis that all four gibbon genera diverged at approximately the same time. Assuming an autosomal mutation rate of 1x10-9/site/year this speciation process occurred [~]5 Mya during a period in the Early Pliocene characterized by climatic shifts and fragmentation of the Sunda shelf forests. Whole genome sequencing of additional individuals will be vital for inferring the extent of gene flow among species after the separation of the gibbon genera.

Genomics