Search bioRxivSearch

Biology subjects

Williams, J. L.

Publications and source records attributed to Williams, J. L..

3 recordsLinked to original sources

FALCON-Phase: Integrating PacBio and Hi-C data for phased diploid genomes

Haplotype-resolved genome assemblies are important for understanding how combinations of variants impact phenotypes. These assemblies can be created in various ways, such as use of tissues that contain single-haplotype (haploid) genomes, or by co-sequencing of parental genomes, but these approaches can be impractical in many situations. We present FALCON-Phase, which integrates long-read sequencing data and ultra-long-range Hi-C chromatin interaction data of a diploid individual to create high-quality, phased diploid genome assemblies. The method was evaluated by application to three datasets, including human, cattle, and zebra finch, for which high-quality, fully haplotype resolved assemblies were available for benchmarking. Phasing algorithm accuracy was affected by heterozygosity of the individual sequenced, with higher accuracy for cattle and zebra finch (>97%) compared to human (82%). In addition, scaffolding with the same Hi-C chromatin contact data resulted in phased chromosome-scale scaffolds.

genomics

Complete assembly of parental haplotypes with trio binning

Reference genome projects have historically selected inbred individuals to minimize heterozygosity and simplify assembly. We challenge this dogma and present a new approach designed specifically for heterozygous genomes. \"Trio binning\" uses short reads from two parental genomes to partition long reads from an offspring into haplotype-specific sets prior to assembly. Each haplotype is then assembled independently, resulting in a complete diploid reconstruction. On a benchmark human trio, this method achieved high accuracy and recovered complex structural variants missed by alternative approaches. To demonstrate its effectiveness on a heterozygous genome, we sequenced an F1 cross between cattle subspecies Bos taurus taurus and Bos taurus indicus, and completely assembled both parental haplotypes with NG50 haplotig sizes >20 Mbp and 99.998% accuracy, surpassing the quality of current cattle reference genomes. We propose trio binning as a new best practice for diploid genome assembly that will enable new studies of haplotype variation and inheritance.

genomics

A Model for Genome-First Care: Returning Secondary Genomic Findings to Participants and Their Healthcare Providers in a Large Research Cohort

BackgroundResearch cohorts with linked genomic data exist, or are being developed, at many research centers. Within any such \"sequenced cohort\" of more than 100 participants, it is likely that there are participants with previously undisclosed risk for life-threatening monogenic diseases that could be identified with targeted analysis of their existing data. Identification of such disease-associated findings are not usually primary to the enrollment research goals. At Geisinger Health System, MyCode(R) Community Health Initiative (MyCode) participants represent one such large sequenced cohort. Since 2013, MyCode participants in discovery research have been consented for secondary analysis of their existing research genomic sequences to allow delivery of medically actionable findings to them and their healthcare providers. This return of genomic results program was developed to manage an anticipated 3.5% of MyCode participants who will receive clinically confirmed genomic variants from an approved gene list out of more than 150,000 total participants. Risk-associated DNA sequences alone without any clinical parameter, prompt \"genome-first\" follow-up encounters.\n\nMethodsThis article describes our process for generating clinical grade results from research-based genomic sequencing data, delivering results to patients and their providers, facilitating targeted clinical evaluations of patients and promoting cascade testing of at-risk relatives. We also summarize our early data about the results generated during this process and our ability to contact patients and their providers to disclose the information.\n\nResultsThis process has been used to generate 343 results on 339 patients. 93% of patients with a result have been successfully contacted about their results as evidenced by direct interaction about their result with the research team or a healthcare provider. 222 healthcare providers have been notified of a result on one or more patient through this result delivery process.\n\nConclusionsHere we describe the existing GHS model to deliver genomic data into the electronic medical record and the clinical interactions that are prompted and supported. Elements of this genome-first care model can be applied in other healthcare settings and in national efforts, such as \"All of Us\", that wish to establish programs for returning genomic results to research participants.

genomics