Search bioRxivSearch

Biology subjects

Gonen, S.

Publications and source records attributed to Gonen, S..

3 recordsLinked to original sources

A heuristic method for fast and accurate phasing and imputation of single nucleotide polymorphism data in bi-parental plant populations

This paper presents a new heuristic method for phasing and imputation of genomic data in diploid plant species. Our method, called AlphaPlantImpute, explicitly leverages features of plant breeding programs to maximise the accuracy of imputation. The features are a small number of parents, which can be inbred and usually have high-density genomic data, and few recombinations separating parents and focal individuals genotyped at low-density (i.e. descendants that are the imputation targets). AlphaPlantImpute works roughly in three steps. First, it identifies informative low-density genotype markers in parents. Second, it tracks the inheritance of parental alleles and haplotypes to focal individuals at informative markers. Finally, it uses this low-density information as anchor points to impute focal individuals to high-density.\n\nWe tested the imputation accuracy of AlphaPlantImpute in simulated bi-parental populations across different scenarios. We also compared its accuracy to existing software called PlantImpute. In general, AlphaPlantImpute had better or equal imputation accuracy as PlantImpute. The computational time and memory requirements of AlphaPlantImpute were tiny compared to PlantImpute. For example, accuracy of imputation was 0.96 for a scenario where both parents were inbred and genotyped at 25,000 markers per chromosome and a focal F2 individual was genotyped with 50 markers per chromosome. The maximum memory requirement for this scenario was 0.08 GB and took 37 seconds to complete.

genomics

Near-Atomic Cryo-EM Imaging of a Small Protein Displayed on a Designed Scaffolding System

Current single particle electron cryo-microscopy (cryo-EM) techniques can produce images of large protein assemblies and macromolecular complexes at atomic level detail without the need for crystal growth. However, proteins of smaller size, typical of those found throughout the cell, are not presently amenable to detailed structural elucidation by cryo-EM. Here we use protein design to create a modular, symmetrical scaffolding system to make protein molecules of typical size amenable to cryo-EM. Using a rigid continuous alpha-helical linker, we connect a small 17 kDa protein (DARPin) to a protein subunit that was designed to self-assemble into a cage with cubic symmetry. We show that the resulting construct is amenable to structural analysis by single particle cryo-EM, allowing us to identify and solve the structure of the attached small protein at near-atomic detail, ranging from 3.5 to 5 [A] resolution. The result demonstrates that proteins considerably smaller than the theoretical limit of 50 kDa for cryo-EM can be visualized clearly when arrayed in a rigid fashion on a symmetric designed protein scaffold. Furthermore, because the amino acid sequence of a DARPin can be chosen to confer tight binding to various other protein or nucleic acid molecules, the system provides a future route for imaging diverse macromolecules, potentially broadening the application of cryoEM to proteins of typical size in the cell.\n\nSignificance statementNew electron microscopy methods are making it possible to view the structures of large proteins and nucleic acid complexes at atomic detail, but the methods are difficult to apply to molecules smaller than about 50 kDa, which is larger than the size of the average protein in the cell. The present work demonstrates that a protein much smaller than that limit can be successfully visualized when it is attached to a large protein scaffold designed to hold 12 copies of the attached protein in symmetric and rigidly defined orientations. The small protein chosen for attachment and visualization can be modified to bind to other diverse proteins, opening up a new avenue for imaging cellular proteins by cryo-EM.

molecular biology

A method for allocating low-coverage sequencing resources by targeting haplotypes rather than individuals

BackgroundThis paper describes a heuristic method for allocating low-coverage sequencing resources by targeting haplotypes rather than individuals. Low-coverage sequencing assembles high-coverage sequence information for every individual by accumulating data from the genome segments that they share with many other individuals into consensus haplotypes. Deriving the consensus haplotypes accurately is critical for achieving a high phasing and imputation accuracy. In order to enable accurate phasing and imputation of sequence information for the whole population we allocate the available sequencing resources among individuals with existing phased genomic data by targeting the sequencing coverage of their haplotypes.\n\nResultsOur method, called AlphaSeqOpt, prioritizes haplotypes using a score function that is based on the frequency of the haplotypes in the sequencing set relative to the target coverage. AlphaSeqOpt has two steps: (1) selection of an initial set of individuals by iteratively choosing the individuals that have the maximum score conditional to the current set, and (2) refinement of the set through several rounds of exchanges of individuals. AlphaSeqOpt is very effective for distributing a fixed amount of sequencing resources evenly across haplotypes, which results in a reduction of the proportion of haplotypes that are sequenced below the target coverage. AlphaSeqOpt can provide a greater proportion of haplotypes sequenced at the target coverage by sequencing less individuals, as compared with other methods that use a score function based on the haplotypes population frequency. A refinement of the initially selected set can provide a larger more diverse set with more unique individuals, which is beneficial in the context of low-coverage sequencing. We extend the method with an approach to filter rare haplotypes based on their flanking haplotypes, so that only those that are likely to derive from a recombination event are targeted.\n\nConclusionsWe present a method for allocating sequencing resources so that a greater proportion of haplotypes are sequenced at a coverage that is sufficiently high for population-based imputation with low-coverage sequencing. The haplotype score function, the refinement step, and the new approach of filtering rare haplotypes make AlphaSeqOpt more effective for that purpose than methods reported previously for reducing sequencing redundancy.

genomics