Search bioRxivSearch

Biology subjects

Kerstin Howe

Publications and source records attributed to Kerstin Howe.

5 recordsLinked to original sources

Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly

The human reference genome assembly plays a central role in nearly all aspects of todays basic and clinical research. GRCh38 is the first coordinate-changing assembly update since 2009 and reflects the resolution of roughly 1000 issues and encompasses modifications ranging from thousands of single base changes to megabase-scale path reorganizations, gap closures and localization of previously orphaned sequences. We developed a new approach to sequence generation for targeted base updates and used data from new genome mapping technologies and single haplotype resources to identify and resolve larger assembly issues. For the first time, the reference assembly contains sequence-based representations for the centromeres. We also expanded the number of alternate loci to create a reference that provides a more robust representation of human population variation. We demonstrate that the updates render the reference an improved annotation substrate, alter read alignments in unchanged regions and impact variant interpretation at clinically relevant loci. We additionally evaluated a collection of new de novo long-read haploid assemblies and conclude that while the new assemblies compare favorably to the reference with respect to continuity, error rate, and gene completeness, the reference still provides the best representation for complex genomic regions and coding sequences. We assert that the collected updates in GRCh38 make the newer assembly a more robust substrate for comprehensive analyses that will promote our understanding of human biology and advance our efforts to improve health.

Genomics

gEVAL - A web based browser for evaluating genome assemblies

MotivationFor most research approaches, genome analyses are dependent on the existence of a high quality genome reference assembly. However, the local accuracy of an assembly remains difficult to assess and improve. The gEVAL browser allows the user to interrogate an assembly in any region of the genome by comparing it to different datasets and evaluating the concordance. These analyses include: a wide variety of sequence alignments, comparative analyses of multiple genome assemblies, and consistency with optical and other physical maps. gEVAL highlights allelic variations, regions of low complexity, abnormal coverage, and potential sequence and assembly errors, and offers strategies for improvement. While gEVAL focuses primarily on sequence integrity, it can also display arbitrary annotation including Ensembl or TrackHub sources. We provide gEVAL web sites for many human, mouse, zebrafish and chicken assemblies to support the Genome Reference Consortium, and gEVAL is also downloadable to enable its use for any organism and assembly.\n\nAvailabilityWeb Browser: http://geval.sanger.ac.uk, Plugin: http://wchow.github.io/wtsi-geval-plugin.\n\nContactkj2@sanger.ac.uk\n\nSupplementary informationSupplementary data are available at Bioinformatics online.

Genomics

Structure and evolutionary history of a large family of NLR proteins in the zebrafish

Animals and plants have evolved a range of mechanisms for recognizing noxious substances and organisms. A particular challenge, most successfully met by the adaptive immune system in vertebrates, is the specific recognition of potential pathogens, which themselves evolve to escape recognition. A variety of genomic and evolutionary mechanisms shape large families of proteins dedicated to detecting pathogens and create the diversity of binding sites needed for epitope recognition. One family involved in innate immunity are the NACHT-domain-and Leucine-Rich-Repeat-containing (NLR) proteins. Mammals have a small number of NLR proteins, which are involved in first-line immune defense and recognize several conserved molecular patterns. However, there is no evidence that they cover a wider spectrum of differential pathogenic epitopes. In other species, mostly those without adaptive immune systems, NLRs have expanded into very large families. A family of nearly 400 NLR proteins is encoded in the zebrafish genome. They are subdivided into four groups defined by their NACHT and effector domains, with a characteristic overall structure that arose in fishes from a fusion of the NLR domains with a domain used for immune recognition, the B30.2 domain. The majority of the genes are located on one chromosome arm, interspersed with other large multi-gene families, including a new family encoding proteins with multiple tandem arrays of Zinc fingers. This chromosome arm may be a hot spot for evolutionary change in the zebrafish genome. NLR genes not on this chromosome tend to be located near chromosomal ends.\n\nExtensive duplication, loss of genes and domains, exon shuffling and gene conversion acting differentially on the NACHT and B30.2 domains have shaped the family. Its four groups, which are conserved across the fishes, are homogenised within each group by gene conversion, while the B30.2 domain is subject to gene conversion across the groups. Evidence of positive selection on diversifying mutations in the B30.2 domain, probably driven by pathogen interactions, indicates that this domain rather than the LRRs acts as a recognition domain. The NLR-B30.2 proteins represent a new family with diversity in the specific recognition module that is present in fishes in spite of the parallel existence of an adaptive immune system.

Genomics

Structure and evolutionary history of a large family of NLR proteins in the zebrafish

NACHT- and Leucine-Rich-Repeat-containing domain (NLR) proteins act as cytoplasmic sensors for pathogen- and danger-associated molecular patterns and are found throughout the plant and animal kingdoms. In addition to having a small set of conserved NLRs, the genomes in some animal lineages contain massive expansions of this gene family. One of these arose in fishes, after the creation of a gene fusion that combined the core NLR domains with another domain used for immune recognition, the PRY/SPRY or B30.2 domain. We have analysed the expanded NLR gene family in zebrafish, which contains 368 genes, and studied its evolutionary history. The encoded proteins share a defining overall structure, but individual domains show different evolutionary trajectories. Our results suggest gene conversion homogenizes NACHT and B30.2 domain sequences among different gene subfamilies, however, the functional implications of its action remains unclear. The majority of the genes are located on the long arm of chromosome 4, interspersed with several other large multi-gene families, including a new family encoding proteins with multiple tandem arrays of Zinc fingers. This suggests that chromosome 4 may be a hotspot for rapid evolutionary change in zebrafish.

Evolutionary Biology

The pig X and Y chromosomes: structure, sequence and evolution

We have generated an improved assembly and gene annotation of the pig X chromosome, and a first draft assembly of the pig Y chromosome, by sequencing BAC and fosmid clones, and incorporating information from optical mapping and fibre-FISH. The X chromosome carries 1,014 annotated genes, 689 of which are protein-coding. Gene order closely matches that found in Primates (including humans) and Carnivores (including cats and dogs), which is inferred to be ancestral. Nevertheless, several protein-coding genes present on the human X chromosome were absent from the pig (e.g. the cancer/testis antigen family) or inactive (e.g. AWAT1), and 38 pig-specific X-chromosomal genes were annotated, 22 of which were olfactory receptors. The pig Y chromosome assembly focussed on two clusters of male-specific low-copy number genes, separated by an ampliconic region including the HSFY gene family, which together make up most of the short arm. Both clusters contain palindromes with high sequence identity, presumably maintained by gene conversion. The long arm of the chromosome is almost entirely repetitive, containing previously characterised sequences. Many of the ancestral X-related genes previously reported in at least one mammalian Y chromosome are represented either as active genes or partial sequences. This sequencing project has allowed us to identify genes - both single copy and amplified - on the pig Y, to compare the pig X and Y chromosomes for homologous sequences, and thereby to reveal mechanisms underlying pig X and Y chromosome evolution.

Genomics