Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 919 records · Page 51Linked to original sources

Population genetic history and polygenic risk biases in 1000 Genomes populations

The vast majority of genome-wide association studies are performed in Europeans, and their transferability to other populations is dependent on many factors (e.g. linkage disequilibrium, allele frequencies, genetic architecture). As medical genomics studies become increasingly large and diverse, gaining insights into population history and consequently the transferability of disease risk measurement is critical. Here, we disentangle recent population history in the widely-used 1000 Genomes Project reference panel, with an emphasis on populations underrepresented in medical studies. To examine the transferability of single-ancestry GWAS, we used published summary statistics to calculate polygenic risk scores for six well-studied traits and diseases. We identified directional inconsistencies in all scores; for example, height is predicted to decrease with genetic distance from Europeans, despite robust anthropological evidence that West Africans are as tall as Europeans on average. To gain deeper quantitative insights into GWAS transferability, we developed a complex trait coalescent-based simulation framework considering effects of polygenicity, causal allele frequency divergence, and heritability. As expected, correlations between true and inferred risk were typically highest in the population from which summary statistics were derived. We demonstrated that scores inferred from European GWAS were biased by genetic drift in other populations even when choosing the same causal variants, and that biases in any direction were possible and unpredictable. This work cautions that summarizing findings from large-scale GWAS may have limited portability to other populations using standard approaches, and highlights the need for generalized risk prediction methods and the inclusion of more diverse individuals in medical genomics.

Genomics

Micro-C XL: assaying chromosome conformation at length scales from the nucleosome to the entire genome

Structural analysis of chromosome folding in vivo has been revolutionized by Chromosome Conformation Capture (3C) and related methods, which use proximity ligation to identify chromosomal loci in physical contact. We recently described a variant 3C technique, Micro-C, in which chromatin is fragmented to mononucleosomes using micrococcal nuclease, enabling nucleosome-resolution folding maps of the genome. Here, we describe an improved Micro-C protocol using long crosslinkers, termed Micro-C XL, which exhibits greatly increased signal to noise, and provides further insight into the folding of the yeast genome. We also find that signal to noise is much improved in Micro-C XL libraries generated from relatively insoluble chromatin as opposed to soluble material, providing a simple method to physically enrich for bona-fide long-range interactions. Micro-C XL maps of the budding and fission yeast genomes reveal both short-range chromosome fiber features such as chromosomally-interacting domains (CIDs), as well as higher-order features such as clustering of centromeres and telomeres, thereby addressing the primary discrepancy between prior Micro-C data and reported 3C and Hi-C analyses. Interestingly, comparison of chromosome folding maps of S. cerevisiae and S. pombe revealed widespread qualitative similarities, yet quantitative differences, between these distantly-related species. Micro-C XL thus provides a single assay suitable for interrogation of chromosome folding at length scales from the nucleosome to the full genome.

Genomics

Error rates, PCR recombination, and sampling depth in HIV-1 Whole Genome Deep Sequencing

Deep sequencing is a powerful and cost-effective tool to characterize the genetic diversity and evolution of virus populations. While modern sequencing instruments readily cover viral genomes many thousand fold and very rare variants can in principle be detected, sequencing errors, amplification biases, and other artifacts can limit sensitivity and complicate data interpretation. Here, we describe several control experiments and error correction methods for whole-genome deep sequencing of viral genomes. We developed many of these in the course of a large scale whole genome deep sequencing study of HIV-1 populations. We measured the substitution and indel errors that arose during sequencing and PCR and quantified PCR-mediated recombination. We find that depending on the viral load in the samples, rare mutations down to 0.2% can be reproducibly detected. PCR recombination can be avoided by consistently working at low amplicon concentrations.

Genomics

Long-read whole genome sequencing identifies causal structural variation in a Mendelian disease

Current clinical genomics assays primarily utilize short-read sequencing (SRS), which offers high throughput, high base accuracy, and low cost per base. SRS has, however, limited ability to evaluate tandem repeats, regions with high [GC] or [AT] content, highly polymorphic regions, highly paralogous regions, and large-scale structural variants. Long-read sequencing (LRS) has complementary strengths and offers a means to discover overlooked genetic variation in patients undiagnosed by SRS. To evaluate LRS, we selected a patient who presented with multiple neoplasia and cardiac myxomata suggestive of Carney complex for whom targeted clinical gene testing and whole genome SRS were negative. Low coverage whole genome LRS was performed on the PacBio Sequel system and structural variants were called, yielding 6,971 deletions and 6,821 insertions > 50bp. Filtering for variants that are absent in an unrelated control and that overlap a coding exon of a disease gene identified three deletions and three insertions. One of these, a heterozygous 2,184 bp deletion, overlaps the first coding exon of PRKAR1A, which is implicated in autosomal dominant Carney complex. This variant was confirmed by Sanger sequencing and was classified as pathogenic using standard criteria for the interpretation of sequence variants. This first successful application of whole genome LRS to identify a pathogenic variant suggests that LRS has significant potential to identify disease-causing structural variation. We recommend larger studies to evaluate the diagnostic yield of LRS, and the development of a comprehensive catalog of common human structural variation to support future studies.

genomics

Pan-Cancer Analysis Reveals Technical Artifacts in The Cancer Genome Atlas (TCGA) Germline Variant Calls

The degree to which germline variation drives cancer development and shapes tumor phenotypes remains largely unexplored, possibly due to a lack of large scale publicly available germline data for a cancer cohort. Here we called germline variants on 9,618 cases from The Cancer Genome Atlas (TCGA) database representing 31 cancer types. We identified batch effects affecting loss of function (LOF) variant calls that can be traced back to differences in the way the sequence data were generated both within and across cancer types. Overall, LOF indel calls were more sensitive to technical artifacts than LOF Single Nucleotide Variant (SNV) calls. In particular, whole genome amplification of DNA prior to sequencing led to an artificially increased burden of LOF indel calls, which confounded association analyses relating germline variants to tumor type despite stringent indel filtering strategies. Due to the inherent noise we chose to remove all 614 amplified DNA samples, including all acute myeloid leukemia and virtually all ovarian cancer samples, from the final dataset. This study demonstrates how insufficient quality control can lead to false positive germlinetumor type associations and draws attention to the need to be sensitive to problems associated with a lack of uniformity in data generation in TCGA data.\n\nAuthor SummaryCancer research to date has largely focused on genetic aberrations specific to tumor tissue. In contrast, the degree to which germline, or inherited, variation contributes to tumorigenesis remains unclear, possibly due to a lack of accessible germline variant data. In this study we identify germline variants in 9,618 samples using raw germline exome data from The Cancer Genome Atlas (TCGA). There are substantial differences in the way exome sequence data was generated both across and within cancer types in TCGA. We observe that differences in sequence data generation introduced batch effects, or variation that is due to technical factors not true biological variation, in our variant data. Most notably, we observe that amplification of DNA prior to sequencing resulted in an excess of predicted damaging indel variants. We show how these batch effects can confound germline association analyses if not properly addressed. Our study highlights the difficulties of working with large public genomic datasets like TCGA where samples are collected over time and across data centers, and particularly cautions the use of amplified DNA samples for genetic association analyses.

genomics

Functional Analysis of All Salmonid Genomes (FAASG): an international initiative supporting future salmonid research, conservation and aquaculture

We describe an emerging initiative - the Functional Analysis of All Salmonid Genomes (FAASG), which will leverage the extensive trait diversity that has evolved since a whole genome duplication event in the salmonid ancestor, to develop an integrative understanding of the functional genomic basis of phenotypic variation. The outcomes of FAASG will have diverse applications, ranging from improved understanding of genome evolution, through to improving the efficiency and sustainability of aquaculture production, supporting the future of fundamental and applied research in an iconic fish lineage of major societal importance.

genomics

Longitudinal samples of bacterial genomes potentially bias evolutionary analyses

Samples of bacteria collected over a period of time are attractive for several reasons, including the ability to estimate the molecular clock rate and to detect fluctuations in allele frequencies over time. However, longitudinal datasets are occasionally used in analyses that assume samples were collected contemporaneously. Using both simulations and genomic data from Neisseria gonorrhoeae, Streptococcus mutans, Campylobacter jejuni, and Helicobacter pylori, we show that longitudinal samples (spanning more than a decade in real data) may suffer from considerable bias that inflates estimates of recombination and the number of rare mutations in a sample of genomic sequences. While longitudinal data are frequently accounted for using the serial coalescent, many studies use other programs or metrics, such as Tajimas D, that are sensitive to these sampling biases and contain genomic data collected across many years. Notably, longitudinal samples from a population of constant size may exhibit evidence of exponential growth. We suggest that population genomic studies of bacteria should routinely account for temporal diversity in samples or provide evidence that longitudinal sampling bias does not affect conclusions.

genomics

De Novo PacBio long-read and phased avian genome assemblies correct and add to genes important in neuroscience research

Reference quality genomes are expected to provide a resource for studying gene structure and function. However, often genes of interest are not completely or accurately assembled, leading to unknown errors in analyses or additional cloning efforts for the correct sequences. A promising solution to this problem is long-read sequencing. Here we tested PacBio-based long-read sequencing and diploid assembly for potential improvements to the Sanger-based intermediate-read zebra finch reference and Illumina-based short-read Annas hummingbird reference, two vocal learning avian species widely studied in neuroscience and genomics. With DNA of the same individuals used to generate the reference genomes, we generated diploid assemblies with the FALCON-Unzip assembler, resulting in contigs with no gaps in the megabase range (N50s of 5.4 and 7.7 Mb, respectively), and representing 150-fold and 200-fold improvements over the current zebra finch and hummingbird references, respectively. These long-read assemblies corrected and resolved what we discovered to be misassemblies, including due to erroneous sequences flanking gaps, complex repeat structure errors in the references, base call errors in difficult to sequence regions, and inaccurate resolution of allelic differences between the two haplotypes. We analyzed protein-coding genes widely studied in neuroscience and specialized in vocal learning species, and found numerous assembly and sequence errors in the reference genes that the PacBio-based assemblies resolved completely, validated by single long genomic reads and transcriptome reads. These findings demonstrate, for the first time in non-human vocal learning species, the impact of higher quality, phased and gap-less assemblies for understanding gene structure and function.

genomics

Statistical Correction of the Winner’s Curse Explains Replication Variability in Quantitative Trait Genome-Wide Association Studies

Genome-wide association studies (GWAS) have identified hundreds of SNPs responsible for variation in human quantitative traits. However, genome-wide-significant associations often fail to replicate across independent cohorts, in apparent inconsistency with their apparent strong effects in discovery cohorts. This limited success of replication raises pervasive questions about the utility of the GWAS field. We identify all 332 studies of quantitative traits from the NHGRI-EBI GWAS Database with attempted replication. We find that the majority of studies provide insufficient data to evaluate replication rates. The remaining papers replicate significantly worse than expected (p < 10-14), even when adjusting for regression-to-the-mean of effect size between discovery- and replication-cohorts termed the Winners Curse (p < 10-16). We show this is due in part to misreporting replication cohort-size as a maximum number, rather than per-locus one. In 39 studies accurately reporting per-locus cohort-size for attempted replication of 707 loci in samples with similar ancestry, replication rate matched expectation (predicted 458, observed 457, p = 0.94). In contrast, ancestry differences between replication and discovery (13 studies, 385 loci) cause the most highly-powered decile of loci to replicate worse than expected, due to difference in linkage disequilibrium.\n\nAuthor SummaryThe majority of associations between common genetic variation and human traits come from genome-wide association studies, which have analyzed millions of single-nucleotide polymorphisms in millions of samples. These kinds of studies pose serious statistical challenges to discovering new associations. Finite resources restrict the number of candidate associations that can brought forward into validation samples, introducing the need for a significance threshold. This threshold creates a phenomenon called the Winners Curse, in which candidate associations close to the discovery threshold are more likely to have biased overestimates of the variants true association in the sampled population. We survey all human quantitative trait association studies that validated at least one signal. We find the majority of these studies do not publish sufficient information to actually support their claims of replication. For studies that did, we computationally correct the Winners Curse and evaluate replication performance. While all variants combined replicate significantly less than expected, we find that the subset of studies that (1) perform both discovery and replication in samples of the same ancestry; and (2) report accurate per-variant sample sizes, replicate as expected. This study provides strong, rigorous evidence for the broad reliability of genome-wide association studies. We furthermore provide a model for more efficient selection of variants as candidates for replication, as selecting variants using cursed discovery data enriches for variants with little real evidence for trait association.

genomics

Genomic analysis of P elements in natural populations of Drosophila melanogaster

The Drosophila melanogaster P transposable element provides one of the best cases of horizontal transfer of a mobile DNA sequence in eukaryotes. Invasion of natural populations by the P element has led to a syndrome of phenotypes known as P-M hybrid dysgenesis that emerges when strains differing in their P element composition mate and produce offspring. Despite extensive research on many aspects of P element biology, many questions remain about the genomic basis of variation in P-M dysgenesis phenotypes in natural populations. Here we compare gonadal dysgenesis phenotypes and genomic P element predictions for isofemale strains obtained from three worldwide populations of D. melanogaster to illuminate the molecular basis of natural variation in cytotype status. We show that the number of predicted P element insertions in genome sequences from isofemale strains is highly correlated across different bioinformatics methods, but the absolute number of insertions per strain is sensitive to method and filtering strategies. Regardless of method used, we find that the number of euchromatic P element insertions predicted per strain varies significantly across populations, with strains from a North American population having fewer P element insertions than strains from populations sampled in Europe or Africa. Despite these geographic differences, numbers of euchromatic P element insertions are not strongly correlated with the degree of gonadal dysgenesis exhibited by an isofemale strain. Thus, variation in P element insertion numbers across different populations does not necessarily lead to corresponding geographic differences in gonadal dysgenesis phenotypes. Additionally, we show that pool-seq samples can uncover population differences in the number of P element insertions observed from isofemale lines, but that efforts to rigorously detect differences in the number of P elements across populations using pool-seq data must properly control for read depth per strain. Our work supports the view that euchromatic P element copy number is not sufficient to explain variation in gonadal dysgenesis across strains of D. melanogaster, and informs future efforts to decode the genomic basis of geographic and temporal differences in P element induced phenotypes.

genomics

Holocene selection for variants associated with cognitive ability: Comparing ancient and modern genomes.

Human populations living in Eurasia during the Holocene experienced significant evolutionary change. It has been predicted that the transition of Holocene populations into agrarianism and urbanization brought about culture-gene co-evolution that favoured via directional selection genetic variants associated with higher general cognitive ability (GCA). Population expansion and replacement has also been proposed as an important source of GCA gene-frequency change during this time period. To examine whether GCA might have risen during the Holocene, we compare a sample of 99 ancient Eurasian genomes (ranging from 4,557 to 1,208 years of age) with a sample of 503 modern European genomes, using three different cognitive polygenic scores. Significant differences favouring the modern genomes were found for all three polygenic scores (Odds Ratio=0.92, p=0.037; 0.81, p=0.001 and 0.81, p=0.02). Furthermore, a significant increase in positive allele count over 3,249 years was found using a sample of 66 ancient genomes (r=0.217, pone-tailed=0.04). These observations are consistent with the expectation that GCA rose during the Holocene.

genomics

Genome-wide protein phylogenies for four African cichlid species

BackgroundThe thousands of species of closely related cichlid fishes in the great lakes of East Africa are a powerful model for understanding speciation and the genetic basis of trait variation. Recently, the genomes of five species of African cichlids representing five distinct lineages were sequenced and used to predict protein products at a genome-wide level. Here we characterize the evolutionary relationship of each cichlid protein to previously sequenced animal species.\n\nResultsWe used the Treefam database, a set of preexisting protein phylogenies built using 109 previously sequenced genomes, to identify Treefam families for each protein annotated from four cichlid species: Metriaclima zebra, Astatotilapia burtoni, Pundamilia nyererei and Neolamporologus brichardi. For each of these Treefam families, we built new protein phylogenies containing each of the cichlid protein hits. Using these new phylogenies we identified the evolutionary relationship of each cichlid protein to its nearest human and zebrafish protein. This data is available either through download or through a webserver we have implemented.\n\nConclusionThese phylogenies will be useful for any cichlid researchers trying to predict biological and protein function for a given cichlid gene, understanding the evolutionary history of a given cichlid gene, identifying recently duplicated cichlid genes, or performing genome-wide analysis in cichlids that relies on using databases generated from other species.

genomics

Comparative genomics of the tardigrades Hypsibius dujardini and Ramazzottius varieornatus

Tardigrada, a phylum of meiofaunal organisms, have been at the center of discussions of the evolution of Metazoa, the biology of survival in extreme environments, and the role of horizontal gene transfer in animal evolution. Tardigrada are placed as sisters to Arthropoda and Onychophora (velvet worms) in the superphylum Ecdysozoa by morphological analyses, but many molecular phylogenies fail to recover this relationship. This tension between molecular and morphological understanding may be very revealing of the mode and patterns of evolution of major groups. Similar to bdelloid rotifers, nematodes and other animals of the water film, limno-terrestrial tardigrades display extreme cryptobiotic abilities, including anhydrobiosis and cryobiosis. These extremophile behaviors challenge understanding of normal, aqueous physiology: how does a multicellular organism avoid lethal cellular collapse in the absence of liquid water? Meiofaunal species have been reported to have elevated levels of HGT events, but how important this is in evolution, and in particular in the evolution of extremophile physiology, is unclear. To address these questions, we resequenced and reassembled the genome of Hypsibius dujardini, a limno-terrestrial tardigrade that can undergo anhydrobiosis only after extensive pre-exposure to drying conditions, and compared it to the genome of Ramazzottius varieornatus, a related species with tolerance to rapid desiccation. The two species had contrasting gene expression responses to anhydrobiosis, with major transcriptional change in H. dujardini but limited regulation in R. varieornatus. We identified few horizontally transferred genes, but some of these were shown to be involved in entry into anhydrobiosis. Whole-genome molecular phylogenies supported a Tardigrada+Nematoda relationship over Tardigrada+Arthropoda, but rare genomic changes tended to support Tardigrada+Arthropoda.

genomics

Type 1 diabetes genome-wide association analysis with imputation identifies five new risk regions

Type 1 diabetes genotype datasets have undergone several well powered genome wide analysis studies (GWAS), identifying 57 associated regions at the time of analysis. There are still many regions of smaller effect size or low frequency left to discover, and better exploitation of existing type 1 diabetes cohorts with meta analysis and imputation can precede the acquisition of new or larger cohorts. An existing dataset of 5,913 case and 8,829 control samples was analysed using genome-wide microarrays (Affymetrix GeneChip 500K and Illumina Infinium 550K) with imputation via IMPUTE2 with the 1000 Genomes Project (phase 3) reference panel. Genotyping coverage was doubled in known association regions, and increased by four fold in other regions compared to previous studies. Our analysis resulted in new index variants for 17/57 regions, an expanded set of plausible candidate SNPs for 17 regions, and five novel type 1 diabetes association regions at 1p31.3, 1q24.3, 1q31.2, 2q11.2 and 11q12.2. Candidate genes for the new loci included ITGB3BP, FASLG, RGS1, AFF3 and CD5/CD6. Further prioritisation of causal genes and causal variants will require detailed RNA and protein expression studies, in conjunction with genome annotation studies including analysis of physical promoter-enhancer interactions.

genomics

Interactome INSIDER: A Multi-Scale Structural Interactome Browser For Genomic Studies

Protein interactions underlie nearly all known cellular function, making knowledge of their binding conformations paramount to understanding the physical workings of the cell. Studying binding conformations has allowed scientists to explore some of the mechanistic underpinnings of disease caused by disruption of protein interactions. However, since experimentally determined interaction structures are only available for a small fraction of the known interactome such inquiry has largely excluded functional genomic studies of the human interactome and broad observations of the inner workings of disease. Here we present Interactome INSIDER, an information center for genomic studies using the first full-interactome map of human interaction interfaces. We applied a new, unified framework to predict protein interaction interfaces for 184,605 protein interactions with previously unresolved interfaces in human and 7 model organisms, including the entire experimentally determined human binary interactome. We find that predicted interfaces share several known functional properties of interfaces, including an enrichment for disease mutations and recurrent cancer mutations, suggesting their applicability to functional genomic studies. We also performed 2,164 de novo mutagenesis experiments and show that mutations of predicted interface residues disrupt interactions at a similar rate to known interface residues and at a much higher rate than mutations outside of predicted interfaces. To spur functional genomic studies in the human interactome, Interactome INSIDER (http://interactomeinsider.yulab.org) allows users to explore known population variants, disease mutations, and somatic cancer mutations, or upload their own set of mutations to find enrichment at the level of protein domains, residues, and 3D atomic clustering in known and predicted interaction interfaces.

genomics

Evolutionary Dynamics Of Genome-Wide Position Effects In Mammals

Conserved noncoding elements (CNEs) have significant regulatory influence on their neighbouring genes. Loss of synteny to CNEs through genomic rearrangements can, therefore, impact the transcriptional states of the cognate genes. Yet, the evolutionary implications of such chromosomal position effects have not been studied. Through genome-wide analysis of CNEs and the cognate genes of representative species from 5 different mammalian orders, we observed significant loss of synteny to CNEs in rat lineage. The CNEs and genes losing synteny had significant association with the fetal, but not the post-natal, brain development as assessed through ontology terms, developmental gene expression, chromatin marks and genetic mutations. The loss of synteny correlated with the independent evolutionary loss of fetus-specific upregulation of genes in rat brain. DNA-breakpoints implicated in brain abnormalities of germ-line origin had significant representation between CNE and the gene that exhibited loss of synteny, signifying the underlying developmental tolerance of genomic rearrangements that had allowed the evolutionary splits of CNEs and the cognate genes in rodent lineage. These observations highlighted the non-trivial impact of chromosomal position-effect in shaping the evolutionary dynamics of mammalian brain development and might explain loss of brain traits, like cerebral folding of cortex, in rodent lineage.\n\nAuthor SummaryExpression of genes is regulated by proximally located non-coding regulatory elements. Loss of linear proximity between gene and its regulatory element thus can alter the expression of gene. Such a phenomenon can be tested at whole genome scale using evolutionary methods. We compared the positions of genes and regulatory elements in 5 different mammals and identified the significant loss of proximities between gene and their regulatory elements in rat during evolution. Brain development related function was selectively enriched among the genes and regulatory elements that had lost the proximity in rat. The observed separation of genes and their regulatory elements was strongly associated with the evolutionary loss of developmental gene expression pattern in rat brain, which coincided with the loss of brain traits in rodents. The study highlighted the importance of relative chromosomal positioning of genes and their gene regulatory elements in the evolution of phenotypes.

genomics

The Genomic Landscape Of Tree Rot In Phellinus noxius And Its Hymenochaetales Members

The order Hymenochaetales of white rot fungi contain some of the most aggressive wood decayers causing tree deaths around the world. Despite their ecological importance and the impact of diseases they cause, little is known about the evolution and transmission patterns of these pathogens. Here, we sequenced and undertook comparative genomics analyses of Hymenochaetales genomes using brown root rot fungus Phellinus noxius, wood-decomposing fungus Phellinus lamaensis, laminated root rot fungus Phellinus sulphurascens, and trunk pathogen Porodaedalea pini. Many gene families of lignin-degrading enzymes were identified from these fungi, reflecting their ability as white rot fungi. Comparing against distant fungi highlighted the expansion of 1,3-beta-glucan synthases in P. noxius, which may account for its fast-growing attribute. We identified 13 linkage groups conserved within Agaricomycetes, suggesting the evolution of stable karyotypes. We determined that P. noxius has a bipolar heterothallic mating system, with unusual highly expanded ~60 kb A locus as a result of accumulating gene transposition. We investigated the population genomics of 60 P. noxius isolates across multiple islands of the Asia Pacific region. Whole-genome sequencing showed this multinucleate species contains abundant poly-allelic single-nucleotide-polymorphisms (SNPs) with atypical allele frequencies. Different patterns of intra-isolate polymorphism reflect mono-/heterokaryotic states which are both prevalent in nature. We have shown two genetically separated lineages with one spanning across many islands despite the geographical barriers. Both populations possess extraordinary genetic diversity and show contrasting evolutionary scenarios. These results provide a framework to further investigate the genetic basis underlying the fitness and virulence of white rot fungi.

genomics

Population Genomics Of Cryptococcus neoformans var. grubii Reveals New Biogeographic Relationships And Finely Maps Hybridization

Cryptococcus neoformans var. grubii is the causative agent of cryptococcal meningitis, a significant source of mortality in immunocompromised individuals, typically HIV/AIDS patients from developing countries. Despite the worldwide emergence of this ubiquitous infection, little is known about the global molecular epidemiology of this fungal pathogen. Here we sequence the genomes of 188 diverse isolates and characterized the major subdivisions, their relative diversity and the level of genetic exchange between them. While most isolates of C. neoformans var. grubii belong to one of three major lineages (VNI, VNII, and VNB), some haploid isolates show hybrid ancestry including some that appear to have recently interbred, based on the detection of large blocks of each ancestry across each chromosome. Many isolates display evidence of aneuploidy, which was detected for all chromosomes. In diploid isolates of C. neoformans var. grubii (serotype A/A) and of hybrids with C. neoformans var. neoformans (serotype A/D) such aneuploidies have resulted in loss of heterozygosity, where a chromosomal region is represented by the genotype of only one parental isolate. Phylogenetic and population genomic analyses of isolates from Brazil revealed that the previously African VNB lineage occurs naturally in the South American environment. This suggests migration of the VNB lineage between Africa and South America prior to its diversification, supported by finding ancestral recombination events between isolates from different lineages and regions. The results provide evidence of substantial population structure, with all lineages showing multi-continental distributions demonstrating the highly dispersive nature of this pathogen.\n\nAuthor SummaryCryptococcus neoformans var. grubii is a human fungal pathogen of immunocompromised individuals that has global clinical impact, causing half a million deaths per year. Substantial genetic substructure exists for this pathogen, with two lineages found globally (VNI, VNII) whereas a third has appeared confined to sub-Saharan Africa (VNB). Here, we utilized genome sequencing of a large set of global isolates to examine the genetic diversity, hybridization, and biogeography of these lineages. We found that while the three major lineages are well separated, recombination between the lineages has occurred, notably resulting in hybrid isolates with segmented ancestry across the genome. In addition, we showed that isolates from South America are placed within the VNB lineage, formerly thought to be confined to Africa, and that there is phylogenetic separation between these geographies that substantially expands the diversity of these lineages. Our findings provide a new framework for further studies of the dynamics of natural populations of C. neoformans var. grubii.

genomics