Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,387 records · Page 77Linked to original sources

Complex Structural Variants Resolved by Short-Read and Long-Read Whole Genome Sequencing in Mendelian Disorders

Complex structural variants (cxSVs) are genomic rearrangements comprising multiple structural variants, typically involving three or more breakpoint junctions. They contribute to human genomic variation and can cause Mendelian disease, however they are not typically considered during genetic testing. Here, we investigate the role of cxSVs in Mendelian disease using short-read whole genome sequencing (WGS) data from 1,324 individuals with neurodevelopmental or retinal disorders from the NIHR BioResource project. We present four cases of individuals with a cxSV affecting Mendelian disease-associated genes. Three of the cxSVs are pathogenic: a de novo duplication-inversion-inversion-deletion affecting ARID1B in an individual with Coffin-Siris syndrome, a deletion-inversion-duplication affecting HNRNPU in an individual with intellectual disability and seizures, and a homozygous deletion-inversion-deletion affecting CEP78 in an individual with cone-rod dystrophy. Additionally, we identified a de novo duplication-inversion-duplication overlapping CDKL5 in an individual with neonatal hypoxic-ischaemic encephalopathy. Long-read sequencing technology used to resolve the breakpoints demonstrated the presence of both a disrupted and an intact copy of CDKL5 on the same allele; therefore, it was classified as a variant of uncertain significance. Analysis of sequence flanking all breakpoint junctions in all the cxSVs revealed both microhomology and longer repetitive sequences, suggesting both replication and homology based processes. Accurate resolution of cxSVs is essential for clinical interpretation, and here we demonstrate that long-read WGS is a powerful technology by which to achieve this. Our results show cxSVs are an important although rare cause of Mendelian disease, and we therefore recommend their consideration during research and clinical investigations.

genomics

Nucleosome remodelling at origins of Global Genome-Nucleotide Excision Repair occurs at the boundaries of higher-order chromatin structure

Repair of UV-induced DNA damage requires chromatin remodeling. How repair is initiated in chromatin remains largely unknown. We recently demonstrated that Global Genome Nucleotide Excision Repair (GG-NER) in chromatin is organized into domains around open reading frames. Here, we identify these domains, and by examining DNA damage-induced changes in the linear structure of nucleosomes, we demonstrate how chromatin remodeling is initiated during repair. In undamaged cells, we show that the GG-NER complex occupies chromatin at nucleosome free regions of specific gene promoters. This establishes the nucleosome structure at these genomic locations, which we refer to as GG-NER complex binding sites (GCBSs). We demonstrate that these sites are frequently located at genomic boundaries that delineate chromasomally interacting domains (CIDs). These boundaries define domains of higher-order nucleosome-nucleosome interaction. We show that efficient repair of DNA damage in chromatin is initiated following disruption of H2A.Z-containing nucleosomes adjacent to GCBSs by the GG-NER complex.

genomics

Recovering signals of ghost archaic admixture in the genomes of present-day Africans

While introgression from Neanderthals and Denisovans has been well-documented in modern humans outside Africa, the contribution of archaic hominins to the genetic variation of present-day Africans remains poorly understood. Using 405 whole-genome sequences from four sub-Saharan African populations, we provide complementary lines of evidence for archaic introgression into these populations. Our analyses of site frequency spectra indicate that these populations derive 2-19% of their genetic ancestry from an archaic population that diverged prior to the split of Neanderthals and modern humans. Using a method that can identify segments of archaic ancestry without the need for reference archaic genomes, we built genome-wide maps of archaic ancestry in the Yoruba and the Mende populations that recover about 482 and 502 megabases of archaic sequence, respectively. Analyses of these maps reveal segments of archaic ancestry at high frequency in these populations that represent potential targets of adaptive introgression. Our results reveal the substantial contribution of archaic ancestry in shaping the gene pool of present-day African populations. One sentence summaryMultiple present-day African populations inherited genes from an unknown archaic population that diverged before modern humans and Neanderthals split.

genomics

The Genomic Prediction of Disease: Example of type 2 diabetes (T2D)

Application of concepts from information theory have revealed new features of Single Nucleotide Polymorphism (SNP) organization.. These features lead to effective classifiers by which to distinguish genomic sequences of contrasting phenotypes; as in case/control cohorts.\n\nWhen applied to a disease/control database, a disease classifier results; a parallel analysis leads to the determination of a wellness classifier. The classifiers have non-intersecting loci, and each involves roughly 100 alleles.\n\nThe effectiveness of this framework is illustrated by application to adult onset, type 2, diabetes (T2D), as represented in the Wellcome Trust ((WT) Case/Control database.\n\nSimultaneous use of the two classifiers on the WT database leads to successful prediction of disease versus wellness; to the extent that near certain genomic forecasting is achieved.\n\nThis framework gives a resolution to the oft posed uncertainty: \"Where is the missing heritability?\"\n\nApplication of both classifiers on two additional T2D databases produced informative consequences.\n\nA fully independent, compelling, confirmation of the present results is obtained by means of the machine learning algorithm, Random Forests.\n\nThe analytical model presented here is generalizable to other diseases.\n\nOne Sentence SummaryDiscovery of intrinsic chromosomal SNP organizations leads to near certain genomic disease prediction.

genomics

Complete genome direct RNA sequencing of influenza A virus

For the first time, a complete genome of an RNA virus has been sequenced in its original form. Previously, RNA was sequenced by the chemical degradation of radiolabelled RNA, a difficult method that produced only short sequences. Instead, RNA has usually been sequenced indirectly by copying it into cDNA, which is often amplified to dsDNA by PCR and subsequently analyzed using a variety of DNA sequencing methods. We designed an adapter to short highly conserved termini of the influenza virus genome to target the (-) sense RNA into a protein nanopore on the Oxford Nanopore MinION sequencing platform. Utilizing this method and total RNA extracted from the allantoic fluid of infected chicken eggs, we demonstrate successful sequencing of the complete influenza virus genome with 100% nucleotide coverage, 99% consensus identity, and 99% of reads mapped to influenza. By utilizing the same methodology we can redesign the adapter in order to expand the targets to include viral mRNA and (+) sense cRNA, which are essential to the viral life cycle. This has the potential to identify and quantify splice variants and base modifications, which are not practically measurable with current methods.

genomics

Sequencing and de novo assembly reveals genomic variations associated with differential responses of Candida albicans ATCC 10231 towards fluconazole, pH and non-invasive growth

The whole genome sequencing generated 8.09 million paired-end reads, which were assembled into 2,262 scaffolds totaling 17,113,050 bp in length. We predict 7,654 coding regions and 7,647 genes, and annotated the coding regions with gene ontology terms based on similarity to other annotated genomes. Genome comparisons revealed variations (including SNP and indels in genes involved in pH response, fluconazole resistance and invasive growth) compared to C. albicans SC5314 and WO1.

genomics

crisprQTL mapping as a genome-wide association framework for cellular genetic screens

Expression quantitative trait locus (eQTL) and genome-wide association studies (GWAS) are powerful paradigms for mapping the determinants of gene expression and organismal phenotypes, respectively. However, eQTL mapping and GWAS are limited in scope (to naturally occurring, common genetic variants) and resolution (by linkage disequilibrium). Here, we present crisprQTL mapping, a framework in which large numbers of CRISPR/Cas9 perturbations are introduced to each cell on an isogenic background, followed by single-cell RNA-seq (scRNA-seq). crisprQTL mapping is analogous to conventional human eQTL studies, but with individual humans replaced by individual cells; genetic variants replaced by unique combinations of unlinked guide RNA (gRNA)-programmed perturbations per cell; and tissue-level RNA-seq of many individuals replaced by scRNA-seq of many cells. By randomly introducing gRNAs, a single population of cells can be leveraged to test for association between each perturbation and the expression of any potential target gene, analogous to how eQTL studies leverage populations of humans to test millions of genetic variants for associations with expression in a genome-wide manner. However, crisprQTL mapping is neither limited to naturally occurring, common genetic variants nor by linkage disequilibrium. As a proof-of-concept, we applied crisprQTL mapping to evaluate 1,119 candidate enhancers with no strong a priori hypothesis as to their target gene(s). Perturbations were made by a nuclease-dead Cas9 (dCas9) tethered to KRAB, and introduced at a mean allele frequency of 1.1% into a population of 47,650 profiled human K562 cells (median of 15 gRNAs identified per cell). We tested for differential expression of all genes within 1 megabase (Mb) of each candidate enhancer, effectively evaluating 17,584 potential enhancer-target gene relationships within a single experiment. At an empirical false discovery rate (FDR) of 10%, we identify 128 cis crisprQTLs (11%) whose targeting resulted in downregulation of 105 nearby genes. crisprQTLs were strongly enriched for proximity to their target genes (median 34.3 kilobases (Kb)) and the strength of H3K27ac, p300, and lineage-specific transcription factor (TF) ChIP-seq peaks. Our results establish the power of the eQTL mapping paradigm as applied to programmed variation in populations of cells, rather than natural variation in populations of individuals. We anticipate that crisprQTL mapping will facilitate the comprehensive elucidation of the cis-regulatory architecture of the human genome.

genomics

Human genomics of acute liver failure due to hepatitis B virus infection: an exome sequencing study in liver transplant recipients

Acute liver failure (ALF) or fulminant hepatitis is a rare, yet severe outcome of infection with hepatitis B virus (HBV) that carries a high mortality rate. The occurrence of a life-threatening condition upon infection with a prevalent virus in individuals without known risk factors is suggestive of pathogen-specific immune dysregulation. In the absence of established differences in HBV virulence, we hypothesized that ALF upon primary infection with HBV could be due to rare deleterious variants in the human genome. To search for such variants, we performed exome sequencing in 21 previously healthy adults who required liver transplantation upon fulminant HBV infection and 172 controls that were positive for anti-HBc and anti-HBs antibodies but had no clinical history of jaundice or liver disease. After a series of hypothesis-driven filtering steps, we searched for putatively pathogenic variants that were significantly associated with case-control status. We did not find any causal variant or gene, a result that does not support the hypothesis of a shared monogenic basis for human susceptibility to HBV-related ALF in adults. This study represents a first attempt at deciphering the human genetic contribution to the most severe clinical presentation of acute HBV infection in previously healthy individuals.\n\nAuthor SummaryInfection with hepatitis B virus (HBV) is very common and causes a variety of liver diseases including acute and chronic hepatitis, cirrhosis and liver carcinoma. Acute HBV infection is often asymptomatic, still about 1% of newly infected people develop a rapid and severe disease known as acute liver failure or fulminant hepatitis. Acute liver failure has a high mortality rate and is an indication for urgent liver transplantation. It is not clear why some people, who are otherwise healthy, develop such severe symptoms upon infection with a common pathogen. Here, we hypothesized that rare DNA variants in the human genome could contribute to this unusual susceptibility. We sequenced the exome (i.e. the regions of the genome that encode the proteins) of 21 previously healthy adults who required liver transplantation upon fulminant HBV infection and searched for rare genetic variants that could explain the clinical presentation. We did not identify any variant that could be convincingly linked to the extreme susceptibility to HBV observed in the study participants. This suggests that HBV-induced acute liver failure is more likely to result from the combined influence of multiple genetic and environmental factors.

genomics

Genome-wide association analysis identifies 27 novel loci associated with uterine leiomyomata revealing common genetic origins with endometriosis

Uterine leiomyomata (UL), also known as uterine fibroids, are the most common neoplasms of the reproductive tract and the primary cause for hysterectomy, leading to considerable impact on womens lives as well as high economic burden1,2. Genetic epidemiologic studies indicate that heritable risk factors contribute to UL pathogenesis3. Previous genome-wide association studies (GWAS) identified five loci associated with UL at genome-wide significance (P < 5 x 10-8)4-6. We conducted GWAS meta-analysis in 20,406 cases and 223,918 female controls of white European ancestry, identifying 24 genome-wide significant independent loci; 17 replicated in an unrelated cohort of 15,068 additional cases and 43,587 female controls. Aggregation of discovery and replication studies (35,474 cases and 267,505 female controls) revealed six additional significant loci. Interestingly, four of the 17 loci identified and replicated in these analyses have also been associated with risk for endometriosis - another common gynecologic disorder. These findings increase our understanding of the biological mechanisms underlying UL development, and suggest overlapping genetic origins with endometriosis.

genomics

Sense-antisense gene overlap causes evolutionary retention of the few introns in Giardia genome and the implications

BackgroundIt is widely accepted that the last eukaryotic common ancestor (LECA) and early eukaryotes were intron-rich and intron loss dominated subsequent evolution, thus the presence of only very few introns in some modern eukaryotes must be the consequence of massive loss. But it is striking that few eukaryotes were found to have completely lost introns. Despite extensive research, the causes of massive intron losses remain elusive, and actually the reverse question - how the few introns are retained under the pressure of loss is equally significant but was rarely studied, except that it was conjectured that the essential functions of some introns prevent their loss. The extremely few (eight) spliceosome-mediated cis-spliced introns in the relatively simple genome of Giardia lamblia provide an excellent opportunity to explore this question.\n\nResultsOur investigation of the intron-containing genes and introns in Giardia found three types of intron distribution patterns: ancient intron in ancient gene, relatively new intron in ancient gene, and relatively new intron in relatively new gene, which can reflect to some extent the dynamic evolution of introns in Giardia. Not finding any special features or functional importance of these introns responsible for the retention, we noticed and experimentally verified that some intron-containing genes form sense-antisense gene pairs with functional genes on their complementary strands, and that the introns just reside in the overlapping regions.\n\nConclusionsIn Giardias evolution, despite constant pressure of intron loss, intron gain can still occur in both ancient and newly-evolved genes, but only a few introns have been retained; the evolutionary retention of introns is most likely not due to the functional constraint of the introns themselves but the causes outside of introns, such as the constraints imposed by other genomic functional elements overlapping with the introns. These findings can not only provide some clues to find new genomic functional elements -- in the areas overlapping with introngs, but suggest that \"functional constraint\" of introns may not be necessarily directly associated with intron loss and gain, or that the real functions or the way of functioning of introns are probably still outside of our current knowledge.

genomics

Genomic meta-analysis of the interplay between 3D chromatin organization and gene expression programs under basal and stress conditions

BackgroundOur appreciation of the critical role of the 3D organization of the genome in gene regulation is steadily increasing. Recent 3C-based deep sequencing techniques elucidated a hierarchy of structures that underlie the spatial organization of the genome in the nucleus. At the top of this hierarchical organization are chromosomal territories and the megabase-scale A/B compartments that correlate with transcriptional activity within cells. Below them are the relatively cell-type invariant topologically associated domains (TADs), characterized by high frequency of physical contacts between loci within the same TAD and are assumed to function as regulatory units. Within TADs, chromatin loops bring enhancers and target promoters to close spatial proximity. Yet, we still have only rudimentary understanding how differences in chromatin organization between different cell types affect cell-type specific gene expression programs that are executed under basal and challenged conditions.\n\nResultsHere, we carried out a large-scale meta-analysis that integrated Hi-C data from thirteen different cell lines and dozens of ChIP-seq and RNA-seq datasets measured on these cells, either under basal conditions or after treatment. Pairwise comparisons between cell lines demonstrated the strong association between modulation of A/B compartmentalization, differential gene expression and transcription factor (TF) binding events. Furthermore, integrating the analysis of transcriptomes of different cell lines in response to various challenges, we show that 3D organization of cells under basal conditions constrains not only gene expression programs and TF binding profiles that are active under the basal condition but also those induced in response to treatment.\n\nConclusionsOur results further elucidate the role of dynamic genome organization in regulation of differential gene expression between different cell types, and indicate the impact of intra-TAD enhancer-promoter interactions that are established under basal conditions on both the basal and treatment-induced gene expression programs.

genomics

A comparative genome analysis of Rift Valley Fever virus isolates from foci of the disease outbreak in South Africa in 2008-2010

Rift Valley fever (RVF) is a re-emerging zoonotic disease responsible for major losses in livestock production, with negative impact on the livelihoods of both commercial and resource-poor farmers in sub-Sahara African countries. The disease remains a threat in countries where its mosquito vectors thrives. Outbreaks of RVF usually follow weather conditions which favour increase in mosquito populations. Such outbreaks are usually cyclical, occurring every 10-15 years.\n\nRecent outbreaks of the disease in South Africa have occurred unpredictably and with increased frequency. In 2008 outbreaks were reported in Mpumalanga, Limpopo and Gauteng provinces, followed by a 2009 outbreak in KwaZulu-Natal, Mpumalanga and Northern Cape provinces and in 2010 in the Eastern Cape, Northern Cape, Western Cape, North West, Free State and Mpumalanga provinces. By August 2010, 232 confirmed infections had been reported in humans, with 26 confirmed deaths.\n\nTo investigate the evolutionary dynamics of RVF viruses (RVFVs) circulating in South Africa, we undertook complete genome sequence analysis of isolates from animals at discrete foci of the 2008-2010 outbreaks. The genome sequences of these viruses were compared with viruses from earlier outbreaks in South Africa and in other countries. The data indicates that one 2009 and all the 2008 isolates from South Africa and Madagascar (M49/08) cluster in Lineage C or Kenya-1. The remaining of the 2009 and 2010 isolates cluster within Lineage H, except isolate M259_RSA_09, a probable segment M reassortant.\n\nAuthor summaryA single RVF virus serotype exists, yet differences in virulence and pathogenicity of the virus have been observed. This necessitates the need for detailed genetic characterization of various isolates of the virus. The RVF virus isolates that caused the 2008-2010 disease outbreaks in South Africa were most probably reassortants. Reassortment results from exchange of portions of the genome, particularly those of segment M. Although clear association between RVFV genotype and phenotype has not been established, various amino acid substitutions have been implicated in the phenotype. Viruses with amino acid substitutions from glycine to glutamic acid at position 277 of segment M have been shown to be more virulent in mice in comparison to viruses with glycine at the same position. Phylogenetic analysis indicated the viruses responsible for the 2008-2010 RVF outbreaks in South Africa were not introduced from outside the country, but mutated in time and caused the outbreaks when environmental conditions became favourable.

genomics

Genome-scale oscillations in DNA methylation during exit from pluripotency

Pluripotency is accompanied by the erasure of parental epigenetic memory with naive pluripotent cells exhibiting global DNA hypomethylation both in vitro and in vivo. Exit from pluripotency and priming for differentiation into somatic lineages is associated with genome-wide de novo DNA methylation. We show that during this phase, coexpression of enzymes required for DNA methylation turnover, DNMT3s and TETs, promotes cell-to-cell variability in this epigenetic mark. Using a combination of single-cell sequencing and quantitative biophysical modelling, we show that this variability is associated with coherent, genome-scale, oscillations in DNA methylation with an amplitude dependent on CpG density. Analysis of parallel single-cell transcriptional and epigenetic profiling provides evidence for oscillatory dynamics both in vitro and in vivo. These observations provide fresh insights into the emergence of epigenetic heterogeneity during early embryo development, indicating that dynamic changes in DNA methylation might influence early cell fate decisions.\n\nHighlightsO_LICo-expression of DNMT3s and TETs drive genome-scale oscillations of DNA methylation\nC_LIO_LIOscillation amplitude is greatest at a CpG density characteristic of enhancers\nC_LIO_LICell synchronisation reveals oscillation period and link with primary transcripts\nC_LIO_LIMultiomic single-cell profiling provides evidence for oscillatory dynamics in vivo\nC_LI

genomics

Whole genome sequencing, de novo assembly and phenotypic profiling for the new budding yeast species Saccharomyces jurei

Saccharomyces sensu stricto complex consist of yeast species, which are not only important in the fermentation industry but are also model systems for genomic and ecological analysis. Here, we present the complete genome assemblies of Saccharomyces jurei, a newly discovered Saccharomyces sensu stricto species from high altitude oaks. Phylogenetic and phenotypic analysis revealed that S. jurei is a sister-species to S. mikatae, than S. cerevisiae, and S. paradoxus. The karyotype of S. jurei presents two reciprocal chromosomal translocations between chromosome VI/VII and I/XIII when compared to S. cerevisiae genome. Interestingly, while the rearrangement I/XIII is unique to S. jurei, the other is in common with S. mikatae strain IFO1815, suggesting shared evolutionary history of this species after the split between S. cerevisiae and S. mikatae. The number of Ty elements differed in the new species, with a higher number of Ty elements present in S. jurei than in S. cerevisiae. Phenotypically, the S. jurei strain NCYC 3962 has relatively higher fitness than the other strain NCYC 3947T under most of the environmental stress conditions tested and showed remarkably increased fitness in higher concentration of acetic acid compared to the other sensu stricto species. Both strains were found to be better adapted to lower temperatures compared to S. cerevisiae.

genomics

Whole-genome sequences suggest long term declines of spotted owl (Strix occidentalis) (Aves: Strigiformes: Strigidae) populations in California

We analyzed whole-genome data of four spotted owls (Strix occidentalis) to provide a broad-scale assessment of the genome-wide nucleotide diversity across S. occidentalis populations in California. We assumed that each of the four samples was representative of its population and we estimated effective population sizes through time for each corresponding population. Our estimates provided evidence of long-term population declines in all California S. occidentalis populations. We found no evidence of genetic differentiation between northern spotted owl (S. o. caurina) populations in the counties of Marin and Humboldt in California. We estimated greater differentiation between populations at the northern and southern extremes of the range of the California spotted owl (S. o. occidentalis) than between populations of S. o. occidentalis and S. o. caurina in northern California. The San Diego County S. o. occidentalis population was substantially diverged from the other three S. occidentalis populations. These whole-genome data support a pattern of isolation-by-distance across spotted owl populations in California, rather than elevated differentiation between currently recognized subspecies.

genomics

Re-identification of genomic data using long range familial searches

Consumer genomics databases reached the scale of millions of individuals. Recently, law enforcement investigators have started to exploit some of these databases to find distant familial relatives, which can lead to a complete re-identification. Here, we leveraged genomic data of 600,000 individuals tested with consumer genomics to investigate the power of such long-range familial searches. We project that half of the searches with European-descent individuals will result with a third cousin or closer match and will provide a search space small enough to permit re-identification using common demographic identifiers. Moreover, in the near future, virtually any European-descent US person could be implicated by this technique. We propose a potential mitigation strategy based on cryptographic signature that can resolve the issue and discuss policy implications to human subject research.

genomics

Whole Proteome Clustering of 2,307 Genomes Reveals Remarkable Conservation of Four Proteins Among Proteobacteria While Revealing Significant Annotation Issues

To explore the concept of a minimal gene set, we clustered 8.76 M protein sequences deduced from 2,307 completely sequenced Proteobacterial genomes. To our knowledge this is the first study of this scale. Clustering resulted in 707,311 clusters of which 224,442 ranged in size from 2 to 2,894 sequences. The resulting clusters allowed us to ask the question: Is a set of proteins conserved across all Proteobacteria? We chose four essential proteins, the chaperonin GroEL, DNA dependent RNA polymerase subunits beta and beta (RpoB/RpoB), and DNA polymerase I (PolA), representing fundamental cellular functions, and examined their distribution in the clusters. We found these proteins to be remarkably conserved. Although the groEL gene was universally conserved in all the organisms in the study, the protein was not represented in all the deduced proteomes. The genes for RpoB and RpoB were missing from two genomes and merged in 88 genomes, and the sequences were sufficiently divergent that they formed separate clusters for 18 RpoB proteins (seven clusters) and 14 RpoB proteins (three clusters). For PolA, 52 organisms lacked an identifiable sequence, and seven sequences were sufficiently divergent that they formed five separate clusters. Interestingly, organisms lacking an identifiable PolA and those with divergent RpoB/RpoB were almost all endosymbionts. Furthermore, we present a range of examples of annotation issues that caused the deduced proteins to be incorrectly represented in the proteome. These annotation issues represent a significant obstacle for high throughput analyses.

genomics

Demographic inference in a spatially-explicit ecological model from genomic data: a proof of concept for the Mojave Desert Tortoise

In this paper, we study the general problem of extracting information from spatially explicit genomic data to inform inference of ecologically and geographically realistic population models. We describe methods and apply them to simulations motivated by the demography of the Mojave desert tortoise (Gopherus agassizii). The tortoise is an example of a long-lived, threatened species for which we have an excellent understanding of range, habitat preference, and certain aspects of demography, but inadequate information on other life history components that are important for conservation management. We use an individual-based model on a discretized geographic landscape with overlapping generations and age and sex-specific dispersal, fecundity, and mortality to develop and test a method that uses genomic data to infer demographic parameters. We do this by seeking parameters that best match a set of spatial statistics of genomes, which we introduce and discuss. We find that for inferring only overall population density and mean migration distance, a simple statistical learning method performs well using simulated training data, inferring parameters to within 10% accuracy. In the process, we introduce spatial analogues of common population genetics statistics, and discuss how and why they are expected to contain signal about the geography of population dynamics that are key for ecological modeling generally and conservation of endangered taxa.

genomics