Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Natural Selection Shapes the Mosaic Ancestry of the Drosophila Genetic Reference Panel and the D. melanogaster Reference Genome

North American populations of Drosophila melanogaster are thought to derive from both European and African source populations, but despite their importance for genetic research, patterns of admixture along their genomes are essentially undocumented. Here, I infer geographic ancestry along genomes of the Drosophila Genetic Reference Panel (DGRP) and the D. melanogaster reference genome. Overall, the proportion of African ancestry was estimated to be 20% for the DGRP and 9% for the reference genome. Based on the size of admixture tracts and the approximate timing of admixture, I estimate that the DGRP population underwent roughly 13.9 generations per year. Notably, ancestry levels varied strikingly among genomic regions, with significantly less African introgression on the X chromosome, in regions of high recombination, and at genes involved in specific processes such as circadian rhythm. An important role for natural selection during the admixture process was further supported by a genome-wide signal of ancestry disequilibrium, in that many between-chromosome pairs of loci showed a deficiency of Africa-Europe allele combinations. These results support the hypothesis that admixture between partially genetically isolated Drosophila populations led to natural selection against incompatible genetic variants, and that this process is ongoing. The ancestry blocks inferred here may be relevant for the performance of reference alignment in this species, and may bolster the design and interpretation of many population genetic and association mapping studies.

Evolutionary Biology

Is there such a thing as Landscape Genetics?

For a scientific discipline to be interdisciplinary it must satisfy two conditions; it must consist of contributions from at least two existing disciplines and it must be able to provide insights, through this interaction, that neither progenitor discipline could address. In this paper, I examine the complete body of peer-reviewed literature self-identified as landscape genetics using the statistical approaches of text mining and natural language processing. The goal here is to quantify the kinds of questions being addressed in landscape genetic studies, the ways in which questions are evaluated mechanistically, and how they are differentiated from the progenitor disciplines of landscape ecology and population genetics. I then circumscribe the main factions within published landscape genetic papers examining the extent to which emergent questions are being addressed and highlighting a deep bifurcation between existing individual- and population-based approaches. I close by providing some suggestions on where theoretical and analytical work is needed if landscape genetics is to serve as a real bridge connecting evolution and ecology sensu lato.

Evolutionary Biology

Application of a dense genetic map for assessment of genomic responses to selection and inbreeding in Heliothis virescens.

Adaptation of pest species to laboratory conditions and selection for resistance to toxins in the laboratory are expected to cause inbreeding and genetic bottlenecks that reduce genetic variation. Heliothis virescens, a major cotton pest, has been colonized in the laboratory many times, and a few laboratory colonies have been selected for Bt resistance. We developed 350 bp Double-Digest Restriction-site Associated DNA-sequencing (ddRAD-seq) molecular markers to examine and compare changes in genetic variation associated with laboratory adaptation, artificial selection, and inbreeding in this non-model insect species. We found that allelic and nucleotide diversity declined dramatically in laboratory-reared H. virescens as compared with field-collected populations. The declines were primarily due to the loss of low frequency alleles present in field-collected H. virescens. A further, albeit modest decline in genetic diversity was observed in a Bt-selected population. The greatest decline was seen in H. virescens that were sib-mated for 10 generations, where more than 80% of loci were fixed for a single allele. To determine which regions of the genome were resistant to fixation in our sib-mated line, we generated a dense intraspecific linkage map containing 3 PCR-based, and 659 ddRAD-seq markers. Markers that retained polymorphism were observed in small clusters spread over multiple linkage groups, but this clustering was not statistically significant. Here, we confirmed and extended the general expectations for reduced genetic diversity in laboratory colonies, provided tools for further genomic analyses, and produced highly homozygous genomic DNA for future whole genome sequencing of H. virescens.

Genomics

Quantitative Genetics Meets Integral Projection Models: Unification of Widely Used Methods from Ecology and Evolution

O_LIMicro-evolutionary predictions are complicated by ecological feedbacks like density dependence, while ecological predictions can be complicated by evolutionary change. A widely used approach in micro-evolution, quantitative genetics, struggles to incorporate ecological processes into predictive models, while structured population modelling, a tool widely used in ecology, rarely incorporates evolution explicitly.\nC_LIO_LIIn this paper we develop a flexible, general framework that links quantitative genetics and structured population models. We use the quantitative genetic approach to write down the phenotype as an additive map. We then construct integral projection models for each component of the phenotype. The dynamics of the distribution of the phenotype are generated by combining distributions of each of its components. Population projection models can be formulated on per generation or on shorter time steps.\nC_LIO_LIWe introduce the framework before developing example models with parameters chosen to exhibit specific dynamics. These models reveal (i) how evolution of a phenotype can cause populations to move from one dynamical regime to another (e.g. from stationarity to cycles), (ii) how additive genetic variances and covariances (the G matrix) are expected to evolve over multiple generations, (iii) how changing heritability with age can maintain additive genetic variation in the face of selection and (iii) life history, population dynamics, phenotypic characters and parameters in ecological models will change as adaptation occurs.\nC_LIO_LIOur approach unifies population ecology and evolutionary biology providing a framework allowing a very wide range of questions to be addressed. The next step is to apply the approach to a variety of laboratory and field systems. Once this is done we will have a much deeper understanding of eco-evolutionary dynamics and feedbacks.\nC_LI

Evolutionary Biology

A genetic test for differential causative pathology in disease subgroups

Many common diseases show wide phenotypic variation. We present a statistical method for determining whether phenotypically defined subgroups of disease cases represent different genetic architectures, in which disease-associated variants have different effect sizes in the two subgroups. Our method models the genome-wide distributions of genetic association statistics with mixture Gaussians. We apply a global test without requiring explicit identification of disease-associated variants, thus maximising power in comparison to a standard variant by variant subgroup analysis. Where evidence for genetic subgrouping is found, we present methods for post-hoc identification of the contributing genetic variants.\n\nWe demonstrate the method on a range of simulated and test datasets where expected results are already known. We investigate subgroups of type 1 diabetes (T1D) cases defined by autoantibody positivity, establishing evidence for differential genetic architecture with thyroid peroxidase antibody positivity, driven generally by variants in known T1D associated regions.

Genomics

The genetic architecture of local adaptation II: The QTL landscape of water-use efficiency for foxtail pine (Pinus balfouriana Grev. & Balf.)

Water availability is an important driver of the geographic distribution of many plant species, although its importance relative to other climatic variables varies across climate regimes and species. A common indirect measure of water-use efficiency (WUE) is the ratio of carbon isotopes ({delta}13C) fixed during photosynthesis, especially when analyzed in conjunction with a measure of leaf-level resource utilization ({delta}15N). Here, we test two hypotheses about the genetic architecture of WUE for foxtail pine (Pinus balfouriana Grev. & Balf.) using a novel mixture of double digest restriction site associated DNA sequencing, species distribution modeling, and quantitative genetics. First, we test the hypothesis that water availability is an important determinant of the geographical range of foxtail pine. Second, we test the hypothesis that variation in {delta}13C and {delta}15N is genetically based, differentiated between regional populations, and has genetic architectures that include loci of large effect. We show that precipitation-related variables structured the geographical range of foxtail pine, climate-based niches differed between regional populations, and {delta}13C and {delta}15N were heritable with moderate signals of differentiation between regional populations. A set of large-effect QTLs (n = 11 for {delta}13C; n = 10 for {delta}15N) underlying {delta}13C and {delta}15N variation, with little to no evidence of pleiotropy, was discovered using multiple-marker, half-sibling regression models. Our results represent a first approximation to the genetic architecture of these phenotypic traits, including documentation of several patterns consistent with {delta}13C being a fitness-related trait affected by natural selection.

Evolutionary Biology

Mixing of Porpoise Ecotypes in South Western UK Waters Revealed by Genetic Profiling

Contact zones between marine ecotypes are of interest for understanding how key pelagic predators may react to climate change. We analysed the fine scale genetic structure and morphological variation in harbour porpoises around the UK, at the proposed northern limit of a contact zone between southern and northern ecotypes in the Bay of Biscay. Using a sample of 591 stranded animals spanning a decade and microsatellite profiling at 9 loci, clustering and spatial analyses revealed that animals stranded around UK are composed of mixed genetic ancestries from two genetic pools. Porpoises from SW England displayed a distinct genetic ancestry, had larger body-sizes and inhabit an environment differentiated from other UK costal areas. Genetic ancestry blends from one group to the other along a SW-NE axis along the UK coastline, and showed a significant association with body size, consistent with morphological differences between the two ecotypes and their mixing around the SW coast. We also found significant isolation-by-distance among juveniles, suggesting that stranded juveniles display reduced intergenerational dispersal, while adults show larger variance. The fine scale structure of this admixture zone raises the question of how it will respond to future climate change and provides a reference point for further study.

Evolutionary Biology

Levels and patterns of genetic diversity differ between two closely related endemic Arabidopsis species

Theory predicts that a small effective population size leads to slower accumulation of mutations, increased levels of genetic drift and reduction in the efficiency of natural selection. Therefore endemic species should harbor low levels of genetic diversity and exhibit a reduced ability of adaptation to environmental changes. Arabidopsis pedemontana and Arabidopsis cebennensis, two endemic species from Italy and France respectively, provide an excellent model to study the adaptive potential of species with small distribution ranges. To evaluate the genome-wide levels and patterns of genetic variation, effective population size and demographic history of both species, we genotyped 53 A. pedemontana and 28 A. cebennensis individuals across the entire species ranges with Genotyping-by-Sequencing. SNPs data confirmed a low genetic diversity for A. pedemontana although its effective population size is relatively high. Only a weak population structure was observed over the small distribution range of A. pedemontana, resulting from an isolation-by-distance pattern of gene flow. In contrary, A. cebennensis individuals clustered in three populations according to their geographic distribution. Despite this and a larger distribution, the overall genetic diversity was even lower for A. cebennensis than for A. pedemontana. A demographic analysis demonstrated that both endemics have undergone a strong population size decline in the past, without recovery. The more drastic decline observed in A. cebennensis partially explains the very small effective population size observed in the present population. In light of these results, we discuss the adaptive potential of these endemic species in the context of rapid climate change.

Genomics

Predicting Drug Synergy and Antagonism from Genetic Interaction Neighborhoods

Although drug combinations have proven efficacious in a variety of diseases, the design of such regimens often involves extensive experimental screening due to the myriad choice of drugs and doses. To address these challenges, we utilize the budding yeast Saccharomyces cerevisiae as a model organism to evaluate whether drug synergy or antagonism is mediated through genetic interactions between their target genes. Specifically, we hypothesize that if the inhibition targets of one chemical compound are in close proximity to those of a second compound in a genetic interaction network, then the compound pair will exhibit synergy or antagonism. Graph metrics are employed to make precise the notion of proximity in a network. Knowledge of genetic interactions and small-molecule targets are compiled through literature sources and curated databases, with predictions validated according to experimentally determined gold standards. Finally, we test whether genetic interactions propagate through networks according to a \"guilt-by-association\" framework. Our results suggest that close proximity between the target genes of one drug and those of another drug does not strongly predict synergy or antagonism. In addition, we find that the extent to which the growth of a double gene mutant deviates from expectation is moderately anti-correlated with their distance in a genetic interaction network.

Systems Biology

LD Hub: a centralized database and web interface to perform LD score regression that maximizes the potential of summary level GWAS data for SNP heritability and genetic correlation analysis

MotivationLD score regression is a reliable and efficient method of using genome-wide association study (GWAS) summary-level results data to estimate the SNP heritability of complex traits and diseases, partition this heritability into functional categories, and estimate the genetic correlation between different phenotypes. Because the method relies on summary level results data, LD score regression is computationally tractable even for very large sample sizes. However, publicly available GWAS summary-level data are typically stored in different databases and have different formats, making it difficult to apply LD score regression to estimate genetic correlations across many different traits simultaneously.\n\nResultsIn this manuscript, we describe LD Hub - a centralized database of summary-level GWAS results for 177 diseases/traits from different publicly available resources/consortia and a web interface that automates the LD score regression analysis pipeline. To demonstrate functionality and validate our software, we replicated previously reported LD score regression analyses of 49 traits/diseases using LD Hub; and estimated SNP heritability and the genetic correlation across the different phenotypes. We also present new results obtained by uploading a recent atopic dermatitis GWAS meta-analysis to examine the genetic correlation between the condition and other potentially related traits. In response to the growing availability of publicly accessible GWAS summary-level results data, our database and the accompanying web interface will ensure maximal uptake of the LD score regression methodology, provide a useful database for the public dissemination of GWAS results, and provide a method for easily screening hundreds of traits for overlapping genetic aetiologies.\n\nAvailability and implementationThe web interface and instructions for using LD Hub are available at http://ldsc.broadinstitute.org/

Bioinformatics

Uplift and erosion of genomic islands with standing genetic variation

Details of the processes that generate biological diversity have long been of interest to evolutionary biologists. A common theme in nature is diversification via divergent selection with gene flow. Empirical studies on this topic find variable genetic differentiation throughout the genome, that genetic differentiation is non-randomly distributed, and that loci of adaptive significance are often found clustered together within \"genomic islands of divergence\". Theoretical models based on new mutations show how these genomic islands can arise and grow as a result of a complex interaction of various evolutionary and genic processes. In the current study, I ask if such genomic islands can alternatively arise from divergent selection from standing genetic variation and I tested this using a simple two locus model of selection. There are numerous ways in which standing genetic variation can be partitioned (e.g., between alleles, between loci, and between populations) and I tested which of these scenarios can give rise to an island pattern compared to no genomic differentiation or complete genomic differentiation. I found that divergent selection, even without reciprocal gene exchange between populations, following a bout of admixture can relatively quickly produce an island pattern. Moreover, I found two pathways in which islands can form from divergence from standing variation: 1) through the build up of islands and 2) through the breakdown of larger, genome-wide differentiation. Lastly, similar to new mutation theory, I found that the frequency of recombination is an important determinant of island formation from standing genetic variation such that mating behavior of a species (e.g., facultative or obligate sexual) can impact the likelihood of island formation.

Evolutionary Biology

Elucidating the genetic basis of an oligogenic birth defect using whole genome sequence data in a non-model organism, Bubalus bubalis

Recent strong selection for dairy traits in water buffalo has been associated with higher levels of inbreeding, leading to an increase in the prevalence of genetic diseases such as transverse hemimelia (TH), a congenital developmental abnormality characterized by the absence of a variable distal portion of the hindlimbs. The limited genomic resources available for water buffalo, in conjunction with an unconfirmed inheritance pattern, required an original approach to identify genetic variants associated with this disease. The genomes of 4 bilaterally affected cases, 7 unilaterally affected cases, and 14 controls were sequenced. Variant calling identified 19.8 million high confidence single nucleotide polymorphisms (SNPs) and 2.8 million insertions/deletions (INDELs). A concordance analysis of SNPs and INDELs requiring all unilateral and bilateral cases and none of the controls to be homozygous for the same allele, revealed two genes, WNT7A and SMARCA4, known to play a role in embryonic hindlimb development. Additionally, SNP alleles in NOTCH1 and RARB were homozygous exclusively in the bilaterally affected cases, suggesting an oligogenic mode of inheritance. Homozygosity mapping by whole genome de novo assembly was then used to identify large contigs representing regions of homozygosity in the cases. This also supported an oligogenic mode of inheritance; implicating 13 genes involved in aberrant hindlimb development in the bilateral cases and 11 in the unilateral cases. A genome-wide association study (GWAS) predicted additional modifier genes. Results from these analyses suggest that mutations in SMARCA4 and WNT7A are required for expression of TH, while several other loci including NOTCH1 act as modifiers and increase the severity of the disease phenotype. Although our data show that the inheritance of TH is complex, we predict that homozygous variants in WNT7A and SMARCA4 are necessary for the expression of TH and selection against these variants and avoidance of carrier-to-carrier matings should eradicate TH.\n\nAuthor SummaryGenetic diseases often occur and are spread through small populations under strong selection where rates of inbreeding can be significant. The use of a limited number of water buffalo males via artificial insemination for genetic improvement of milk and milk composition has increased the frequency of the genetic disease, transverse hemimelia (TH). Transverse hemimelia affected calves are normally developed except for malformation of one or both hindlimbs or both hindlimbs and one or both forelimbs. Little is known about the inheritance pattern of TH. We discovered genetic variants present in cases where both hindlimbs and one forelimb were affected, cases were both hindlimbs were affected, cases where only one hindlimb was affected, and in non-affected water buffalo that predict TH to be inherited as an oligogenic disease with two driver loci necessary for disease expression and several additional modifier genes that are responsible for the severity of the disease phenotype. We predict that selection against mutations in the two major loci and the avoidance of mating animals that are heterozygous for these mutations will eliminate TH from water buffalo.

Genomics

Quantitative Seq-LGS: Genome-Wide Identification of Genetic Drivers of Multiple Phenotypes in Malaria Parasites

Identifying the genetic determinants of phenotypes that impact on disease severity is of fundamental importance for the design of new interventions against malaria. Traditionally, such discovery has relied on labor-intensive approaches that require significant investments of time and resources. By combining Linkage Group Selection (LGS), quantitative whole genome population sequencing and a novel mathematical modeling approach (qSeq-LGS), we simultaneously identified multiple genes underlying two distinct phenotypes, identifying novel alleles for growth rate and strain specific immunity (SSI), while removing the need for traditionally required steps such as cloning, individual progeny phenotyping and marker generation. The detection of novel variants, verified by experimental phenotyping methods, demonstrates the remarkable potential of this approach for the identification of genes controlling selectable phenotypes in malaria and other apicomplexan parasites for which experimental genetic crosses are amenable.\n\nSignificance StatementThis paper describes a powerful and rapid approach to the discovery of genes underlying medically important phenotypes in malaria parasites. This is crucial for the design of new drug and vaccine interventions. The approach bypasses the most time-consuming steps required by traditional genetic linkage studies and combines Mendelian genetics, quantitative deep sequencing technologies, genome analysis and mathematical modeling. We demonstrate that the approach can simultaneously identify multigenic drivers of multiple phenotypes, thus allowing complex genotyping studies to be conducted concomitantly. This methodology will be particularly useful for discovering the genetic basis of medically important phenotypes such as drug resistance and virulence in malaria and other apicomplexan parasites, as well as potentially in any organism undergoing sexual recombination.

Genomics

Identifying Migrant Origins Using Genetics, Isotopes, and Habitat Suitability

O_LIIdentifying migratory connections across the annual cycle is important for studies of migrant ecology, evolution, and conservation. While recent studies have demonstrated the utility of high-resolution SNP-based genetic markers for identifying population-specific migratory patterns, the accuracy of this approach relative to other intrinsic tagging techniques has not yet been assessed.\nC_LIO_LIHere, using a straightforward application of Bayes' Rule, we develop a method for combining inferences from high-resolution genetic markers, stable isotopes, and habitat suitability models, to spatially infer the breeding origin of migrants captured anywhere along their migratory pathway. Using leave-one-out cross validation, we compare the accuracy of this combined approach with the accuracy attained using each source of data independently.\nC_LIO_LIOur results indicate that when each method is considered in isolation, the accuracy of genetic assignments far exceeded that of assignments based on stable isotopes or habitat suitability models. However, our joint assignment method consistently resulted in small, but informative increases in accuracy and did help to correct misassignments based on genetic data alone. We demonstrate the utility of the combined method by identifying previously undetectable patterns in the timing of migration in a North American migratory songbird, the Wilson's warbler.\nC_LIO_LIOverall, our results support the idea that while genetic data provides the most accurate method for tracking animals using intrinsic markers when each method is considered independently, there is value in combining all three methods. The resulting methods are provided as part of a new computationally-efficient R-package, GIAIH, allowing broad application of our statistical framework to other migratory animal systems.\nC_LI

ecology

Identifying genetic variants that affect viability in large cohorts

A number of open questions in human evolutionary genetics would become tractable if we were able to directly measure evolutionary fitness. As a step towards this goal, we developed a method to examine whether individual genetic variants, or sets of genetic variants, currently influence viability. The approach consists in testing whether the frequency of an allele varies across ages, accounting for variation in ancestry. We applied it to the Genetic Epidemiology Research on Aging (GERA) cohort and to the parents of participants in the UK Biobank. Across the genome, we find only a few common variants with large effects on age-specific mortality: tagging the APOE {varepsilon}4 allele and near CHRNA3. These results suggest that when large, even late onset effects are kept at low frequency by purifying selection. Testing viability effects of sets of genetic variants that jointly influence one of 42 traits, we detect a number of strong signals. In participants of the UK Biobank study of British ancestry, we find that variants that delay puberty timing are enriched in longer-lived parents (P~6x10-6 for fathers and P~2x10-3 for mothers), consistent with epidemiological studies. Similarly, in mothers, variants associated with later age at first birth are associated with a longer lifespan (P~1x10-3). Signals are also observed for variants influencing cholesterol levels, risk of coronary artery disease, body mass index, as well as risk of asthma. These signals exhibit consistent effects in the GERA cohort and among participants of the UK Biobank of non-British ancestry. Moreover, we see marked differences between males and females, most notably at the CHRNA3 locus, and variants associated with risk of coronary artery disease and cholesterol levels. Beyond our findings, the analysis serves as a proof of principle for how upcoming biomedical datasets can be used to learn about selection effects in contemporary humans.

evolutionary biology

A computational method for the investigation of multistable systems and its application to genetic switches

Genetic switches exhibit multistability, form the basis of epigenetic memory, and are found in natural decision making systems, such as cell fate determination in developmental pathways. Synthetic genetic switches can be used for recording the presence of different environmental signals, for changing phenotype using synthetic inputs and as building blocks for higher-level sequential logic circuits. Understanding how multistable switches can be constructed and how they function within larger biological systems is therefore key to synthetic biology. Here we present a new computational tool, called StabilityFinder, that takes advantage of sequential Monte Carlo methods to identify regions of parameter space capable of producing multistable behaviour, while handling uncertainty in biochemical rate constants and initial conditions. The algorithm works by clustering trajectories in phase space, and iteratively minimizing a distance metric. Here we examine a collection of models of genetic switches, ranging from the deterministic Gardner toggle switch to stochastic models containing different positive feedback connections. We uncover the design principles behind making bistable, tristable and quadristable switches, and find that rate of gene expression is a key parameter. We demonstrate the ability of the framework to examine more complex systems and examine the design principles of a three gene switch. Our framework allows us to relax the assumptions that are often used in genetic switch models and we show that more complex abstractions are still capable of multistable behaviour. Our results suggest many ways in which genetic switches can be enhanced and offer designs for the construction of novel switches. Our analysis also highlights subtle changes in correlation of experimentally tunable parameters that can lead to bifurcations in deterministic and stochastic systems. Overall we demonstrate that StabilityFinder will be a valuable tool in the future design and construction of novel gene networks.

synthetic biology

Genetic footprint of population fragmentation and contemporary collapse in a freshwater cetacean

Understanding demographic trends and patterns of gene flow in an endangered species is crucial for devising conservation strategies. Here, we examined the extent of population structure and recent evolution of the critically endangered Yangtze finless porpoise (Neophocaena asiaeorientalis asiaeorientalis). By analysing genetic variation at the mitochondrial and nuclear microsatellite loci for 148 individuals, we identified three populations along the Yangtze River, each one connected to a group of admixed ancestry. Each population displayed extremely low genetic diversity, consistent with extremely small effective size ([≤]92 individuals). Habitat degradation and distribution gaps correlated with highly asymmetric gene-flow that was inefficient in maintaining connectivity between populations. Genetic inferences of historical demography revealed that the populations in the Yangtze descended from a small number of founders colonizing the river from the sea during the last Ice Age. The colonization was followed by a rapid population split during the last millennium predating the Chinese Modern Economy Development. However, genetic diversity showed a clear footprint of population contraction over the last 50 years leaving only ~2% of the pre-collapsed size, consistent with the population collapses reported from field studies. This genetic perspective provides background information for devising mitigation strategies to prevent this species from extinction.

evolutionary biology

Sex matters in Massive Parallel Sequencing: Evidence for biases in genetic parameter estimation and investigation of sex determination systems

Using massively parallel sequencing data from two species with different life history traits -- American lobster (Homarus americanus) and Arctic Char (Salvelinus alpinus) -- we highlighted how an unbalanced sex ratio in the samples combined with a few sex-linked markers may lead to false interpretations of population structure and thus to potentially erroneous management recommendations. Multivariate analyses revealed two genetic clusters that separated males and females instead of showing the expected pattern of genetic differentiation among ecologically divergent (inshore vs. offshore in lobster) or geographically distant (east vs. west in Arctic Char) sampling locations. We created several subsamples artificially varying the sex ratio in the inshore/offshore and east/west groups, and then demonstrated that significant genetic differentiation could be observed despite panmixia for lobster, and that Fst values were overestimated for Arctic Char. This pattern was due to 12 and 94 sex-linked markers driving differentiation for lobster and Arctic Char, respectively. Removing sex-linked markers led to nonsignificant genetic structure (lobster) and a more accurate estimation of Fst (Arctic Char). We further characterized the putative functions of sex-linked markers. Given that only 9.6% of all marine/diadromous population genomic studies to date reported sex information, we urge researchers to collect and consider individual sex information. In summary, we argue that sex information is useful to (i) control sex ratio in sampling, (ii) overcome \"sex-ratio bias\" that can lead to spurious genetic differentiation signals and (iii) fill knowledge gaps regarding sex determining systems.

genomics