Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Hallmarks of early sex-chromosome evolution in the dioecious plant Mercurialis annua revealed by de novo genome assembly, genetic mapping and transcriptome analysis

Suppressed recombination around a sex-determining locus allows divergence between homologous sex chromosomes and the functionality of their genes. Here, we reveal patterns of the earliest stages of sex-chromosome evolution in the diploid dioecious herb Mercurialis annua on the basis of cytological analysis, de novo genome assembly and annotation, genetic mapping, exome resequencing of natural populations, and transcriptome analysis. Both genetic mapping and exome resequencing of individuals across the species range independently identified the largest linkage group, LG1, as the sex chromosome. Although the sex chromosomes of M. annua are karyotypically homomorphic, we estimate that about a third of the Y chromosome has ceased recombining, a region containing 568 transcripts and spanning 22.3 cM in the corresponding female map. Patterns of gene expression hint at the possible role of sexually antagonistic selection in having favored suppressed recombination. In total, the genome assembly contained 34,105 expressed genes, of which 10,076 were assigned to linkage groups. There was limited evidence of Y-chromosome degeneration in terms of gene loss and pseudogenization, but sequence divergence between the X and Y copies of many sex-linked genes was higher than between M. annua and its dioecious sister species M. huetii with which it shares a sex-determining region. The Mendelian inheritance of sex in interspecific crosses, combined with the other observed pattern, suggest that the M. annua Y chromosome has at least two evolutionary strata: a small old stratum shared with M. huetii, and a more recent larger stratum that is probably unique to M. annua and that stopped recombining about one million years ago. Article summaryPlants that evolved separate sexes (dioecy) recently are ideal models for studying the early stages of sex-chromosome evolution. Here, we use karyological, whole genome and transcriptome data to characterize the homomorphic sex chromosomes of the annual dioecious plant Mercurialis annua. Our analysis reveals many typical hallmarks of dioecy and sex-chromosome evolution, including sex-biased gene expression and high X/Y sequence divergence, yet few premature stop codons in Y-linked genes and very little outright gene loss, despite 1/3 of the sex chromosome having ceased recombination in males. Our results confirm that the M. annua species complex is a fertile system for probing early stages in the evolution of sex chromosomes.

genomics

Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome

A primary goal of The Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) is to develop an African Diaspora Power Chip (ADPC), a genotyping array consisting of tagging SNPs, useful in comprehensively identifying African specific genetic variation. This array is designed based on the novel variation identified in 642 CAAPA samples of African ancestry with high coverage whole genome sequence data (~30x depth). This novel variation extends the pattern of variation catalogued in the 1000 Genomes and Exome Sequencing Projects to a spectrum of populations representing the wide range of West African genomic diversity. These individuals from CAAPA also comprise a large swath of the African Diaspora population and incorporate historical genetic diversity covering nearly the entire Atlantic coast of the Americas. Here we show the results of designing and producing such a microchip array. This novel array covers African specific variation far better than other commercially available arrays, and will enable better GWAS analyses for researchers with individuals of African descent in their study populations. A recent study1 cataloging variation in continental African populations suggests this type of African-specific genotyping array is both necessary and valuable for facilitating large-scale GWAS in populations of African ancestry.

genomics

Systematic identification of pleiotropic genes from genetic interactions

Modular structures in biological networks are ubiquitous and well-described, yet this organization does not capture the complexity of genes individually influencing many modules. Pleiotropy, the phenomenon of a single genetic locus with multiple phenotypic effects, has previously been measured according to many definitions, which typically count phenotypes associated with genes. We take the perspective that, because genes work in complex and interconnected modules, pleiotropy can be treated as a network-derived characteristic. Here, we use the complete network of yeast genetic interactions (GI) to measure pleiotropy of nearly 2700 essential and nonessential genes. Our method uses frequent item set mining to discover GI modules, annotates them with high-level processes, and uses entropy to measure the functional spread of each genes set of containing modules. We classify genes whose modules indicate broad functional influence as having high pleiotropy, while genes with focused functional influence have low pleiotropy. These pleiotropy classes differed in a number of ways: high-pleiotropy genes have comparatively higher expression variance, higher protein abundance, more domains, and higher copy number, while low pleiotropy genes are more likely to be in protein complexes and have many curated phenotypes. We discuss the implications of these results regarding the nature and evolution of pleiotropy.

genomics

Genetic Diversity in Circulating Tumor Cell Clusters

Genetic diversity plays a central role in tumor progression, metastasis, and resistance to treatment. Experiments are shedding light on this diversity at ever finer scales, but interpretation is challenging. Using recent progress in numerical models, we simulate macroscopic tumors to investigate the interplay between growth dynamics, microscopic composition, and circulating tumor cell cluster diversity. We find that modest differences in growth parameters can profoundly change microscopic diversity. Simple outwards expansion leads to spatially segregated clones and low diversity, as expected. However, a modest cell turnover can result in an increased number of divisions and mixing among clones resulting in increased microscopic diversity in the tumor core. Using simulations to estimate power to detect such spatial trends, we find that multiregion sequencing data from contemporary studies is marginally powered to detect the predicted effects. Slightly larger samples, improved detection of rare variants, or sequencing of smaller biopsies or circulating tumor cell clusters would allow one to distinguish between leading models of tumor evolution. The genetic composition of circulating tumor cell clusters, which can be obtained from non-invasive blood draws, is therefore informative about tumor evolution and its metastatic potential.\n\nHighlightsO_LINumerical and theoretical models show interaction of front expansion, mutation, and clonal mixing in shaping tumor heterogeneity.\nC_LIO_LICell turnover increases intratumor heterogeneity.\nC_LIO_LISimulated circulating tumor cell clusters and microbiopsies exhibit substantial diversity with strong spatial trends.\nC_LIO_LISimulations suggest attainable sampling schemes able to distinguish between prevalent tumor growth models.\nC_LI

cancer biology

Extensive hidden genetic variation shapes the structure of functional elements in Drosophila

Mutations that add, subtract, rearrange, or otherwise refashion genome structure often affect phenotypes, though the fragmented nature of most contemporary assemblies obscure them. To discover such mutations, we assembled the first reference quality genome of Drosophila melanogaster since its initial sequencing. By comparing this genome to the existing D. melanogaster assembly, we create a structural variant map of unprecedented resolution, revealing extensive genetic variation that has remained hidden until now. Many of these variants constitute strong candidates underlying phenotypic variation, including tandem duplications and a transposable element insertion that dramatically amplifies the expression of detoxification genes associated with nicotine resistance. The abundance of important genetic variation that still evades discovery highlights how crucial high quality references are to deciphering phenotypes.

genomics

Orienting The Causal Relationship Between Imprecisely Measured Traits Using Genetic Instruments

Inference of the causal structure that induces correlations between two traits can be achieved by combining genetic associations with a mediation-based approach, as is done in the causal inference test (CIT) and others. However, we show that measurement error in the phenotypes can lead to mediation-based approaches inferring the wrong causal direction, and that increasing sample sizes has the adverse effect of increasing confidence in the wrong answer. Here we introduce an extension to Mendelian randomisation, a method that uses genetic associations in an instrumentation framework, that enables inference of the causal direction between traits, with some advantages. First, it is less susceptible to bias in the presence of measurement error; second, it is more statistically efficient; third, it can be performed using only summary level data from genome-wide association studies; and fourth, its sensitivity to measurement error can be evaluated. We apply the method to infer the causal direction between DNA methylation and gene expression levels. Our results demonstrate that, in general, DNA methylation is more likely to be the causal factor, but this result is highly susceptible to bias induced by systematic differences in measurement error between the platforms. We emphasise that, where possible, implementing MR and appropriate sensitivity analyses alongside other approaches such as CIT is important to triangulate reliable conclusions about causality.

systems biology

Patterns of cancer somatic mutations predict genes involved in phenotypic abnormalities and genetic diseases

Genomic sequence mutations in both the germline and somatic cells can be pathogenic. Several authors have observed that often the same genes are involved in cancer when mutated in somatic cells and in genetic diseases when mutated in the germline. Recent advances in high-throughput sequencing techniques have provided us with large databases of both types of mutations, allowing us to investigate this issue in a systematic way. Here we show that high-throughput data about the frequency of somatic mutations in the most common cancers can be used to predict the genes involved in abnormal phenotypes and diseases. The predictive power of somatic mutation patterns is largely independent of that of methods based on germline mutation frequency, so that they can be fruitfully integrated into algorithms for the prioritization of causal variants. Our results confirm the deep relationship between pathogenic mutations in somatic and germline cells, provide new insight into the common origin of cancer and genetic diseases and can be used to improve the identification of new disease genes.

genomics

Genetic costs of domestication and improvement

The cost of domestication hypothesis posits that the process of domesticating wild species can result in an increase in the number, frequency, and/or proportion of deleterious genetic variants that are fixed or segregating in the genomes of domesticated species. This cost may limit the efficacy of selection and thus reduce genetic gains in breeding programs for these species. Understanding when and how deleterious mutations accumulate can also provide insight into fundamental questions about the interplay of demography and selection. Here we describe the evolutionary processes that may contribute to deleterious variation accrued during domestication and improvement, and review the available evidence for the cost of domestication in animal and plant genomes. We identify gaps and explore opportunities in this emerging field, and finally offer suggestions for researchers and breeders interested in understanding or avoiding the consequences of an increased number or frequency of deleterious variants in domesticated species.

evolutionary biology

Loss And Gain Of Function Experiments Implicate TMEM18 As A Mediator Of The Strong Association Between Genetic Variants At Human Chromosome 2p25.3 And Obesity

An intergenic region of human Chromosome 2 (2p25.3) harbours genetic variants which are among those most strongly and reproducibly associated with obesity. The molecular mechanisms mediating these effects remain entirely unknown. The gene closest to these variants is TMEM18, encoding a transmembrane protein localised to the nuclear membrane. The expression of Tmem18 within the murine hypothalamic paraventricular nucleus was altered by changes in nutritional state, with no significant change seen in three other closest genes. Germline loss of Tmem18 in mice resulted in increased body weight, which was exacerbated by high fat diet and driven by increased food intake. Selective overexpression of Tmem18 in the PVN of wild-type mice reduced food intake and also increased energy expenditure. We confirmed the nuclear membrane localisation of TMEM18 but provide new evidence that it is has four, not three, transmembrane domains and that it physically interacts with key components of the nuclear pore complex. Our data support the hypothesis that TMEM18 itself, acting within the central nervous system, is a plausible mediator of the impact of adjacent genetic variation on human adiposity.

physiology

Association Between Genetically Elevated Levels Of Inflammatory Biomarkers And Risk Of Schizophrenia: A Two-Sample Mendelian Randomisation Study

BackgroundPositive associations between inflammatory biomarkers and risk of psychiatric disorders, including schizophrenia, have been reported in observational studies. However, conventional observational studies are prone to bias such as reverse causation and residual confounding.\n\nMethodsIn this study, we used summary data to evaluate the association of genetically elevated C reactive protein (CRP), interleukin-1 receptor antagonist (IL-1Ra) and soluble interleukin-6 receptor (IL-6R) levels with schizophrenia in a two-sample Mendelian randomisation design.\n\nResultsThe pooled odds ratio estimate using 18 CRP genetic instruments was 0.90 (95% CI: 0.84; 0.97) per two-fold increment in CRP levels; consistent results were obtained using different Mendelian randomisation methods and a more conservative set of instruments. The odds ratio for soluble IL-6R was 1.06 (95% CI: 1.01; 1.12) per two-fold increment. Estimates for IL-1Ra were inconsistent among instruments and pooled estimates were imprecise and centred on the null.\n\nConclusionUnder Mendelian randomisation assumptions, our findings suggest a protective causal effect of CRP and a risk-increasing causal effect of soluble IL-6R (potentially mediated at least in part by CRP) on schizophrenia risk.

epidemiology

The genetic legacy of Zoroastrianism in Iran and India: Insights into population structure, gene flow and selection.

Zoroastrianism is one of the oldest extant religions in the world, originating in Persia (present-day Iran) during the second millennium BCE. Historical records indicate that migrants from Persia brought Zoroastrianism to India, but there is debate over the timing of these migrations. Here we present novel genome-wide autosomal, Y-chromosome and mitochondrial data from Iranian and Indian Zoroastrians and neighbouring modern-day Indian and Iranian populations to conduct the first genome-wide genetic analysis in these groups. Using powerful haplotype-based techniques, we show that Zoroastrians in Iran and India show increased genetic homogeneity relative to other sampled groups in their respective countries, consistent with their current practices of endogamy. Despite this, we show that Indian Zoroastrians (Parsis) intermixed with local groups sometime after their arrival in India, dating this mixture to 690-1390 CE and providing strong evidence that the migrating group was largely comprised of Zoroastrian males. By exploiting the rich information in DNA from ancient human remains, we also highlight admixture in the ancestors of Iranian Zoroastrians dated to 570 BCE-746 CE, older than admixture seen in any other sampled Iranian group, consistent with a long-standing isolation of Zoroastrians from outside groups. Finally, we report genomic regions showing signatures of positive selection in present-day Zoroastrians that might correlate to the prevalence of particular diseases amongst these communities.

genomics

Fine-mapping of genetic loci driving spontaneous clearance of hepatitis C virus infection

Approximately three quarters of acute HCV infections evolve to a chronic state, while one quarter are spontaneously cleared. Genetic predispositions strongly contribute to the development of chronicity. We have conducted a genome-wide association study to identify genomic variants underlying HCV spontaneous clearance using Immunochip in European and African ancestries. We confirmed two previously reported significant associations, in the IL28B/IFNL41,2 and MHC regions, with spontaneous clearance in the European population. We further fine-mapped the MHC association to a region of about 50 kilo base pairs, down from 1 mega base pairs in the previous study. Additional analyses suggested that the association in the MHC locus might be significantly stronger for virus subtype 1a than 1b, suggesting that viral subtype may have influenced the genetic mechanism underlying the clearance of HCV.

immunology

Estimating Wildlife Vaccination Coverage Using Genetic Methods

Vaccination is a potentially useful approach for the control of disease in wildlife populations. The effectiveness of vaccination is contingent in part on obtaining adequate vaccine coverage at the population level. However, measuring vaccine coverage in wild animal populations is challenging and so there is a need to develop robust approaches to estimate coverage and so contribute to understanding the likely efficacy of vaccination.\n\nWe used a modified capture mark recapture technique to estimate vaccine coverage in a wild population of European badgers (Meles meles) vaccinated by live-trapping and injecting with Bacillus Calmette-Guerin as part of a bovine tuberculosis control initiative in Wales, United Kingdom. Our approach used genetic matching of vaccinated animals to a sample of the wider population to estimate the percentage of badgers that had been vaccinated. Individual-specific genetic profiles were obtained using microsatellite genotyping of hair samples which were collected both directly from trapped and vaccinated badgers and non-invasively from the wider population using hair traps deployed at badger burrows.\n\nWe estimated the percentage of badgers vaccinated in a single year and applied this to a simple model to estimate cumulative vaccine coverage over a four year period, corresponding to the total duration of the vaccination campaign.\n\nIn the year of study, we estimated that between 44-65% (95% confidence interval, mean 55%) of the badger population received a vaccine dose. Using the model, we estimated that 70-85% of the total population would have received at least one vaccine dose over the course of the four year vaccination campaign.\n\nThis study represents the first application of this novel approach for measuring vaccine coverage in wildlife. This is also the first attempt at quantifying the level of vaccine coverage achieved by trapping and injecting badgers. The results therefore have specific application to bovine tuberculosis control policy, and the approach is of significance to the wider field of wildlife vaccination.

ecology

Integrated Computing And Tracking System For Centralized High-Throughput Genetic Analysis: A Case Study

The Genetic Analysis Center (GAC) of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) developed an Integrated Computing and Tracking system (ICT) in order to perform genome-wide and other genetic association studies automatically and efficiently, while documenting all analysis specifications. This system provides easy-to-use analysis set-up and computing procedures, automatic reports, and analysis search functionality due to integration with an on-site database. In this paper we describe the ICT and demonstrate how it satisfies key principles of reproducible research, while respecting constraints and challenges arising from using very large, restricted access, human-subjects data. This case study may benefit other groups that have similar requirements for high-throughput analysis execution and management.

bioinformatics

Pair Matcher (PaM): Fast Model-Based Optimisation Of Treatment/Case-Control Matches Using Demographic And Genetic Data

In clinical trials, individuals are matched for demographic criteria, paired, and then randomly assigned to treatment and control groups to determine a drugs efficacy. The successful completion of pilot trials is a prerequisite to larger and more expensive Phase III trials. One of the chief causes for the irreproducibility of results across pilot to Phase III trials is population stratification bias caused by the uneven distribution of ancestries in the treatment and control groups. Pair Matcher (PaM) addresses stratification bias by optimising pairing assignments a priori- and\\or posteriori to the trial using both genetic and demographic criteria. Using simulated and real datasets, we show that PaM identifies ideal and near-ideal pairs that are more genetically homogeneous than those identified based on racial criteria or Principal Component Analysis (PCA) alone. Homogenising the treatment (or case) and control groups can be expected to improve the accuracy and reproducibility of the study. PaMs ability to infer the ancestry of the participants further allows identifying subgroup of responders and developing a precision medicine approach to treatment. PaM is simple to execute, fast, and can be used for clinical trials and association studies. PaM is freely available via R scripts and a web interface.

bioinformatics

Evidence for "inter- and intraspecific horizontal genetic transfers" between anciently asexual bdelloid rotifers is explained by cross-contamination

Bdelloid rotifers are microscopic invertebrates thought to have evolved for millions of years without sexual reproduction. They have attracted the attention of biologists puzzled by the maintenance of sex among nearly all other eukaryotes. Bdelloid genomes have an unusually high proportion of horizontally acquired non-metazoan genes. This well-substantiated finding has invited speculation that homologous horizontal transfer between rotifers also may occur, perhaps even 'replacing' sex. A 2016 study in Current Biology claimed to supply evidence for this hypothesis. The authors sampled rotifers of the genus Adineta from natural populations and sequenced one mitochondrial and four nuclear loci. For several samples, species assignments were incongruent among loci, which the authors interpreted as evidence of \"interspecific genetic exchanges\". Here, we use sequencing chromatograms supplied by the authors to demonstrate that samples treated as individuals actually contained two or more divergent mitochondrial and ribosomal sequences, indicating contamination with DNA from additional animals belonging to the supposed \"donor species\". We also show that \"exchanged\" molecules share only 75% sequence homology, a degree of divergence incompatible with established mechanisms of recombination and genomic features of Adineta. These findings are parsimoniously explained by cross-contamination of tubes with animals or DNA from different species. Given the proportion of tubes contaminated in this way, we show by calculation that evidence for \"intraspecific horizontal exchange\" in the same dataset is explained by contamination with conspecific DNA. On the clear evidence of these analyses, the 2016 study provides no reliable support for the hypothesis of horizontal genetic transfer between or within these bdelloid species.

evolutionary biology

The genetic overlap between mood disorders and cardio-metabolic diseases: A systematic review of genome wide and candidate gene studies

Meta-analyses of genome-wide association studies (meta-GWAS) and candidate gene studies have identified genetic variants associated with cardiovascular diseases, metabolic diseases, and mood disorders. Although previous efforts were successful for individual disease conditions (single disease), limited information exists on shared genetic risk between these disorders. This article presents a detailed review and analysis of cardio-metabolic diseases risk (CMD-R) genes that are also associated with mood disorders. Firstly, we reviewed meta-GWA studies published until January 2016, for the diseases \"type 2 diabetes, coronary artery disease, hypertension\" and/or for the risk factors \"blood pressure, obesity, plasma lipid levels, insulin and glucose related traits\". We then searched the literature for published associations of these CMD-R genes with mood disorders. We considered studies that reported a significant association of at least one of the CMD-R genes and \"depressive disorder\" OR \"depressive symptoms\" OR \"bipolar disorder\" OR \"lithium treatment\", OR \"serotonin reuptake inhibitors treatment\". Our review revealed 24 potential pleiotropic genes that are likely to be shared between mood disorders and CMD-Rs. These genes include MTHFR, CACNA1D, CACNB2, GNAS, ADRB1, NCAN, REST, FTO, POMC, BDNF, CREB, ITIH4, LEP, GSK3B, SLC18A1, TLR4, PPP1R1B, APOE, CRY2, HTR1A, ADRA2A, TCF7L2, MTNR1B, and IGF1. A pathway analysis of these genes revealed significant pathways: corticotrophin-releasing hormone signaling, AMPK signaling, cAMP-mediated or G-protein coupled receptor signaling, axonal guidance signaling, serotonin and dopamine receptors signaling, dopamine-DARPP32 feedback in cAMP signaling, circadian rhythm signaling and leptin signaling. Our findings provide insights in to the shared biological mechanisms of mood disorders and cardio-metabolic diseases.

genomics

Genetic context and proliferation determine contrasting patterns of copy number amplification and loss for 5S and 45S ribosomal DNA arrays in cancer

The multicopy 45S ribosomal DNA (45S rDNA) array gives origin to the nucleolus, the first discovered nuclear organelle, site of Poll I 45S rRNA transcription and key regulator of cellular metabolism, DNA repair response, genome stability, and global epigenetic states. The multicopy 5S ribosomal DNA array (5S rDNA) is located on a separate chromosome, encodes the 5S rRNA transcribed by Pol III, and exhibits concerted copy number variation (cCNV) with the 45S rDNA array in human blood. Here we combined genomic data from >700 tumors and normal tissues to provide a portrait of ribosomal DNA variation in human tissues and cancers of diverse mutational signatures. We show that most cancers undergo coupled 5S rDNA array amplification and 45S rDNA loss, with abundant inter-individual variation in rDNA copy number of both arrays, and concerted modulation of 5S-45S copy number in some but not all tissues. Analysis of genetic context revealed associations between the presence of specific somatic alterations, such as P53 mutations in stomach and lung adenocarcinomas, and coupled 5S gain / 45S loss. Finally, we show that increased proliferation rates along cancer lineages can partially explain contrasting copy number changes in the 5S and 45S rDNA arrays. We suggest that 5S rDNA amplification facilitates increased ribosomal synthesis in cancer, whereas 45S rDNA loss emerges as a byproduct of transcription-replication conflict in highly proliferating tumor cells. Our results highlight the tissue- specificity of concerted copy number variation and uncover contrasting changes in 5S and 45S rDNA copy number along rapidly proliferating cell lineages.\n\nLay SummaryThe 45S and 5S ribosomal DNA (rDNA) arrays contain hundreds of rDNA copies, with substantial variability across individuals and species. Although physically unlinked, both arrays exhibit concerted copy number variation. However, whether concerted copy number is universally observed across all tissues is unknown. It also remains unknown if rDNA copy number may vary in tissues and cancer lineages. Here we showed that most cancers undergo coupled 5S rDNA array amplification and 45S rDNA loss, and concerted 5S-45S copy number variation in some but not all tissues. The coupled 5S amplification and 45S loss is associated with the presence of certain somatic genetic variations, as well as increased cancerous cell proliferation rate. Our research highlights the tissue- specificity of concerted copy number variation and uncover contrasting changes in rDNA copy number along rapidly proliferating cell lineages. Our observations raise the prospects of using 5S and 45S ribosomal DNA states as indicators of cancer status and targets in new strategies for cancer therapy.

genomics