Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genetics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

High throughput characterization of genetic effects on DNA:protein binding and gene transcription

Many variants associated with complex traits are in non-coding regions, and contribute to phenotypes by disrupting regulatory sequences. To characterize these variants, we developed a streamlined protocol for a high-throughput reporter assay, BiT-STARR-seq (Biallelic Targeted STARR-seq), that identifies allele-specific expression (ASE) while accounting for PCR duplicates through unique molecular identifiers. We tested 75,501 oligos (43,500 SNPs) and identified 2,720 SNPs with significant ASE (FDR 10%). To validate disruption of binding as one of the mechanisms underlying ASE, we developed a new high throughput allele specific binding assay for NFKB-p50. We identified 2,951 SNPs with allele-specific binding (ASB) (FDR 10%); 173 of these SNPs also had ASE (OR=1.97, p-value=0.0006). Of variants associated with complex traits, 1,531 resulted in ASE and 1,662 showed ASB. For example, we characterized that the Crohns disease risk variant for rs3810936 increases NFKB binding and results in altered gene expression.

genetics

Tasman-PCR: A genetic diagnostic assay for Tasmanian devil facial tumour diseases

Tasmanian devils have spawned two transmissible cancer clones, known as devil facial tumour 1 (DFT1) and devil facial tumour 2 (DFT2). DFT1 and DFT2 are transmitted between animals by the transfer of allogeneic contagious cancer cells by biting, and both cause facial tumours. DFT1 and DFT2 tumours are grossly indistinguishable, but can be differentiated using histopathology, cytogenetics or genotyping of polymorphic markers. However, standard diagnostic methods require specialist skills and equipment and entail long processing times. Here, we describe Tasman-PCR: a simple PCR-based diagnostic assay that distinguishes DFT1 and DFT2 by amplification of DNA spanning tumour-specific interchromosomal translocations. We demonstrate the high sensitivity and specificity of this assay by testing DNA from 557 tumours and 818 normal devils. A temporal-spatial screen confirmed the reported geographic ranges of DFT1 and DFT2 and did not provide evidence of additional DFT clones. DFT2 affects disproportionately more males than females, and devils can be co-infected with DFT1 and DFT2. Overall, we present a PCR-based assay that delivers rapid, accurate and high-throughput diagnosis of DFT1 and DFT2. This tool provides an additional resource for devil disease management and may assist with ongoing conservation efforts.

genetics

Genetic sequences are two-dimensional

In attempting to align divergent homologs of a conserved developmental enhancer, a flaw in the homology concept embedded in gapped alignment (GA) was discovered. To correct this flaw, we developed a methodological approach called maximal homology alignment (MHA). The goal of MHA is to rescue internal microparalogy of biological sequences rather than to insert a pattern of gaps (null characters), which transform homologous sequences into strings of uniform size (1-dimensional lengths). The core operation in MHA is the \"cinch\", whereby inferred tandem microparalogy is represented in multiple rows across the same span of alignment columns. Thus, MHAs have a second (vertical) paralogy dimension, which re-categorizes most indel mutations as replication slippage and attenuates the indel problem. Furthermore, internally-cinched, inferred microparalogy in a self-MHA can later be relaxed to restore uniformity to 2-dimensional widths in a multiple sequence alignment. This de-cinching operation is used as a first resort before artificial null characters are used. We implement MHA in a program called maximal, which is composed of a series of modules for cinching and cyclelizing divergent tandem repeats. In conclusion, we find that the MHA approach is of higher utility than GA in non-protein-coding regulatory sequences, which are unconstrained by codon-based reading frames and are enriched in dense microparalogical content.

genetics

Transcriptomic analysis reveals similarities in genetic activation of detoxification mechanisms resulting from imidacloprid and chlorothalonil exposure

The Colorado potato beetle, Leptinotarsa decemlineata (Say), is an agricultural pest of commercial potatoes in parts of North America, Europe, and Asia. Plant protection strategies within this geographic range employ a variety of pesticides to combat not only the insect, but also plant pathogens. Previous research has shown that field populations of Leptinotarsa decemlineata have a chronological history of resistance development to a suite of insecticides, including the Group 4A neonicotinoids. The aim of this study is to contextualize the transcriptomic response of Leptinotarsa decemlineata when exposed to the neonicotinoid insecticide imidacloprid, or the fungicides boscalid or chlorothalonil, in order to determine whether these compounds induce similar detoxification mechanisms. We found that chlorothalonil and imidacloprid induced similar patterns of transcript expression, including the up-regulation of a cytochrome p450 and a UDP-glucuronosyltransferase transcript, which are often associated with xenobiotic metabolism. Further, transcriptomic responses varied among individuals within the same treatment group, suggesting individual insects responses vary within a population and may cope with chemical stressors in a variety of manners. These results further our understanding of the mechanisms involved in insecticide resistance in Leptinotarsa decemlineata.\n\nAuthor Contribution StatementConceived and designed the experiments: JC, CB, RLG. Performed the experiments: JC, BSS. Analyzed the data: JC, BSS. Wrote the paper: JC, BSS, RLG.

genetics

A Rigorous Interlaboratory Examination of the Need to Confirm NGS-Detected Variants by an Orthogonal Method in Clinical Genetic Testing

Orthogonal confirmation of NGS-detected germline variants has been standard practice, although published studies have suggested that confirmation of the highest quality calls may not always be necessary. The key question is how laboratories can establish criteria that consistently identify those NGS calls that require confirmation. Most prior studies addressing this question have limitations: These studies are generally small, omit statistical justification, and explore limited aspects of the underlying data. The rigorous definition of criteria that separate high-accuracy NGS calls from those that may or may not be true remains a critical issue.\n\nWe analyzed five reference samples and over 80,000 patient specimens from two laboratories. We examined quality metrics for approximately 200,000 NGS calls with orthogonal data, including 1662 false positives. A classification algorithm used these data to identify a battery of criteria that flag 100% of false positives as requiring confirmation (CI lower bound: 98.5-99.8% depending on variant type) while minimizing the number of flagged true positives. These criteria identify false positives that the previously published criteria miss. Sampling analysis showed that smaller datasets resulted in less effective criteria.\n\nOur methodology for determining test and laboratory-specific criteria can be generalized into a practical approach that can be used by many laboratories to help reduce the cost and time burden of confirmation without impacting clinical accuracy.

genetics

A Fully-Adjusted Two-Stage Procedure for Rank Normalization in Genetic Association Studies

When testing genotype-phenotype associations using linear regression, departure of the trait distribution from normality can impact both Type I error rate control and statistical power, with worse consequences for rarer variants. While it has been shown that applying a rank-normalization transformation to trait values before testing may improve these statistical properties, the factor driving them is not the trait distribution itself, but its residual distribution after regression on both covariates and genotype. Because genotype is expected to have a small effect (if any) investigators now routinely use a two-stage method, in which they first regress the trait on covariates, obtain residuals, rank-normalize them, and then secondly use the rank-normalized residuals in association analysis with the genotypes. Potential confounding signals are assumed to be removed at the first stage, so in practice no further adjustment is done in the second stage. Here, we show that this widely-used approach can lead to tests with undesirable statistical properties, due to both a combination of a mis-specified mean-variance relationship, and remaining covariate associations between the rank-normalized residuals and genotypes. We demonstrate these properties theoretically, and also in applications to genome-wide and whole-genome sequencing association studies. We further propose and evaluate an alternative fully-adjusted two-stage approach that adjusts for covariates both when residuals are obtained, and in the subsequent association test. This method can reduce excess Type I errors and improve statistical power.

genetics

Identification of a genetic element required for spore killing in Neurospora

Meiotic drive elements like Spore killer-2 (Sk-2) in Neurospora are transmitted through sexual reproduction to the next generation in a biased manner. Sk-2 achieves this biased transmission through spore killing. Here, we identify rfk-1 as a gene required for the spore killing mechanism. The rfk-1 gene is associated with a 1,481 bp DNA interval (called AH36) near the right border of the 30 cM Sk-2 element, and its deletion eliminates the ability of Sk-2 to kill spores. The rfk-1 gene also appears to be sufficient for spore killing because its insertion into a non-Sk-2 isolate disrupts sexual reproduction after the initiation of meiosis. Although the complete rfk-1 transcript has yet to be defined, our data indicate that rfk-1 encodes a protein of at least 39 amino acids and that rfk-1 has evolved from a partial duplication of gene ncu07086. We also present evidence that rfk-1s location near the right border of Sk-2 is critical for the success of spore killing. Increasing the distance of rfk-1 from the right border of Sk-2 causes it to be inactivated by a genome defense process called meiotic silencing by unpaired DNA (MSUD), adding to accumulating evidence that MSUD exists, at least in part, to protect genomes from meiotic drive.

genetics

Genetic Association of The CYP7A1 Gene with Duck Lipid Traits

Cholesterol 7-hydroxylase (Cyp7a1) participates in lipid metabolism of liver, and its pathway involves catabolism of cholesterol to bile acids and excretion from the body. However, little is known about the effect of the polymorphisms of CYP7A1 gene on duck lipid traits. In the present study, seven novel synonymous mutations loci in exon 2 and exon 3 of CYP7A1 gene in Cherry Valley ducks were identified using PCR production direct sequencing. One novel SNP g.1033130 C>T was predicted in exon 2. Six novel SNPs g.1034076 C>T, g.1034334 G>A, g.1034373 G>A, g.1034448 T>C, g.1034541 C>G, and g.1034550 G>A were discovered in exon 3. Six haplotypes were detected using SHEsis online analysis software, and five loci (g.1034334G>A, g.1034373G>A, g.1034448T>C, g.1034541C>G, and g.1034550G>A) were in complete linkage disequilibrium, and as a block named Locus C3. By single SNP association analysis, we found that the g.1033130 C>T locus was significantly associated with IMF, AFP, TG, and TC (P<0.01 or P<0.05) respectively, the g.1034076 C>T locus was significantly associated with AFP (P<0.05), and the locus C3 was significantly associated with TCH (P<0.05). Sixteen dipoltypes were detected by the combination of haplotypes, and demonstrated strong association with IMF, AFP, TG, and TCH (P<0.01). Therefore, our data suggested that the seven SNPs of CYP7A1 gene are potential markers for lipid homeostasis, and may be used for early breeding and selection of duck.

genetics

First estimation of the scale of canonical 5’ splice site GT&gt;GC mutations generating wild-type transcripts and their medical genetic implications

It has long been known that canonical 5 splice site (5SS) GT>GC mutations may be compatible with normal splicing. However, to date, the true scale of canonical 5SS GT>GC mutations generating wild-type transcripts, both in the context of the frequency of such mutations and the level of wild-type transcripts generated from the mutation alleles, remain unknown. Herein, combining data derived from a meta-analysis of 45 informative disease-causing 5SS GT>GC mutations (from 42 genes) and a cell culture-based full-length gene splicing assay of 103 5SS GT>GC mutations (from 30 genes), we estimate that [~]15-18% of the canonical GT 5SSs are capable of generating between 1 and 84% normal transcripts as a consequence of the substitution of GT by GC. We further demonstrate that the canonical 5SSs whose substitutions of GT by GC generated normal transcripts show stronger complementarity to the 5 end of U1 snRNA than those sites whose substitutions of GT by GC did not lead to the generation of normal transcripts. We also observed a correlation between the generation of wild-type transcripts and a milder than expected clinical phenotype but found that none of the available splicing prediction tools were able to accurately predict the functional impact of 5SS GT>GC mutations. Our findings imply that 5SS GT>GC mutations may not invariably cause human disease but should also help to improve our understanding of the evolutionary processes that accompanied GT>GC subtype switching of U2-type introns in mammals.

genetics

Clinical and Functional Characterization of Melanocortin 4 Receptor genetic variants in African American and/or Hispanic children with severe early onset obesity.

ContextMutations in melanocortin receptor (MC4R) are the most frequent cause of monogenic obesity in children of European ancestry, but little is known about their prevalence in children from minority populations in the United States.\n\nObjectiveThis study aims to identify the prevalence of MC4R mutations in children with severe early onset obesity of African-American and/or Latina ancestry.\n\nDesign and SettingIndividuals were recruited from the weight management clinics at two hospitals and from the institutional biobank at a third hospital. Sequencing of the MC4R gene was performed by whole exome and/or Sanger sequencing. Functional testing was performed to establish the surface expression of the receptor and cAMP response to its cognate ligand -melanocyte stimulating hormone.\n\nParticipantsThree hundred and twelve children (1-18 years, 50% girls) with body mass index (BMI) > 120% of 95th percentile of CDC 2000 growth charts at an age < 6 years, with no known pathological cause of obesity were enrolled.\n\nResultsEight rare MC4R mutations (2.6%) were identified in this study (R7S, F202L (n=2), M215I, G252D, V253I, I269N, F284I), three of which have not been previously reported (M215I, G252D, F284I). The pathogenicity of the variants was confirmed either by prior literature reports, or by functional testing. There was no significant difference in the BMI or height trajectories of children with or without MC4R mutations in this cohort.\n\nConclusionsWhile the prevalence of MC4R mutations in this cohort was similar to that reported in obese children of European ancestry, some of the variants were novel.

genetics

A genome wide dosage suppressor network reveals genetic robustness and a novel mechanism for Huntington’s disease

Mutational robustness is the extent to which an organism has evolved to withstand the effects of deleterious mutations. We explored the extent of mutational robustness in the budding yeast by genome wide dosage suppressor analysis of 53 conditional lethal mutations in cell division cycle and RNA synthesis related genes, revealing 660 suppressor interactions of which 642 are novel. This collection has several distinctive features, including high co-occurrence of mutant-suppressor pairs within protein modules, highly correlated functions between the pairs, and higher diversity of functions among the co-suppressors than previously observed. Dosage suppression of essential genes encoding RNA polymerase subunits and chromosome cohesion complex suggest a surprising degree of functional plasticity of macromolecular complexes and the existence of degenerate pathways for circumventing potentially lethal mutations. The utility of dosage-suppressor networks is illustrated by the discovery of a novel connection between chromosome cohesion-condensation pathways involving homologous recombination, and Huntingtons disease.

Genomics

A Genetic Screen Identifies Two Novel Rice Cysteine-rich Receptor-like Kinases That Are Required for the Rice NH1-mediated Immune Response

Over-expression of rice NH1 (NH1ox), the ortholog of Arabidopsis NPR1, confers immunity to bacterial and fungal pathogens and induces the appearance of necrotic lesions due to activation of defense genes at the pre-flowering stage. This lesion-mimic phenotype can be enhanced by the application of benzothiadiazole (BTH). To identify genes regulating these responses, we screened a fast neutron-irradiated NH1ox rice population. We identified one mutant, called sn11 (suppressor of NH1-mediated lesion-mimic 1), which is impaired both in BTH-induced necrotic lesion formation and in the immune response. Using a comparative genome hybridization approach employing rice whole genome tiling array, we identified 11 genes associated with the sn11 phenotype. Transgenic analysis revealed that RNA interference of two of the genes, encoding previously uncharacterized cysteine-rich receptor-like kinases (CRK6 and CRK10), re-created the sn11 phenotype. Elevated expression of CRK10 using an inducible expression system resulted in enhanced immunity. Quantitative PCR revealed that BTH treatment and elevated levels of rice NH1 and its paralog NH3 induced expression of CRK10 and CRK6 RNA. These results indicate that CRK6 and CRK10 are required for the BTH-activated immune response mediated by NH1.

Plant Biology

Sequencing of the human IG light chain loci from a hydatidiform mole BAC library reveals locus-specific signatures of genetic diversity

Germline variation at immunoglobulin gene (IG) loci is critical for pathogen-mediated immunity, but establishing complete reference sequences in these regions is problematic because of segmental duplications and somatically rearranged source DNA. We sequenced BAC clones from the essentially haploid hydatidiform mole, CHM1, across the light chain IG loci, kappa (IGK) and lambda (IGL), creating single haplotype representations of these regions. The IGL haplotype is 1.25Mb of contiguous sequence with four novel V gene and one novel C gene alleles and an 11.9kbp insertion. The IGK haplotype consists of two 644kbp proximal and 466kbp distal contigs separated by a gap also present in the reference genome sequence. Our effort added an additional 49kbp of unique sequence extending into this gap. The IGK haplotype contains six novel V gene and one novel J gene alleles and a 16.7kbp region with increased sequence identity between the two IGK contigs, exhibiting signatures of interlocus gene conversion. Our data facilitated the first comparison of nucleotide diversity between the light and IG heavy (IGH) chain haplotypes within a single genome, revealing a three to six fold enrichment in the IGH locus, supporting the theory that the heavy chain may be more important in determining antigenic specificity.

Genomics

Reproductive isolation of hybrid populations driven by genetic incompatibilities

Despite its role in homogenizing populations, hybridization has also been proposed as a means to generate new species. The conceptual basis for this idea is that hybridization can result in novel phenotypes through recombination between the parental genomes, allowing a hybrid population to occupy ecological niches unavailable to parental species. A key feature of these models is that these novel phenotypes ecologically isolate hybrid populations from parental populations, precipitating speciation. Here we present an alternative model of the evolution of reproductive isolation in hybrid populations that occurs as a simple consequence of selection against incompatibilities. Unlike previous models, our model does not require small population sizes, the availability of new niches for hybrids or ecological or sexual selection on hybrid traits. We show that reproductive isolation between hybrids and parents evolves frequently and rapidly under this model, even in the presence of substantial ongoing migration with parental species and strong selection against hybrids. Our model predicts that multiple distinct hybrid species can emerge from replicate hybrid populations formed from the same parental species, potentially generating patterns of species diversity and relatedness that resemble an adaptive radiation.

Evolutionary Biology

Thinking too positive? Revisiting current methods of population-genetic selection inference

In the age of next-generation sequencing, the availability of increasing amounts and quality of data at decreasing cost ought to allow for a better understanding of how natural selection is shaping the genome than ever before. Yet, alternative forces such as demography and background selection obscure the footprints of positive selection that we would like to identify. Here, we illustrate recent developments in this area, and outline a roadmap for improved selection inference. We argue (1) that the development and obligatory use of advanced simulation tools is necessary for improved identification of selected loci, (2) that genomic information from multiple-time points will enhance the power of inference, and (3) that results from experimental evolution should be utilized to better inform population-genomic studies.

Genomics

Bridging psychology and genetics using large-scale spatial analysis of neuroimaging and neurogenetic data

Understanding how microscopic molecules give rise to complex cognitive processes is a major goal of the biological sciences. The countless hypothetical molecule-cognition relationships necessitate discovery-based techniques to guide scientists toward the most productive lines of investigation. To this end, we present a novel discovery tool that uses spatial patterns of neural gene expression from the Allen Brain Institute (ABI) and large-scale functional neuroimaging meta-analyses from the Neurosynth framework to bridge neurogenetic and neuroimaging data. We quantified the spatial similarity between over 20,000 genes from the ABI and 48 psychological topics derived from lexical analysis of neuroimaging articles, producing a comprehensive set of gene/cognition mappings that we term the Neurosynth-gene atlas. We demonstrate the ability to independently replicate known gene/cognition associations (e.g., between dopamine and reward), and subsequently use it to identify a range of novel associations between individual molecules or genes and complex psychological phenomena such as reward, memory and emotion. Our results complement existing discovery-based methods such as GWAS, and provide a novel means of generating hypotheses about the neurogenetic substrates of complex cognitive functions.

Neuroscience

The design and analysis of binary variable traits in common garden genetic experiments of highly fecund species to assess heritability

Many biologically important traits are binomially distributed, with their key phenotypes being presence or absence. Despite their prevalence, estimating the heritability of binomial traits presents both experimental and statistical challenges. Here we develop both an empirical and computational methodology for estimating the narrow-sense heritability of binary traits for highly fecund species. Our experimental approach controls for undesirable culturing effects, while minimizing culture numbers, increasing feasibility in the field. Our statistical approach accounts for known issues with model-selection by using a permutation test to calculate significance values and includes both fitting and power calculation methods. We illustrate our methodology by estimating the narrow-sense heritability for larval settlement, a key life-history trait, in the reef-building coral Orbicella faveolata. The experimental, statistical and computational methods, along with all of the data from this study, were deployed in the R package multiDimBio.

Evolutionary Biology

Estimating K in Genetic Mixture Models

A key quantity in the analysis of structured populations is the parameter K, which describes the number of subpopulations that make up the total population. Inference of K ideally proceeds via the model evidence, which is equivalent to the likelihood of the model. However, the evidence in favour of a particular value of K cannot usually be computed exactly, and instead programs such as SO_SCPLOWTRUCTUREC_SCPLOW make use of simple heuristic estimators to approximate this quantity. We show - using simulated data sets small enough that the true evidence can be computed exactly - that these simple heuristics often fail to estimate the true evidence, and that this can lead to incorrect conclusions about K. Our proposed solution is to use thermodynamic integration (TI) to estimate the model evidence. After outlining the TI methodology we demonstrate the effectiveness of this approach using a range of simulated data sets. We find that TI can be used to obtain estimates of the model evidence that are orders of magnitude more accurate and precise than those based on simple heuristics. Furthermore, estimates of K based on these values are found to be more reliable than those based on a suite of model comparison statistics. Our solution is implemented for models both with and without admixture in the software TO_SCPLOWRUEC_SCPLOWK.

Evolutionary Biology