Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 883 records · Page 49Linked to original sources

The effects of training population design on genomic prediction accuracy in wheat

Genomic selection offers several routes for increasing genetic gain or efficiency of plant breeding programs. In various species of livestock there is empirical evidence of increased rates of genetic gain from the use of genomic selection to target different aspects of the breeders equation. Accurate predictions of genomic breeding value are central to this and the design of training sets is in turn central to achieving sufficient levels of accuracy. In summary, small numbers of close relatives and very large numbers of distant relatives are expected to enable accurate predictions.\n\nTo quantify the effect of some of the properties of training sets on the accuracy of genomic selection in crops we performed an extensive field-based winter wheat trial. In summary, this trial involved the construction of 44 F2:4 bi- and triparental populations, from which 2992 lines were grown on four field locations and yield was measured. For each line, genotype data were generated for 25,000 segregating single nucleotide polymorphism markers. The overall heritability of yield was estimated to 0.65, and estimates within individual families ranged between 0.10 and 0.85. Within cross genomic prediction accuracies of yield BLUEs were 0.125 - 0.127 using two different cross-validation approaches, and generally increased with training set size. Using related crosses in training and validation sets generally resulted in higher prediction accuracies than using unrelated crosses. The results of this study emphasize the importance of the training set design in relation to the genetic material to which the resulting prediction model is to be applied.

genetics

Highly Structured Homolog Pairing Reflects Functional Organization of the Drosophila Genome

Trans-homolog interactions encompass potent regulatory functions, which have been studied extensively in Drosophila, where homologs are paired in somatic cells and pairing-dependent gene regulation, or transvection, is well-documented. Nevertheless, the structure of pairing and whether its functional impact is genome-wide have eluded analysis. Accordingly, we generated a diploid cell line from divergent parents and applied haplotype-resolved Hi-C, discovering that homologs pair relatively precisely genome-wide in addition to establishing trans-homolog domains and compartments. We also elucidated the structure of pairing with unprecedented detail, documenting significant variation across the genome. In particular, we characterized two forms: tight pairing, consisting of contiguous small domains, and loose pairing, consisting of single larger domains. Strikingly, active genomic regions (A-type compartments, active chromatin, expressed genes) correlated with tight pairing, suggesting that pairing has a functional role genome-wide. Finally, using RNAi and haplotype-resolved Hi-C, we show that disruption of pairing-promoting factors results in global changes in pairing.\n\nOne Sentence SummaryHaplotype-resolved Hi-C reveals structures of homolog pairing and global implications for gene activity in hybrid PnM cells.

genetics

A computational framework for systematic exploration of biosynthetic diversity from large-scale genomic data

Genome mining has become a key technology to explore and exploit natural product diversity through the identification and analysis of biosynthetic gene clusters (BGCs). Initially, this was performed on a single-genome basis; currently, the process is being scaled up to large-scale mining of pan-genomes of entire genera, complete strain collections and metagenomic datasets from which thousands of bacterial genomes can be extracted at once. However, no bioinformatic framework is currently available for the effective analysis of datasets of this size and complexity. Here, we provide a streamlined computational workflow, tightly integrated with antiSMASH and MIBiG, that consists of two new software tools, BiG-SCAPE and CORASON. BiG-SCAPE facilitates rapid calculation and interactive visual exploration of BGC sequence similarity networks, grouping gene clusters at multiple hierarchical levels, and includes a glocal alignment mode that accurately groups both complete and fragmented BGCs. CORASON employs a phylogenomic approach to elucidate the detailed evolutionary relationships between gene clusters by computing high-resolution multi-locus phylogenies of all BGCs within and across gene cluster families (GCFs), and allows researchers to comprehensively identify all genomic contexts in which particular biosynthetic gene cassettes are found. We validate BiG-SCAPE by correlating its GCF output to metabolomic data across 403 actinobacterial strains. Furthermore, we demonstrate the discovery potential of the platform by using CORASON to comprehensively map the phylogenetic diversity of the large detoxin/rimosamide gene cluster clan, prioritizing three new detoxin families for subsequent characterization of six new analogs using isotopic labeling and analysis of tandem mass spectrometric data.

bioinformatics

Structure and genome ejection mechanism of Podoviridae phage P68 infecting Staphylococcus aureus

Phages infecting S. aureus have the potential to be used as therapeutics against antibiotic-resistant bacterial infections. However, there is limited information about the mechanism of genome delivery of phages that infect Gram-positive bacteria. Here we present the structures of S. aureus phage P68 in its native form, genome ejection intermediate, and empty particle. The P68 head contains seventy-two subunits of inner core protein, fifteen of which bind to and alter the structure of adjacent major capsid proteins and thus specify attachment sites for head fibers. Unlike in the previously studied phages, the head fibers of P68 enable its virion to position itself at the cell surface for genome delivery. P68 genome ejection is triggered by disruption of the interaction of one of the portal protein subunits with phage DNA. The inner core proteins are released together with the DNA and enable the translocation of phage genome across the bacterial membrane into the cytoplasm.

microbiology

Comparative genomics of clinical isolates of Pseudomonas aeruginosa from cystic fibrosis patients in Mexico

Pseudomonas aeruginosa (P. aeruginosa) is the primary pathogen responsible for morbidity and mortality in patients with cystic fibrosis (CF). Its genomic plasticity and constant selective pressure from antimicrobial treatments have favored the emergence of multidrug-resistant clones. This study conducted a comparative genomic analysis of 41 P. aeruginosa isolated from pediatric patients with CF in Mexico from 2015 to 2024, with the aim of characterizing their evolutionary dynamics, resistome, and virulome. Whole-genome sequencing (MGI, Illumina, and PacBio platforms) was used, with de novo assemblies performed using Unicycler v0.4.8 on the BV-BRC platform. The databases used for the resistome were CARD and NDARO, and for the virulome, VFDB. Phylogenetic reconstruction was based on core-genome alignments generated with Roary v3.13.0, with maximum likelihood reconstruction performed in IQ-TREE v2.1.2. The statistical significance of the segregation of resistance and virulence patterns was evaluated using PERMANOVA analysis. The results revealed a significant clonal prevalence of sequence types (ST) 307 and ST 167. Phylogenomic analysis grouped the isolates into three main clades; Clade 1 stood out for having the highest resistance gene load (mean of 75 genes/genome), establishing itself as the main reservoir of multidrug-resistant profiles. Genotype-phenotype concordance reached 65.5% overall, with high accuracy for aminoglycosides (87.8%) and fluoroquinolones (82.9%). Furthermore, virulome analysis identified 67 distinct patterns that were significantly segregated among the clades (PERMANOVA: R2=0.31, p=0.001). These findings demonstrate that the evolution of P. aeruginosa lineages in the pediatric clinical setting involves parallel and coordinated adaptations in both their resistance potential and their virulence arsenal. This study underscores the need to adopt a multidisciplinary approach to the clinical management of chronic P. aeruginosa infections in pediatric patients. The persistence of extensively drug-resistant (XDR) strains calls for the integration of genomic surveillance and functional diagnostics, as well as the search for therapeutic alternatives for the clinical management of patients with cystic fibrosis.

microbiology

Evolutionary replay of duplicate-gene retention across independent whole-genome duplications

Whole-genome duplications repeatedly expose ancestral gene lineages to the same broad evolutionary outcome-retention or loss of duplicated copies-but it remains unclear whether this history replays similarly across evolutionary scales. We placed duplicate retention in shared hierarchical orthologous-group coordinates and compared percentile ranks defined within each event-wide mapped universe. Three independent angiosperm whole-genome duplications showed reproducible replay (global rank effect T-replay = 0.210, bootstrap 95% confidence interval 0.172-0.248; permutation P = 1/100,001). A plant reference-panel score specified before target outcomes were examined predicted retention after the Apple/Pear duplication ({rho} = 0.169, n = 373). Deep transfer was heterogeneous: the teleost-genome-duplication estimate was positive but unresolved ({rho} = 0.107, n = 151, 95% confidence interval -0.050 to 0.260), whereas transfer to the ancient budding-yeast whole-genome duplication (yeast WGD) was supported ({rho} = 0.280, n = 186). Independently reconstructed animal outcomes also replayed between teleost and Stylommatophora duplications (r = 0.226, n = 146, P = 0.00326), although the effect remained below a prespecified strong-effect threshold. A strict plant-animal comparison was limited to 25 deeply one-to-one lineages and was unresolved (r = 0.033, 95% confidence interval -0.303 to 0.340). Thus, ancestral gene-lineage identity contributes reproducibly to duplicate retention after independent whole-genome duplications, but replay is structured by evolutionary lineage and modified by event-specific history rather than governed by one universal gene-fate ranking.

evolutionary biology

Early establishment acts as a selective filter shaping climate-associated genomic variation in European beech

Climate change is increasing drought and heat stress in European forests, raising concerns about the capacity of long-lived tree species to respond to rapidly changing environmental conditions. While local adaptation has been documented in many forest trees, it remains unclear whether newly established seedlings, which form the forests of the future, are able to persist and adapt to these new climatic conditions. Here, we investigated genomic differences between naturally regenerated seedlings and trees of European beech (Fagus sylvatica) across the three regions of the German Biodiversity Exploratories using low-coverage whole-genome sequencing (5x) of 1,032 individuals. Population structure was primarily driven by geographic region, whereas genetic diversity was similar across life stages. Despite this genome-wide similarity, we detected allele frequency shifts between trees and seedlings, concentrated in narrow genomic windows. These shifts were strongest in surviving seedlings, suggesting that environmental filtering during early establishment may contribute to shaping the genetic composition of regenerating populations. The strongest signals were observed within the Swabian Alb, where sampled seedlings were 2-years old and had experienced a longer period of potential filtering prior to sampling. Genotype - environment association analyses identified loci associated with climatic variables, and subsequent GO enrichment analyses of genes linked to these loci revealed significantly more enriched GO terms in seedlings than in trees, suggesting stronger environmental filtering by the current climate in seedlings. In particular, we found associations with maximum air temperature, relative humidity, soil moisture, and precipitation, affecting genes involved in stress responses, growth, metabolism, and developmental processes. Together, our results demonstrate that young cohorts of European beech differ genetically from trees and reveal genomic patterns consistent with life-stage-dependent environmental filtering. These findings suggest that the genetic composition of early life-stages is already altered by current environmental conditions, possibly contributing to adaptation to new climatic conditions.

ecology

Population Structure Analysis of Globally Diverse Bull Genomes

Since domestication, population bottlenecks, breed formation, and selective breeding have radically shaped the genealogy and genetics of Bos taurus. In turn, characterization of population structure among globally diverse bull genomes enables detailed assessment of genetic resources and origins. By analyzing 432 unrelated bull genomes from 13 breeds and 16 countries, we demonstrate genetic diversity and structural complexity among the global bull population. Importantly, we relaxed a strong assumption of discrete or admixed population, by adapting latent variable models for individual-specific allele frequencies that directly capture a wide range of complex structure from genome-wide genotypes. We identified a highly complex population structure that defies the conventional hypothesis based on discrete membership and contributes to pervasive genetic differentiation in bull genomes. As measured by magnitude of differentiation, selection pressure on SNPs within genes is substantially greater than that on intergenic regions. Additionally, broad regions of chromosome 6 harboring largest genetic differentiation suggest positive selection underlying population structure. We carried out gene set analysis using SNP annotations to identify enriched functional categories such as energy-related processes and multiple development stages. Our comprehensive analysis of bull population structure can support genetic management strategies that capture structural complexity and promote sustainable genetic breadth.

Genomics

Evidence of causal effect of major depression on alcohol dependence: Findings from the Psychiatric Genomics Consortium

BackgroundDespite established clinical associations among major depression (MD), alcohol dependence (AD), and alcohol consumption (AC), the nature of the causal relationship between them is not completely understood.\n\nMethodsThis study was conducted using genome-wide data from the Psychiatric Genomics Consortium (MD: 135,458 cases and 344,901 controls; AD: 10,206 cases and 28,480 controls) and UK Biobank (AC-Frequency: from \"daily or almost daily\" to \"never\", 438,308 individuals; AC-Quantity: total units of alcohol per week, 307,098 individuals). Linkage disequilibrium score regression and Mendelian Randomization (MR) analyses were applied to investigate shared genetic mechanisms (horizontal pleiotropy) and causal relationships (mediated pleiotropy) among these traits.\n\nOutcomesPositive genetic correlation was observed between MD and AD (rgMD-AD=+0.47, P=6.6x10-10). AC-Quantity showed positive genetic correlation with both AD (rgAD-AC-Quantity=+0.75, P=1.8x10-14) and MD (rgMD-AC-Quantity=+0.14, P=2.9x10-7), while there was negative correlation of AC-Frequency with MD (rgMD-AC-Frequency=-0.17, P=1.5x10-10) and a non-significant result with AD. MR analyses confirmed the presence of pleiotropy among these traits. However, the MD-AD results reflect a mediated-pleiotropy mechanism (i.e., causal relationship) with a causal role of MD on AD (beta=0.28, P=1.29x10-6) that does not appear to be biased by confounding such as horizontal pleiotropy. No evidence of reverse causation was observed as the AD genetic instrument did not show a causal effect on MD.\n\nInterpretationResults support a causal role for MD on AD based on genetic datasets including thousands of individuals. Understanding mechanisms underlying MD-AD comorbidity not only addresses important public health concerns but also has the potential to facilitate prevention and intervention efforts.\n\nFundingNational Institute of Mental Health and National Institute on Drug Abuse.\n\nPutting data into contextO_ST_ABSEvidence before this studyC_ST_ABSWe searched PubMed up to August 24, 2018, for research studies that investigated causality among alcohol-and depression related phenotypes using Mendelian randomization approaches. We used the search terms \"alcohol\" AND \"depression\" AND \"Mendelian Randomization\". No restrictions were applied to language, date, or article type. Ten articles were retrieved, but only two were focused on alcohol consumption and depression-related traits. The studies were based on genetic variants in alcohol dehydrogenase (ADH) genes only, did not find evidence for a causal effect of alcohol consumption on depression phenotypes, with one study finding a causal effect of alcohol consumption on alcoholism. Both studies noted that future studies are needed with increased sample sizes and clinically derived phenotypes. To our knowledge, no previous study has applied two-sample Mendelian randomization to investigate causal relationships between alcohol dependence and major depression.\n\nTwin studies show genetic factors influence susceptibility to MD, AD, and alcohol consumption. Differently from observational approaches where several studies have investigated the relationship between alcohol-and depression-related phenotypes, very limited use of molecular genetic data has been applied to investigate this issue. Additionally, the use of genetic information has been shown to be less biased by confounders and reverse causation than observation data. However, genetic approaches, like Mendelian randomization, require large sample sizes to be informative.\n\nAdded value of this studyIn this study, we used genome-wide data from the Psychiatric Genomic Consortium and UK Biobank, which include information regarding hundred thousands of individuals, to test the presence of shared genetic mechanisms and causal relationships among major depression, alcohol dependence, and alcohol consumption. The results support a causal influence of MD on AD, while alcohol consumption showed shared genetic mechanisms with respect to both major depression and alcohol dependence.\n\nImplications of all the available evidenceGiven the significant morbidity and mortality associated with MD, AD, and the comorbid condition, understanding mechanisms underlying these associations not only address important public health concerns but also has the potential to facilitate prevention and intervention efforts.

genetics

Natural CMT2 variation is associated with genome-wide methylation changes and temperature seasonality

As Arabidopsis thaliana has colonized a wide range of habitats across the world it is an attractive model for studying the genetic mechanisms underlying environmental adaptation. Here, we used public data from two collections of A. thaliana accessions to associate genetic variability at individual loci with differences in climates at the sampling sites. We use a novel method to screen the genome for plastic alleles that tolerate a broader climate range than the major allele. This approach reduces confounding with population structure and increases power compared to standard genome-wide association methods. Sixteen novel loci were found, including an association between Chromomethylase 2 (CMT2) and temperature seasonality where the genome-wide CHH methylation was different for the group of accessions carrying the plastic allele. Cmt2 mutants were shown to be more tolerant to heat-stress, suggesting genetic regulation of epigenetic modifications as a likely mechanism underlying natural adaptation to variable temperatures, potentially through differential allelic plasticity to temperature-stress.\n\nAUTHOR SUMMARYA central problem when studying adaptation to a new environment is the interplay between genetic variation and phenotypic plasticity. Arabidopsis thaliana has colonized a wide range of habitats across the world and it is therefore an attractive model for studying the genetic mechanisms underlying environmental adaptation. Here, we study two collections of A. thaliana accessions from across Eurasia to identify loci associated with differences in climates at the sampling sites. A new genome-wide association analysis method was developed to detect adaptive loci where the alleles tolerate different climate ranges. Sixteen novel such loci were found including a strong association between Chromomethylase 2 (CMT2) and temperature seasonality. The reference allele dominated in areas with less seasonal variability in temperature, and the alternative allele existed in both stable and variable regions. Our results thus link natural variation in CMT2 and epigenetic changes to temperature adaptation. We showed experimentally that plants with a defective CMT2 gene tolerate heat-stress better than plants with a functional gene. Together this strongly suggests a role for genetic regulation of epigenetic modifications in natural adaptation to temperature and illustrates the importance of re-analyses of existing data using new analytical methods to obtain deeper insights into the underlying biology from available data.

Genomics

Genome sequencing of the perciform fish Larimichthys crocea provides insights into stress adaptation

The large yellow croaker Larimichthys crocea (L. crocea) is one of the most economically important marine fish in China and East Asian countries. It also exhibits peculiar behavioral and physiological characteristics, especially sensitive to various environmental stresses, such as hypoxia and air exposure. These traits may render L. crocea a good model for investigating the response mechanisms to environmental stress. To understand the molecular and genetic mechanisms underlying the adaptation and response of L. crocea to environmental stress, we sequenced and assembled the genome of L. crocea using a bacterial artificial chromosome and whole-genome shotgun hierarchical strategy. The final genome assembly was 679 Mb, with a contig N50 of 63.11 kb and a scaffold N50 of 1.03 Mb, containing 25,401 protein-coding genes. Gene families underlying adaptive behaviours, such as vision-related crystallins, olfactory receptors, and auditory sense-related genes, were significantly expanded in the genome of L. crocea relative to those of other vertebrates. Transcriptome analyses of the hypoxia-exposed L. crocea brain revealed new aspects of neuro-endocrine-immune/metabolism regulatory networks that may help the fish to avoid cerebral inflammatory injury and maintain energy balance under hypoxia. Proteomics data demonstrate that skin mucus of the air-exposed L. crocea had a complex composition, with an unexpectedly high number of proteins (3,209), suggesting its multiple protective mechanisms involved in antioxidant functions, oxygen transport, immune defence, and osmotic and ionic regulation. Our results provide novel insights into the mechanisms of fish adaptation and response to hypoxia and air exposure.

Genomics

RNA-Seq analysis and annotation of a draft blueberry genome assembly identifies candidate genes involved in fruit ripening, biosynthesis of bioactive compounds, and stage-specific alternative splicing

BackgroundBlueberries are a rich source of antioxidants and other beneficial compounds that can protect against disease. Identifying genes involved in synthesis of bioactive compounds could enable breeding berry varieties with enhanced health benefits.\n\nResultsToward this end, we annotated a draft blueberry genome assembly using RNA-Seq data from five stages of berry fruit development and ripening. Genome-guided assembly of RNA-Seq read alignments combined with output from ab initio gene finders produced around 60,000 gene models, of which more than half were similar to proteins from other species, typically the grape Vitis vinifera. Comparison of gene models to the PlantCyc database of metabolic pathway enzymes identified candidate genes involved in synthesis of bioactive compounds, including bixin, an apocarotenoid with potential disease-fighting properties, and defense-related cyanogenic glycosides, which are toxic.\n\nCyanogenic glycoside (CG) biosynthetic enzymes were highly expressed in green fruit, and a candidate CG detoxification enzyme was up regulated during fruit ripening. Candidate genes for ethylene, anthocyanin, and 400 other biosynthetic pathways were also identified. Homology-based annotation using Blast2GO and InterPro assigned Gene Ontology terms to around 15,000 genes. RNA-Seq expression profiling showed that blueberry growth, maturation, and ripening involve dynamic gene expression changes, including coordinated up and down regulation of metabolic pathway enzymes and transcriptional regulators. Analysis of RNA-seq alignments identified developmentally regulated alternative splicing, promoter use, and 3 end formation.\n\nConclusionsWe report genome sequence, gene models, functional annotations, and RNA-Seq expression data that provide an important new resource enabling high throughput studies in blueberry. RNA-Seq data are freely available for visualization in Integrated Genome Browser, and analysis code is available from the git repository at http://bitbucket.org/lorainelab/blueberrygenome.

Genomics

Genomic DNA transposition induced by human PGBD5

Transposons are mobile genetic elements that are found in nearly all organisms, including humans. Mobilization of DNA transposons by transposase enzymes can cause genomic rearrangements, but our knowledge of human genes derived from transposases is limited. Here, we find that the protein encoded by human PGBD5, the most evolutionarily conserved transposable element-derived gene in chordates, can induce stereotypical cut-and-paste DNA transposition in human cells. Genomic integration activity of PGBD5 requires distinct aspartic acid residues in its transposase domain, and specific DNA sequences with inverted terminal repeats with similarity to piggyBac transposons. DNA transposition catalyzed by PGBD5 in human cells occurs genome-wide, with precise transposon excision and preference for insertion at TTAA sites. The apparent conservation of DNA transposition activity by PGBD5 raises the possibility that genomic remodeling may contribute to its biological function.

Genomics

Clinical metagenomic identification of Balamuthia mandrillaris encephalitis and assembly of the draft genome: the critical need for reference strain sequencing

Primary amoebic meningoencephalitis (PAM) is a rare, often lethal cause of encephalitis, for which early diagnosis and prompt initiation of combination antimicrobials may improve clinical outcomes. In this study, we present the first draft assembly of the Balamuthia mandrillaris genome recovered from a rare survivor of PAM, in total comprising 49 Mb of sequence. Comparative analysis of the mitochondrial genome and high-copy number genes from 6 additional Balamuthia mandrillaris strains demonstrated remarkable sequence variation, with the closest homologs corresponding to other amoebae, hydroids, algae, slime molds, and peat moss. We also describe the use of unbiased metagenomic next-generation sequencing (NGS) and SURPI bioinformatics analysis to diagnose an ultimately fatal case of Balamuthia mandrillaris encephalitis in a 15-year old girl. Real-time NGS testing of a hospital day 6 CSF sample detected Balamuthia on the basis of high-quality hits to 16S and 18S ribosomal RNA sequences present in the National Center for Biotechnology Information (NCBI) nt reference database. Retrospective analysis of a day 1 CSF sample revealed that more timely identification of Balamuthia by metagenomic NGS, potentially resulting in a better outcome, would have required availability of the complete genome sequence. These results underscore the diverse evolutionary origins underpinning this eukaryotic pathogen, and the critical importance of whole-genome reference sequences for microbial detection by NGS.

Genomics

Hybridization capture using RAD probes (hyRAD), a new tool for performing genomic analyses on museum collection specimens.

In the recent years, many protocols aimed at reproducibly sequencing reduced-genome subsets in non-model organisms have been published. Among them, RAD-sequencing is one of the most widely used. It relies on digesting DNA with specific restriction enzymes and performing size selection on the resulting fragments. Despite its acknowledged utility, this method is of limited use with degraded DNA samples, such as those isolated from museum specimens, as these samples are less likely to harbor fragments long enough to comprise two restriction sites making possible ligation of the adapter sequences (in the case of double-digest RAD) or performing size selection of the resulting fragments (in the case of single-digest RAD). Here, we address these limitations by presenting a novel method called hybridization RAD (hyRAD). In this approach, biotinylated RAD fragments, covering a random fraction of the genome, are used as baits for capturing homologous fragments from genomic shotgun sequencing libraries. This simple and cost-effective approach allows sequencing of orthologous loci even from highly degraded DNA samples, opening new avenues of research in the field of museum genomics. Not relying on the restriction site presence, it improves among-sample loci coverage. In a trial study, hyRAD allowed us to obtain a large set of orthologous loci from fresh and museum samples from a non-model butterfly species, with a high proportion of single nucleotide polymorphisms present in all eight analyzed specimens, including 58-year-old museum samples. The utility of the method was further validated using 49 museum and fresh samples of a Palearctic grasshopper species for which the spatial genetic structure was previously assessed using mtDNA amplicons. The application of the method is eventually discussed in a wider context. As it does not rely on the restriction site presence, it is therefore not sensitive to among-sample loci polymorphisms in the restriction sites that usually causes loci dropout. This should enable the application of hyRAD to analyses at broader evolutionary scales.

Genomics

Application of a dense genetic map for assessment of genomic responses to selection and inbreeding in Heliothis virescens.

Adaptation of pest species to laboratory conditions and selection for resistance to toxins in the laboratory are expected to cause inbreeding and genetic bottlenecks that reduce genetic variation. Heliothis virescens, a major cotton pest, has been colonized in the laboratory many times, and a few laboratory colonies have been selected for Bt resistance. We developed 350 bp Double-Digest Restriction-site Associated DNA-sequencing (ddRAD-seq) molecular markers to examine and compare changes in genetic variation associated with laboratory adaptation, artificial selection, and inbreeding in this non-model insect species. We found that allelic and nucleotide diversity declined dramatically in laboratory-reared H. virescens as compared with field-collected populations. The declines were primarily due to the loss of low frequency alleles present in field-collected H. virescens. A further, albeit modest decline in genetic diversity was observed in a Bt-selected population. The greatest decline was seen in H. virescens that were sib-mated for 10 generations, where more than 80% of loci were fixed for a single allele. To determine which regions of the genome were resistant to fixation in our sib-mated line, we generated a dense intraspecific linkage map containing 3 PCR-based, and 659 ddRAD-seq markers. Markers that retained polymorphism were observed in small clusters spread over multiple linkage groups, but this clustering was not statistically significant. Here, we confirmed and extended the general expectations for reduced genetic diversity in laboratory colonies, provided tools for further genomic analyses, and produced highly homozygous genomic DNA for future whole genome sequencing of H. virescens.

Genomics

Accelerating gene discovery by phenotyping whole-genome sequenced multi- mutation strains and using the sequence kernel association test (SKAT)

Forward genetic screens represent powerful, unbiased approaches to uncover novel components in any biological process. Such screens suffer from a major bottleneck, however, namely the cloning of corresponding genes causing the phenotypic variation. Reverse genetic screens have been employed as a way to circumvent this issue, but can often be limited in scope. Here we demonstrate an innovative approach to gene discovery. Using C. elegans as a model system, we used a whole-genome sequenced multi-mutation library, from the Million Mutation Project, together with the Sequence Kernel Association Test (SKAT), to rapidly screen for and identify genes associated with a phenotype of interest, namely defects in dye-filling of ciliated sensory neurons. Such anomalies in dye-filling are often associated with the disruption of cilia, organelles which in humans are implicated in sensory physiology (including vision, smell and hearing), development and disease. Beyond identifying several well characterised dye-filling genes, our approach uncovered three genes not previously linked to ciliated sensory neuron development or function. From these putative novel dye-filling genes, we confirmed the involvement of BGNT-1.1 in ciliated sensory neuron function and morphogenesis. BGNT-1.1 functions at the trans-Golgi network of sheath cells (glia) to influence dye-filling and cilium length, in a cell non-autonomous manner. Notably, BGNT-1.1 is the orthologue of human B3GNT1/B4GAT1, a glycosyltransferase associated with Walker-Warburg syndrome (WWS). WWS is a multigenic disorder characterised by muscular dystrophy as well as brain and eye anomalies. Together, our work unveils an effective and innovative approach to gene discovery, and provides the first evidence that B3GNT1-associated Walker-Warburg syndrome may be considered a ciliopathy.\n\nAuthor SummaryModel organisms are useful tools for uncovering new genes involved in a biological process via genetic screens. Such an approach is powerful, but suffers from drawbacks that can slow down gene discovery. In forward genetics screens, difficult-to-map phenotypes present daunting challenges, and whole-genome coverage can be equally challenging for reverse genetic screens where typically only a single genes function is assayed per strain. Here, we show a different approach which includes positive aspects of forward (high-coverage, randomly-induced mutations) and reverse genetics (prior knowledge of gene disruption) to accelerate gene discovery. We paired a whole-genome sequenced multi-mutation C. elegans library with a rare-variant associated test to rapidly identify genes associated with a phenotype of interest: defects in sensory neurons bearing sensory organelles called cilia, via a simple dye-filling assay to probe the form and function of these cells. We found two well characterised dye-filling genes and three genes, not previously linked to ciliated sensory neuron development or function, that were associated with dye-filling defects. We reveal that disruption of one of these (BGNT-1.1), whose human orthologue is associated with Walker-Warburg syndrome, results in abrogated uptake of dye and cilia length defects. We believe that our novel approach is useful for any organism with a small genome that can be quickly sequenced and where many mutant strains can be easily isolated and phenotyped, such as Drosophila and Arabidopsis.

Genomics

Fractality and Entropic Scaling in the Chromosomal Distribution of Conserved Noncoding Elements in the Human Genome

Conserved, ultraconserved and other classes of constrained non-coding elements (referred as CNEs) represent one of the mysteries of current comparative genomics. These elements are defined using various degrees of sequence similarity between organisms and several thresholds of minimal length and are often marked by extreme conservation that frequently exceeds the one observed for protein-coding sequences. We here explore the distribution of different classes of CNEs in entire chromosomes, in the human genome. We employ two complementary methodologies, the scaling of block entropy and box-counting, with the aim to assess fractal characteristics of different CNE datasets. Both approaches converge to the conclusion that well-developed fractality is characteristic of elements that are either marked by extreme conservation between two or more organisms or are of ancient origin, i.e. conserved between distant organisms across evolution. Given that CNEs are often clustered around genes, especially those that regulate developmental processes, we verify by appropriate gene masking that fractal-like patterns emerge irrespectively of whether elements found in proximity or inside genes are excluded or not. An evolutionary scenario is proposed, involving genomic events, that might account for fractal distribution of CNEs in the human genome as indicated through numerical simulations.

Genomics