Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,531 records · Page 85Linked to original sources

Reverse-complement parameter sharing improves deep learning models for genomics

Deep learning approaches that have produced breakthrough predictive models in computer vision, speech recognition and machine translation are now being successfully applied to problems in regulatory genomics. However, deep learning architectures used thus far in genomics are often directly ported from computer vision and natural language processing applications with few, if any, domain-specific modifications. In double-stranded DNA, the same pattern may appear identically on one strand and its reverse complement due to complementary base pairing. Here, we show that conventional deep learning models that do not explicitly model this property can produce substantially different predictions on forward and reverse-complement versions of the same DNA sequence. We present four new convolutional neural network layers that leverage the reverse-complement property of genomic DNA sequence by sharing parameters between forward and reverse-complement representations in the model. These layers guarantee that forward and reverse-complement sequences produce identical predictions within numerical precision. Using experiments on simulated and in vivo transcription factor binding data, we show that our proposed architectures lead to improved performance, faster learning and cleaner internal representations compared to conventional architectures trained on the same data.\n\nAvailabilityOur implementation is available at\n\nContactavanti@stanford.edu, pgreens@stanford.edu, akundaje@stanford.edu

bioinformatics

Genome-wide Search for Zelda-like Chromatin Signatures Identifies GAF as a Pioneer Factor in Early Fly Development

MotivationThe protein Zelda was shown to play a key role in early Drosophila development, binding thousands of promoters and enhancers prior to maternal-to-zygotic transition (MZT), and marking them for transcriptional activation. Recently, we showed that Zelda acts through specific chromatin patterns of histone modifications to mark developmental enhancers and active promoters. Intriguingly, some Zelda sites still maintain these chromatin patterns in Drosophila embryos lacking maternal Zelda protein. This suggests that additional Zelda-like pioneer factors may act in early fly embryos.\n\nResultsWe developed a computational method to analyze and refine the chromatin landscape surrounding early Zelda peaks, using a multi-channel spectral clustering. This allowed us to characterize their chromatin patterns through MZT (mitotic cycles 8-14). Specifically, we focused on H3K4me1, H3K4me3, H3K18ac, H3K27ac, and H3K27me3 and identified three different classes of chromatin signatures, matching \"promoters\", \"enhancers\" and \"transiently bound\" Zelda peaks.\n\nWe then further scanned the genome using these chromatin patterns and identified additional loci - with no Zelda binding - that show similar chromatin patterns, resulting with hundreds of Zelda-independent putative enhancers. These regions were found to be enriched with GAGA factor (GAF, Trl), and are typically located near early developmental zygotic genes. Overall our analysis suggests that GAF, together with Zelda, plays an important role in activating the zygotic genome.\n\nAs we show, our computational approach offers an efficient algorithm for characterizing chromatin signatures around some loci of interest, and allows a genome-wide identification of additional loci with similar chromatin patterns.\n\nContact: tommy@cs.huji.ac.il

bioinformatics

Genome-Wide Association Analysis Identifies Genetic Correlates of Immune Infiltrates in Solid Tumors.

Therapeutic options for the treatment of an increasing variety of cancers have been expanded by the introduction of a new class of drugs, commonly referred to as checkpoint blocking agents, that target the host immune system to positively modulate anti-tumor immune response. Although efficacy of these agents has been linked to a pre-existing level of tumor immune infiltrate, it remains unclear why some patients exhibit deep and durable responses to these agents therapy while others do not benefit. To examine the influence of tumor genetics on tumor immune state, we interrogated the relationship between somatic mutation and copy number alteration with infiltration levels of 7 immune cell types across 40 tumor cohorts in The Cancer Genome Atlas. Levels of cytotoxic T, regulatory T, total T, natural killer, and B cells, as well as monocytes and M2 macrophages, were estimated using a novel set of transcriptional signatures that were designed to resist interference from the cellular heterogeneity of tumors. Tumor mutational load and estimates of tumor purity were included in our association models to adjust for biases in multi-modal genomic data. Copy number alterations, mutations summarized at the gene level, and position-specific mutations were evaluated for association with tumor immune infiltration. We observed a strong relationship between copy number loss of a large region of chromosome 9p and decreased lymphocyte estimates in melanoma, pancreatic, and head/neck cancers. Mutations in the oncogenes PIK3CA, FGFR3, and RAS/RAF family members, as well as the tumor supressor TP53, were linked to changes in immune infiltration, usually in restricted tumor types. Associations of specific WNT/beta-catenin pathway genetic changes with immune state were limited, but we noted a link between 9p loss and the expression of the WNT receptor FZD3, suggesting that there are interactions between 9p alteration and WNT pathways. Finally, two different cell death regulators, CASP8 and DIDO1, were often mutated in head/neck tumors that had higher lymphocyte infiltrates. In summary, our study supports the relevance of tumor genetics to questions of efficacy and resistance in checkpoint blockade therapies. It also highlights the need to assess genome-wide influences during exploration of any specific tumor pathway hypothesized to be relevant to therapeutic response. Some of the observed genetic links to immune state, like 9p loss, may influence response to cancer immune therapies. Others, like mutations in cell death pathways, may help guide combination therapeutic approaches.

cancer biology

Whole genome linkage analysis in a large Brazilian multigenerational family reveals distinct linkage signals for Bipolar Disorder and Depression.

Both common and rare genetic variation play a role in the causes for mood disorders. Very large families pose unique opportunities and analytical challenges but may provide a way to identify regions and mutations associated with mood disorders. We identified a family with a high prevalence (~30%) of mood disorders in a rural village in Brazil, featuring decreasing age of onset over generations. The pattern of inheritance was complex with 32 Bipolar type I cases, 11 Bipolar type II and 59 recurrent and/or severe Depression cases in addition to other phenotypes. We enrolled 333 participants with DNA samples from a broader pedigree of 960 subjects for genotyping using the Affymetrix 10K array. Non-parametric linkage was carried out via MERLIN and parametric with both MERLIN and MCLINKAGE. We exome sequenced a subset of the family (n=27) in order to identify rare variation within the linkage regions shared by affected family members. We identified four genome wide significant and four suggestive linkage regions on chromosomes 1, 2, 3, 11 and 12 for different phenotype definitions. However, no region received strong joint support in both the parametric and non-parametric analyses. Exome sequencing revealed potential deleterious variants in 11p15.4 for MDD and 1q21.1-1q21.3 and 12p23.1-p22.3, implicated in cell signaling, adhesion, translation and neurogenesis processes. Overall, our results suggest promising, but not definitive or confirmed evidence, that rare genetic variation contributes to the high prevalence of mood disorders in this multi-generational family. We note that a substantial role for common genetic variation is likely given the strength of the linkage signals observed.\n\nThe World Health Organisation reports depression and bipolar disorder as the second and seventh most important causes of years lost due to disability worldwide[1]. The heritability of bipolar disorder is between 60-90% with a lower but still substantial heritability for major depression (40-45%) [2]; [3]. First-degree relatives of bipolar disorder probands have a 5-10 fold increase in risk of developing the illness compared to relatives of controls but also show a three fold increase in unipolar depression, indicating that bipolar disorder does not \"breed true\" [4]. Large collaborative genome-wide association studies (GWAS) have uncovered several common genetic variants of small effect [5]. Genomewide estimates of heritability suggest that up to 60% of the genetic risk is contributed by common variants [6]. Overall, the current picture for bipolar disorder (and almost all complex traits) is a genetic architecture formed of both common and rare variants.\n\nLinkage studies have been pursued on the basis that there may be variants of greater effect shared between and within affected families. However these studies have usually focused on collections of comparatively small families or sib pairs and few consistent findings have emerged [7]. Large multigenerational families (e. g. of >30 affected individuals) theoretically offer a powerful means for mapping complex disease loci that are individually rare but common in a single family. These loci may be more highly penetrant and of larger effect than loci found with GWAS [8]. Here we report the results of the Brazilian Bipolar Family (BBF) study on a five-generation family of 639 members of which 333 were enrolled in the current analyses. Our objectives were to perform a linkage analysis with genome coverage and try to identify new genes/mutations related to bipolar and other mood disorders in the family. Here we report our findings and preliminary results of sequencing of linkage regions.

genetics

Precise insertion and guided editing of higher plant genomes using Cpf1 CRISPR nucleases

Introduction Paragraph Introduction Paragraph Main Body Online Methods References Precise genome editing of plants has the potential to reshape global agriculture through the targeted engineering of endogenous pathways or the introduction of new traits. To develop a CRISPR nuclease-based platform that would enable higher efficiencies of precise gene insertion or replacement, we screened the Cpf1 nucleases from Francisella novicida and Lachnospiraceae bacterium ND2006 for their capacity to induce targeted gene insertions via homology directed repair. Both nucleases, in the presence of guide RNA and repairing DNA template, were demonstrated to generate precise gene insertions as well as indel mutations at the target site in the rice genome. The frequency of targeted insertions for these Cpf1 nucleases, up to 8%, is higher than most other genome editi ...

plant biology

Evolinc: a comparative transcriptomics and genomics pipeline for quickly identifyingsequence conserved lincRNAs for functional analysis.

Long intergenic non-coding RNAs (lincRNAs) are an abundant and functionally diverse class of eukaryotic transcripts. Reported lincRNA repertoires in mammals vary, but are commonly in the thousands to tens of thousands of transcripts, covering ~90% of the genome. In addition to elucidating function, there is particular interest in understanding the origin and evolution of lincRNAs. Aside from mammals, lincRNA populations have been sparsely sampled, precluding evolutionary analyses focused on lincRNA emergence and persistence. Here we present Evolinc, a two-module pipeline designed to facilitate lincRNA discovery and characterize aspects of lincRNA evolution. The first module (Evolinc-I) is a lincRNA identification workflow that also facilitates downstream differential expression analysis and genome browser visualization of identified lincRNAs. The second module (Evolinc-II) is a genomic and transcriptomic comparative analyses workflow that determines the phylogenetic depth to which a lincRNA locus is conserved within a user-defined group of related species. Evolinc-II builds families of homologous lincRNA loci, aligns constituent sequences, infers gene trees, and then uses gene tree / species tree reconciliation to reconstruct evolutionary processes such as gain, loss, or duplication of the locus. Here we demonstrate that Evolinc-I is agnostic to target organism by validating against previously annotated Arabidopsis and human lincRNA data. Using Evolinc-II, we examine ways in which conservation can rapidly be used to winnow down large lincRNA datasets to a small set of candidates for functional analysis. Finally, we show how Evolinc-II can be used to recover the evolutionary history of a known lincRNA, the human telomerase RNA (TERC). The analyses revealed unexpected duplication events as well as the loss and subsequent acquisition of a novel TERC locus in the lineage leading to mice and rats. The Evolinc pipeline is currently integrated in CyVerses Discovery Environment and is free to use by researchers.

bioinformatics

A Mixture Copula Bayesian Network Model for Multimodal Genomic Data

Gaussian Bayesian networks have become a widely used framework to estimate directed associations between joint Gaussian variables, where the network structure encodes decomposition of multivariate normal density into local terms. However, the resulting estimates can be inaccurate when normality assumption is moderately or severely violated, making it unsuitable to deal with recent genomic data such as the Cancer Genome Atlas data. In the present paper, we propose a mixture copula Bayesian network model which provides great flexibility in modeling non-Gaussian and multimodal data for causal inference. The parameters in mixture copula functions can be efficiently estimated by a routine Expectation-Maximization algorithm. A heuristic search algorithm based on Bayesian information criterion is developed to estimate the network structure, and prediction can be further improved by the best-scoring network out of multiple predictions from random initial values. Our method outperforms Gaussian Bayesian networks and regular copula Bayesian networks in terms of modeling flexibility and prediction accuracy, as demonstrated using a cell signaling dataset. We apply the proposed methods to the Cancer Genome Atlas data to study the genetic and epigenetic pathways that underlie serous ovarian cancer.

bioinformatics

GRIDSS: sensitive and specific genomic rearrangement detection using positional de Bruijn graph assembly

The identification of genomic rearrangements, particularly in cancers, with high sensitivity and specificity using massively parallel sequencing remains a major challenge. Here, we describe the Genome Rearrangement IDentification Software Suite (GRIDSS), a high-speed structural variant (SV) caller that performs efficient genome-wide break-end assembly prior to variant calling using a novel positional de Bruijn graph assembler. By combining assembly, split read and read pair evidence using a probabilistic scoring, GRIDSS achieves high sensitivity and specificity on simulated, cell line and patient tumour data, recently winning SV sub-challenge #5 of the ICGC-TCGA DREAM Somatic Mutation Calling Challenge. On human cell line data, GRIDSS halves the false discovery rate compared to other recent methods. GRIDSS identifies non-template sequence insertions, micro-homologies and large imperfect homologies, and supports multi-sample analysis. GRIDSS is freely available at https://github.com/PapenfussLab/gridss.

bioinformatics

So many genes, so little time: comments on divergence-time estimation in the genomic era

Phylogenomic datasets have been successfully used to address questions involving evolutionary relationships, patterns of genome structure, signatures of selection, and gene and genome duplications. However, despite the recent explosion in genomic and transcriptomic data, the utility of these data sources for efficient divergence-time inference remains unexamined. Phylogenomic datasets pose two distinct problems for divergence-time estimation: (i) the volume of data makes inference of the entire dataset intractable, and (ii) the extent of underlying topological and rate heterogeneity across genes makes model mis-specification a real concern. \"Gene shopping\", wherein a phylogenomic dataset is winnowed to a set of genes with desirable properties, represents an alternative approach that holds promise in alleviating these issues. We implemented an approach for phylogenomic datasets (available in SortaDate) that filters genes by three criteria: (i) clock-likeness, (ii) reasonable tree length (i.e., discernible information content), and (iii) least topological conflict with a focal species tree (presumed to have already been inferred). Such a winnowing procedure ensures that errors associated with model (both clock and topology) mis-specification are minimized, therefore reducing error in divergence-time estimation. We demonstrated the efficacy of this approach through simulation and applied it to published animal (Aves, Diplopoda, and Hymenoptera) and plant (carnivorous Caryophyllales, broad Caryophyllales, and Vitales) phylogenomic datasets. By quantifying rate heterogeneity across both genes and lineages we found that every empirical dataset examined included genes with clock-like, or nearly clock-like, behavior. Moreover, many datasets had genes that were clock-like, exhibited reasonable evolutionary rates, and were mostly compatible with the species tree. We identified overlap in age estimates when analyzing these filtered genes under strict clock and uncorrelated lognormal (UCLN) models. However, this overlap was often due to imprecise estimates from the UCLN model. We find that \"gene shopping\" can be an efficient approach to divergence-time inference for phylogenomic datasets that may otherwise be characterized by extensive gene tree heterogeneity.

evolutionary biology

Evaluation and Design of Genome-wide CRISPR/Cas9 Knockout Screens

The adaptation of CRISPR/Cas9 technology to mammalian cell lines is transforming the study of human functional genomics. Pooled libraries of CRISPR guide RNAs (gRNAs), targeting human protein-coding genes and encoded in viral vectors, have been used to systematically create gene knockouts in a variety of human cancer and immortalized cell lines, in an effort to identify whether these knockouts cause cellular fitness defects. Previous work has shown that CRISPR screens are more sensitive and specific than pooled library shRNA screens in similar assays, but currently there exists significant variability across CRISPR library designs and experimental protocols. In this study, we re-analyze 17 genome-scale knockout screens in human cell lines from three research groups using three different genome-scale gRNA libraries, using the Bayesian Analysis of Gene Essentiality (BAGEL) algorithm to identify essential genes, to refine and expand our previously defined set of human core essential genes, from 360 to 684 genes. We use this expanded set of reference Core Essential Genes (CEG2), plus empirical data from six CRISPR knockout screens, to guide the design of a sequence-optimized gRNA library, the Toronto KnockOut version 3.0 (TKOv3) library. We demonstrate the high effectiveness of the library relative to reference sets of essential and nonessential genes as well as other screens using similar approaches. The optimized TKOv3 library, combined with the CEG2 reference set, provide an efficient, highly optimized platform for performing and assessing gene knockout screens in human cell lines.

systems biology

HUGIn: Hi-C Unifying Genomic Interrogator

MotivationHigh throughput chromatin conformation capture (3C) technologies, such as Hi-C and ChlA-PET, have the potential to elucidate the functional roles of non-coding variants. However, most of published genome-wide unbiased chromatin organization studies have used cultured cell lines, limiting their generalizability.\n\nResultsWe developed a web browser, HUGIn, to visualize Hi-C data generated from 21 human primary tissues and cell liens. HUGIn enables assessment of chromatin contacts both constitutive across and specific to tissue(s) and/or cell line(s) at any genomic loci, including GWAS SNPs, eQTLs and cis-regulatory elements, facilitating the understanding of both GWAS and eQTLs results and functional genomics data.\n\nAvailabilityHUGIn is available at http://yunliweb.its.unc.edu/HUGIn.\n\nContactyunli@med.unc.edu and hum@ccf.org\n\nSupplementary information:

genetics

Genomic inferences of domestication events are corroborated by written records in Brassica rapa

Demographic modeling is often used with population genomic data to infer the relationships and ages among populations. However, relatively few analyses are able to validate these inferences with independent data. Here, we leverage written records that describe distinct Brassica rapa crops to corroborate demographic models of domestication. Brassica rapa crops are renowned for their outstanding morphological diversity, but the relationships and order of domestication remains unclear. We generated genome-wide SNPs from 126 accessions collected globally using high-throughput transcriptome data. Analyses of more than 31,000 SNPs across the B. rapa genome revealed evidence for five distinct genetic groups and supported a European-Central Asian origin of B. rapa crops. Our results supported the traditionally recognized South Asian and East Asian B. rapa groups with evidence that pak choi, Chinese cabbage, and yellow sarson are likely monophyletic groups. In contrast, the oil-type B. rapa subsp. oleifera and brown sarson were polyphyletic. We also found no evidence to support the contention that rapini is the wild type or the earliest domesticated subspecies of B. rapa. Demographic analyses suggested that B. rapa was introduced to Asia 2400-4100 years ago, and that Chinese cabbage originated 1200-2100 years ago via admixture of pak choi and European-Central Asian B. rapa. We also inferred significantly different levels of founder effect among the B. rapa subspecies. Written records from antiquity that document these crops are consistent with these inferences. The concordance between our age estimates of domestication events with historical records provides unique support for our demographic inferences.

genetics

Phandango: an interactive viewer for bacterial population genomics.

SummaryFully exploiting the wealth of data in current bacterial population genomics datasets requires synthesising and integrating different types of analysis across millions of base pairs in hundreds or thousands of isolates. Current approaches often use static representations of phylogenetic, epidemiological, statistical and evolutionary analysis results that are difficult to relate to one another. Phandango is an interactive application running in a web browser allowing fast exploration of large-scale population genomics datasets combining the output from multiple genomic analysis methods in an intuitive and interactive manner.\n\nAvailabilityPhandango is a web application freely available for use at https://jameshadfield.github.io/phandango and includes a diverse collection of datasets as examples. Source code together with a detailed wiki page is available on GitHub at https://github.com/jameshadfield/phandango\n\nContactjh22@sanger.ac.uk, sh16@sanger.ac.uk

bioinformatics

Distinguishing among modes of convergent adaptation using population genomic data

Geographically separated populations can convergently adapt to the same selection pressure. Convergent evolution at the level of a gene may arise via three distinct modes. The selected alleles can (1) have multiple independent mutational origins, (2) be shared due to shared ancestral standing variation, or (3) spread throughout subpopulations via gene flow. We present a model-based, statistical approach that utilizes genomic data to detect cases of convergent adaptation at the genetic level, identify the loci involved and distinguish among these modes. To understand the impact of convergent positive selection on neutral diversity at linked loci, we make use of the fact that hitchhiking can be modeled as an increase in the variance in neutral allele frequencies around a selected site within a population. We build on coalescent theory to show how shared hitchhiking events between subpopulations act to increase covariance in allele frequencies between subpopulations at loci near the selected site, and extend this theory under different models of migration and selection on the same standing variation. We incorporate this hitchhiking effect into a multivariate normal model of allele frequencies that also accounts for population structure. Based on this theory, we present a composite-likelihood-based approach that utilizes genomic data to identify loci involved in convergence, and distinguishes among alternate modes of convergent adaptation. We illustrate our method on genome-wide polymorphism data from two distinct cases of convergent adaptation. First, we investigate the adaptation for copper toxicity tolerance in two populations of the common yellow monkey flower, Mimulus guttatus. We show that selection has occurred on an allele that has been standing in these populations prior to the onset of copper mining in this region. Lastly, we apply our method to data from four populations of the killifish, Fundulus heteroclitus, that show very rapid convergent adaptation for tolerance to industrial pollutants. Here, we identify a single locus at which both independent mutation events and selection on an allele shared via gene flow, either slightly before or during selection, play a role in adaptation across the species range.

evolutionary biology

Conserved roles of RECQ-like helicases Sgs1 and BLM in preventing R-loop induced genome instability

Sgs1 is a yeast DNA helicase functioning in DNA replication and repair, and is the orthologue of the human Blooms syndrome helicase BLM. Here we analyze the mutation signature associated with SGS1 deletion in yeast, and find frequent copy number changes flanked by regions of repetitive sequence and high R-loop forming potential. We show that loss of SGS1 increases R-loop accumulation and sensitizes cells to replication-transcription collisions. Accordingly, in sgs1{Delta} cells the genome-wide distribution of R-loops shifts to known sites of Sgs1 action, replication pausing regions, and to long genes. Depletion of the orthologous BLM helicase from human cancer cells also increases R-loop levels, and R-loop-associated genome instability. In support of a direct effect, BLM is found physically proximal to DNA:RNA hybrids in human cells, and can efficiently unwind R-loops in vitro. Together our data describe a conserved role for Sgs1/BLM in R-loop suppression and support an increasingly broad view of DNA repair and replication fork stabilizing proteins as modulators of R-loop mediated genome instability.

cell biology

Stimulated Raman Scattering Micro-dissection Sequencing (SMD-Seq) for Morphology-specific Genomic Analysis of Oral Squamous Cell Carcinoma

Both the composition of cell types and their spatial distribution in a tissue play a critical role in cellular function, organ development, and disease progression. For example, intratumor heterogeneity and the distribution of transcriptional and genetic events in single cells drive the genesis and development of cancer. However, it can be challenging to fully characterize the molecular profile of cells in a tissue with high spatial resolution because microscopy has limited ability to extract comprehensive genomic information, and the spatial resolution of genomic techniques tends to be limited by dissection. There is a growing need for tools that can be used to explore the relationship between histological features, gene expression patterns, and spatially correlated genomic alterations in healthy and diseased tissue samples. Here, we present a technique that combines label-free histology with spatially resolved multi-omics in un-fixed and unstained tissue sections. This approach leverages stimulated Raman scattering microscopy to provide chemical contrast that reveals histological tissue architecture, allowing for high-resolution in situ laser micro-dissection of regions of interests. These micro-tissue samples are then processed for DNA and RNA sequencing to identify unique genetic profiles that correspond to distinct anatomical regions. We demonstrate the capabilities of this technique by mapping gene expression and copy number alterations to histologically defined regions in human squamous cell carcinoma (OSCC). Our approach provides complementary insights in tumorigenesis and offers an integrative tool for macroscale cancer tissues with spatial multi-omics assessments.

bioengineering

Personal Cancer Genome Reporter: Variant Interpretation Report For Precision Oncology

SummaryIndividual tumor genomes pose a major challenge for clinical interpretation due to their unique sets of acquired mutations. There is a general scarcity of tools that can i) systematically interrogate cancer genomes in the context of diagnostic, prognostic, and therapeutic biomarkers, ii) prioritize and highlight the most important findings, and iii) present the results in a format accessible to clinical experts. We have developed a stand-alone, open-source software package for somatic variant annotation that integrates a comprehensive set of knowledge resources related to tumor biology and therapeutic biomarkers, both at the gene and variant level. Our application generates a tiered report that will aid the interpretation of individual cancer genomes in a clinical setting.\n\nAvailability and ImplementationThe software is implemented in Python/R, and is freely available through Docker technology. Documentation, example reports, and installation instructions are accessible via the project GitHub page: https://github.com/sigven/pcgr)\n\nContactsigven@ifi.uio.no

bioinformatics

Genomic regression of claw keratin, taste receptor and light-associated genes inform biology and evolutionary origins of snakes

Regressive evolution of anatomical traits corresponds with the regression of genomic loci underlying such characters. As such, studying patterns of gene loss can be instrumental in addressing questions of gene function, resolving conflicting results from anatomical studies, and understanding the evolutionary history of clades. The origin of snakes coincided with the regression of a number of anatomical traits, including limbs, taste buds and the visual system. By studying the genomes of snakes, I was able to test three hypotheses associated with the regression of these features. The first concerns two keratins that are putatively specific to claws. Both genes that encode these keratins were pseudogenized/deleted in snake genomes, providing additional evidence of claw- specificity. The second hypothesis is whether snakes lack taste buds, an issue complicated by unequivocal, conflicting results in the literature. I found evidence that different snakes have lost one or more taste receptors, but all snakes examined retained at least some capacity for taste. The final hypothesis I addressed is that the earliest snakes were adapted to a dim light niche. I found evidence of deleted and pseudogenized genes with light- associated functions in snakes, demonstrating a pattern of gene loss similar to other historically nocturnal clades. Together these data also provide some bearing on the ecological origins of snakes, including molecular dating estimates that suggest dim light adaptation preceded the loss of limbs.

evolutionary biology