Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,513 records · Page 84Linked to original sources

Quality control analysis of the 1000 Genomes Project Omni2.5 genotypes

CitationFor any use of the 1000 Genomes Project data, please use the citation as noted here: http://www.1000genomes.org/faq/how-do-i-cite-1000-genomes-project. To cite this report or the lists described here, please use the following:\n\nRoslin NM, Li W, Paterson AD, Strug LJ. Quality control analysis of the 1000 Genomes Project Omni2.5 genotypes (Abstract/Program #576/F). Presented at the 66th Annual Meeting of The American Society of Human Genetics, October 18-22, 2016, Vancouver, Canada.\n\nData SummaryChips: IlluminaHumanOmni2.5-4v1_B and Illumina HumanOmni25M-8v1-1_B\n\nInitial number of SNPs: 2 458 861\n\nInitial number of samples: 2318\n\nNumber of SNPs passing QC: 1 989 184 (80.9%)\n\nNumber of samples passing QC: 2318 (100%)\n\nNumber of quasi-unrelated samples with consistent ethnicity and well inferred sex: 1736\n\nAbstractThe 1000 Genomes Project genotype 2318 individuals (48.1% male) from 19 populations in 5 continental groups on the Illumina Omni2.5 platform. The data are publicly available, and will prove a valuable resource to obtain ethnic-specific allele frequencies, as well as exploring population histories through principal components analysis (PCA), estimation of inbreeding coefficients, and admixture analysis. As in any study, the data should be cleaned prior to analysis, to remove individuals or markers of questionable quality. Furthermore, a thorough understanding of the relationships between individuals must be established. Here we report our findings after comprehensive examination of the data for quality control.\n\nThe basic quality of the genotypes was assessed using standard procedures. KING version 1.4 was used to confirm the relationships in the provided pedigrees, and also to detect undeclared relationships. PCA was used to examine the similarities and differences between individuals among and between population groups.\n\nIn general, the data was found to be of high quality. No samples were removed due to low call rate (<97%) or excess heterozygosity. Sex chromosome genotypes showed two individuals with discrepancies between reported and inferred sex, and were unable to determine sex in an additional 20 individuals; the sex for these was changed to unknown. Relationship checking found discrepancies between first-degree relationships in the provided pedigrees and the genotypes in 9 families, including one instance where a reported parent/child pair was unrelated, two instances where full sibs were unrelated, and one set of three individuals who formed a newly defined trio. A set of 1756 individuals who were inferred to be more distant than 3rd degree relatives was extracted and used in PCA. These individuals clustered in a pattern that is consistent with other published reports of global populations. We identified 4 individuals whose genotypes clustered more closely with a different geographic region than the one in the provided data.\n\nAlthough the genotype data is of high quality, errors exist in the publicly available dataset that require attention prior to using the genotypes. PLINK-format files including SNPs with good quality metrics and revised pedigree structures is available at http://tcag.ca. Files with distantly related or unrelated individuals, with sex inference consistent with provided gender, and with PCA consistent with continental group are also available.

Genetics

Assessing Pathogens for Natural versus Laboratory Origins Using Genomic Data and Machine Learning

Pathogen genomic data is increasingly important in investigations of infectious disease outbreaks. The objective of this study is to develop methods for using large-scale genomic data to determine the type of the environment an outbreak pathogen came from. Specifically, this study focuses on assessing whether an outbreak strain came from a natural environment or experienced substantial laboratory culturing. The approach uses phylogenetic analyses and machine learning to identify DNA changes that are characteristic of laboratory culturing. The analysis methods include parallelized sequence read alignment, variant identification, phylogenetic tree construction, ancestral state reconstruction, semi-supervised classification, and random forests. These methods were applied to 902 Salmonella enterica serovar Typhimurium genomes from the NCBI Sequence Read Archive database. The analyses identified candidate signatures of laboratory culturing that are highly consistent with genes identified in published laboratory passage studies. In particular, the analysis identified mutations in rpoS, hfq, rfb genes, acrB, and rbsR as strong signatures of laboratory culturing. In leave-one-out cross-validation, the classifier had an area under the receiver operating characteristic (ROC) curve of 0.89 for strains from two laboratory reference sets collected in the 1940s and 1980s. The classifier was also used to assess laboratory culturing in foodborne and laboratory acquired outbreak strains closely related to laboratory reference strain serovar Typhimurium 14028. The classifier detected some evidence of laboratory culturing on the phylogeny branch leading to this clade, suggesting all of these strains may have a common ancestor that experienced laboratory culturing. Together, these results suggest that phylogenetic analysis and machine learning could be used to assess whether pathogens collected from patients are naturally occurring or have been extensively cultured in laboratories. The data analysis methods can be applied to any bacterial pathogen species, and could be adapted to assess viral pathogens and other types of source environments.

Bioinformatics

Genome-level parameters describe the pan-nuclear fractal nature of eukaryotic interphase chromosomal arrangement

Long-range inter-chromosomal interactions in the interphase nucleus subsume critical genome-level regulatory functions such as transcription and gene expression. To decipher the physical basis of diverse pan-nuclear patterns of chromosomal arrangement that facilitates these processes, we investigate the scaling effects within disparate genomes and compared their total number of genes with chromosome size. First, we derived the pan-nuclear average fractal dimension of inter-chromosomal arrangement in interphase nuclei of different species and corroborated our predictions with independently reported results. Then, we described the different patterns across disparate unicellular and multicellular eukaryotes. We report that, unicellular lower eukaryotes have inter-chromosomal fractal dimension = 1 at the pan-nuclear scales, which is analogous to the multi-polymer crumpled globule model. Multi-fractal dimensions, corresponding to different inter-chromosomal arrangements emerged from multicellular eukaryotes, such that closely related species have relatively similar patterns. Using this theoretical approach, we could distinguish fractal patterns from human acrocentric versus metacentric chromosomes, implying that the multi-fractal nature of inter-chromosomal geometry facilitates viable large-scale chromosomal aberrations, such as Robertsonian translocations. We report that the nature of such an average multi-fractal dimension for nocturnal mammals is very different in diurnal mammals, which suggests a greatly enhanced plasticity in arrangement across different cell types, for example retinal versus dermal fibroblasts. Altogether, our results substantiate that genome-level constraints have also co-evolved with the average pan-nuclear fractal dimension of inter-chromosomal folding during eukaryotic evolution.

evolutionary biology

An Optimized Approach for Annotation of Large Eukaryotic Genomic Sequences using Genetic Algorithm

Detection of important functional and/or structural elements and identifying their positions in a large eukaryotic genome is an active research area. Gene is an important functional and structural unit of DNA. The computation of gene prediction is essential for detailed genome annotation. In this paper, we propose a new gene prediction technique based on Genetic Algorithm (GA) for determining the optimal positions of exons of a gene in a chromosome or genome. The correct identification of the coding and non-coding regions are difficult and computationally demanding. The proposed genetic-based method, named Gene Prediction with Genetic Algorithm (GPGA), reduces this problem by searching only one exon at a time instead of all exons along with its introns. The advantage of this representation is that it can break the entire gene-finding problem into a number of smaller subspaces and thereby reducing the computational complexity. We tested the performance of the GPGA with some benchmark datasets and compared the results with the well-known and relevant techniques. The comparison shows the better or comparable performance of the proposed method (GPGA). We also used GPGA for annotating the human chromosome 21 (HS21) using cross species comparison with the mouse orthologs.

bioinformatics

Metafounders are Fst fixation indices and reduce bias in single step genomic evaluations

BACKGROUNDMetafounders are pseudo-individuals that condense the genetic heterozygosity and relationships within and across base pedigree populations, i.e. ancestral populations. This work addresses estimation and usefulness of metafounder relationships in Single Step GBLUP.\n\nRESULTSWe show that the ancestral relationship parameters are proportional to standardized covariances of base allelic frequencies across populations, like Fst fixation indexes. These covariances of base allelic frequencies can be estimated from marker genotypes of related recent individuals, and pedigree. Simple methods for estimation include naive computation of allele frequencies from marker genotypes or a method of moments equating average pedigree-based and marker-based relationships. Complex methods include generalized least squares or maximum likelihood based on pedigree relationships. To our knowledge, methods to infer Fst coefficients and Fst differentiation have not been developed for related populations.\n\nA compatible genomic relationship matrix constructed as a crossproduct of {-1,0,1} codes, and equivalent (up to scale factors) to an identity by state relationship matrix at the markers, is derived. Using a simulation with a single population under selection, in which only males and youngest animals were genotyped, we observed that generalized least squares or maximum likelihood gave accurate and unbiased estimates of the ancestral relationship parameter (true value: 0.40) whereas the other two (naive and method of moments) were biased (estimates of 0.43 and 0.35). We also observed that genomic evaluation by Single Step GBLUP using metafounders was less biased in terms of accurate genetic trend (0.01 instead of 0.12 bias), slightly overdispersed (0.94 instead of 0.99) and as accurate (0.74) than the regular Single Step GBLUP. Single Step GBLUP using metafounders also provided consistent estimates of heritability.\n\nCONCLUSIONSEstimation of metafounder relationship can be achieved using BLUP-like methods with pedigree and markers. Inclusion of metafounder relationships improves bias of genomic predictions with no loss in accuracy.

genetics

Phylogenetic relationships and genome size evolution within the genus Amaranthus indicate the ancestors of an ancient crop

The genus Amaranthus consists of 50 to 70 species and harbors several cultivated and weedy species of great economic importance. A small number of suitable traits, phenotypic plasticity, gene flow and hybridization made it difficult to establish the taxonomy and phylogeny of the whole genus despite various studies using molecular markers. We inferred the phylogeny of the Amaranthus genus using genotyping by sequencing (GBS) of 94 genebank accessions representing 35 Amaranthus species and measured their genome sizes. SNPs were called by de novo and reference-based methods, for which we used the distant sugarbeet Beta vulgaris and the closely related Amaranthus hypochondriacus as references. SNP counts and proportions of missing data differed between methods, but the resulting phylogenetic trees were highly similar. A distance-based neighbor joing tree of individual accessions and a species tree calculated with the multispecies coalescent supported a previous taxonomic classification into three subgenera although the subgenus A. Acnida consists of two highly differentiated clades. The analysis of the Hybridus complex within the A. Amaranthus subgenus revealed insights on the history of cultivated grain amaranths. The complex includes the three cultivated grain amaranths and their wild relatives and was well separated from other species in the subgenus. Wild and cultivated amaranth accessions did not differentiate according to the species assignment but clustered by their geographic origin from South and Central America. Different geographically separated populations of Amaranthus hybridus appear to be the common ancestors of the three cultivated grain species and A. quitensis might be additionally be involved in the evolution of South American grain amaranth (A. caudatus). We also measured genome sizes of the species and observed little variation with the exception of two lineages that showed evidence for a recent polyploidization. With the exception of two lineages, genome sizes are quite similar and indicate that polyploidization did not play a major role in the history of the genus.

evolutionary biology

Genome-Wide Fitness Analyses of the Foodborne Pathogen Campylobacter jejuni in In Vitro and In Vivo Models

Infection by Campylobacter is recognised as the most common cause of foodborne bacterial illness worldwide. Faecal contamination of meat, especially chicken, during processing represents a key route of transmission to humans. There is currently no licenced vaccine and no Campylobacter-resistant chickens. In addition, preventative measures aimed at reducing environmental contamination and exposure of chickens to Campylobacter jejuni (biosecurity) have been ineffective. There is much interest in the factors/mechanisms that drive C. jejuni colonisation and infection of animals, and survival in the environment. It is anticipated that understanding these mechanisms will guide the development of effective intervention strategies to reduce the burden of C. jejuni infection. Here we present a comprehensive analysis of C. jejuni fitness during growth and survival within and outside hosts. A comparative analysis of transposon (Tn) gene inactivation libraries in three C. jejuni strains by Tn-seq demonstrated that a large proportion, 331 genes, of the C. jejuni genome is dedicated to (in vitro) growth. An extensive Tn library in C. jejuni M1cam (~10,000 mutants) was screened for the colonisation of commercial broiler chickens, survival in houseflies and under nutrient-rich and-poor conditions at low temperature, and infection of human gut epithelial cells. We report C. jejuni factors essential throughout its life cycle and we have identified genes that fulfil important roles across multiple conditions, including maf3, fliW, fliD, pflB and capM, as well as novel genes uniquely implicated in survival outside hosts. Taking a comprehensive screening approach has confirmed previous studies, that the flagella are central to the ability of C. jejuni to interact with its hosts. Future efforts should focus on how to exploit this knowledge to effectively control infections caused by C. jejuni.\n\nAuthor SummaryCampylobacter jejuni is the leading bacterial cause of human diarrhoeal disease. C. jejuni encounters and has to overcome a wide range of \"stress\" conditions whilst passing through the gastrointestinal tract of humans and other animals, during processing of food products, on/in food and in the environment. We have taken a comprehensive approach to understand the basis of C. jejuni growth and within/outside host survival, with the aim to inform future development of intervention strategies. Using a genome-wide transposon gene inactivation approach we identified genes core to the growth of C. jejuni. We also determined genes that were required during the colonisation of chickens, survival in the housefly and under nutrient-rich and -poor conditions at low temperature, and during interaction with human gut epithelial tissue culture cells. This study provides a comprehensive dataset linking C. jejuni genes to growth and survival in models relevant to its life cycle. Genes important across multiple models were identified as well as genes only required under specific conditions. We identified that a large proportion of the C. jejuni genome is dedicated to growth and that the flagella fulfil a prominent role in the interaction with hosts. Our data will aid development of effective control strategies.

microbiology

lncRNA-screen: an interactive platform for computationally screening long non-coding RNAs in large genomics datasets

Long non-coding RNAs (lncRNAs) have emerged as a class of factors that are important for regulating development and cancer. Computational prediction of lncRNAs from ultra-deep RNA sequencing has been successful in identifying candidate lncRNAs. However, the complexity of handling and integrating different types of genomics data poses significant challenges to experimental laboratories that lack extensive genomics expertise. To address this issue, we have developed lncRNA-screen, a comprehensive pipeline for computationally screening putative lncRNA transcripts over large multimodal datasets. The main objective of this work is to facilitate the computational discovery of lncRNA candidates to be further examined by functional experiments. lncRNA-screen provides a fully automated easy-to-run pipeline which performs data download, RNA-seq alignment, assembly, quality assessment, transcript filtration, novel lncRNA identification, coding potential estimation, expression level quantification, histone mark enrichment profile integration, differential expression analysis, annotation with other type of segmented data (CNVs, SNPs, Hi-C, etc.) and visualization. Importantly, lncRNA-screen generates an interactive report summarizing all interesting lncRNA features including genome browser snapshots and lncRNA-mRNA interactions based on Hi-C data. In summary, our pipeline provides a comprehensive solution for lncRNA discovery and an intuitive interactive report for identifying promising lncRNA candidates. lncRNA-screen is available as free open-source software on GitHub.

bioinformatics

Genome-scale transcriptional regulatory network models for the mouse and human striatum predict roles for SMAD3 and other transcription factors in Huntington’s disease

Transcriptional changes occur presymptomatically and throughout Huntingtons Disease (HD), motivating the study of transcriptional regulatory networks (TRNs) in HD. We reconstructed a genome-scale model for the target genes of 718 TFs in the mouse striatum by integrating a model of the genomic binding sites with transcriptome profiling of striatal tissue from HD mouse models. We identified 48 differentially expressed TF-target gene modules associated with age- and Htt allele-dependent gene expression changes in the mouse striatum, and replicated many of these associations in independent transcriptomic and proteomic datasets. Strikingly, many of these predicted target genes were also differentially expressed in striatal tissue from human disease. We experimentally validated a key model prediction that SMAD3 regulates HD-related gene expression changes using chromatin immunoprecipitation and deep sequencing (ChIP-seq) of mouse striatum. We found Htt allele-dependent changes in the genomic occupancy of SMAD3 and confirmed our models prediction that many SMAD3 target genes are down-regulated early in HD. Importantly, our study provides a mouse and human striatal-specific TRN and prioritizes a hierarchy of transcription factor drivers in HD.

systems biology

Genomic determinants of protein abundance variation in colorectal cancer cells

Assessing the extent to which genomic alterations compromise the integrity of the proteome is fundamental in identifying the mechanisms that shape cancer heterogeneity. We have used isobaric labelling and tribrid mass spectrometry to characterize the proteomic landscapes of 50 colorectal cancer cell lines and to decipher the relationships between genomic and proteomic variation. The robust quantification of 12,000 proteins and 27,000 phosphopeptides revealed how protein symbiosis translates to a co-variome which is subjected to a hierarchical order and exposes the collateral effects of somatic mutations on protein complexes. Targeted depletion of key chromatin modifiers confirmed the transmission of variation and the directionality as characteristics of protein interactions. Protein level variation was leveraged to build drug response predictive models towards a better understanding of pharmacoproteomic interactions in colorectal cancer. Overall, we provide a deep integrative view of the molecular structure underlying the variation of colorectal cancer cells.\n\nHighlightsO_LIThe cancer cell functional \"co-variome\" is a strong attribute of the proteome.\nC_LIO_LIMutations can have a direct impact on protein levels of chromatin modifiers.\nC_LIO_LITransmission of genomic variation is a characteristic of protein interactions.\nC_LIO_LIPharmacoproteomic models are strong predictors of response to DNA damaging agents.\nC_LI\n\nAbbreviations

systems biology

SNVPhyl: A Single Nucleotide Variant Phylogenomics pipeline for microbial genomic epidemiology

MotivationThe recent widespread application of whole-genome sequencing (WGS) for microbial disease investigations has spurred the development of new bioinformatics tools, including a notable proliferation of phylogenomics pipelines designed for infectious disease surveillance and outbreak investigation. Transitioning the use of WGS data out of the research lab and into the front lines of surveillance and outbreak response requires user-friendly, reproducible, and scalable pipelines that have been well validated.\n\nResultsSNVPhyl (Single Nucleotide Variant Phylogenomics) is a bioinformatics pipeline for identifying high-quality SNVs and constructing a whole genome phylogeny from a collection of WGS reads and a reference genome. Individual pipeline components are integrated into the Galaxy bioinformatics framework, enabling data analysis in a user-friendly, reproducible, and scalable environment. We show that SNVPhyl can detect SNVs with high sensitivity and specificity and identify and remove regions of high SNV density (indicative of recombination). SNVPhyl is able to correctly distinguish outbreak from non-outbreak isolates across a range of variant-calling settings, sequencing-coverage thresholds, or in the presence of contamination.\n\nAvailabilitySNVPhyl is available as a Galaxy workflow, Docker and virtual machine images, and a Unix-based command-line application. SNVPhyl is released under the Apache 2.0 license and available at http://snvphyl.readthedocs.io/ or at https://github.com/phac-nml/snvphyl-galaxy.

bioinformatics

Whole genome sequence accuracy is improved by replication in a population of mutagenized sorghum.

The accurate detection of induced mutations is critical for both forward and reverse genetics studies. Experimental chemical mutagenesis induces relatively few single base changes per individual. In a complex eukaryotic genome, false positive detection of mutations can occur at or above this mutagenesis rate. We demonstrate here, using a population of ethyl methanesulfonate (EMS) treated Sorghum bicolor BTx623 individuals, that using replication to detect false positive induced variants in next-generation sequencing data permits higher throughput variant detection with greater accuracy. We used a lower sequence coverage depth (average of 7X) from 586 independently mutagenized individuals and detected 5,399,493 homozygous SNPs. Of these, 76% originated from only 57,872 genomic positions prone to false positive variant calling. These positions are characterized by high copy number paralogs where the error-prone SNP positions are at copies containing a variant at the SNP position. The ability of short stretches of homology to generate these error prone positions suggests that incompletely assembled or poorly mapped repeated sequences are one driver of these error prone positions. Removal of these false positives left 1,275,872 homozygous and 477,531 heterozygous EMS-induced SNPs which, congruent with the mutagenic mechanism of EMS, were greater than 98% G:C to A:T transitions. Through this analysis we generated a database of sequence indexed mutants of Sorghum. This collection contains 4,035 high impact homozygous mutations in 3,637 genes and 56,514 homozygous missense mutations in 23,227 genes. Each line contains, on average, 2,177 annotated homozygous SNPs per genome, including seven likely gene knockouts and 96 missense mutations. The number of mutations in a transcript was linearly correlated with the transcript length and also the G+C count, but not with the GC/AT ratio. Analysis of the detected mutagenized positions identified CG-rich patches, and flanking sequences strongly influenced EMS-induced mutation rates. Our method for detecting false-positive induced mutations is generally applicable to any organism, is independent of the choice of in silico variant-calling algorithm, and is most valuable when the true mutation rate is likely to be low, such as in laboratory induced mutations or somatic mutation detection in medicine.

genetics

Unicycler: resolving bacterial genome assemblies from short and long sequencing reads

1.The Illumina DNA sequencing platform generates accurate but short reads, which can be used to produce accurate but fragmented genome assemblies. Pacific Biosciences and Oxford Nanopore Technologies DNA sequencing platforms generate long reads that can produce more complete genome assemblies, but the sequencing is more expensive and error prone. There is significant interest in combining data from these complementary sequencing technologies to generate more accurate \"hybrid\" assemblies. However, few tools exist that truly leverage the benefits of both types of data, namely the accuracy of short reads and the structural resolving power of long reads. Here we present Unicycler, a new tool for assembling bacterial genomes from a combination of short and long reads, which produces assemblies that are accurate, complete and cost-effective. Unicycler builds an initial assembly graph from short reads using the de novo assembler SPAdes and then simplifies the graph using information from short and long reads. Unicycler utilises a novel semi-global aligner, which is used to align long reads to the assembly graph. Tests on both synthetic and real reads show Unicycler can assemble larger contigs with fewer misassemblies than other hybrid assemblers, even when long read depth and accuracy are low. Unicycler is open source (GPLv3) and available at github.com/rrwick/Unicycler.

bioinformatics

grID: A CRISPR-Cas9 guide RNA Database and Resource for Genome-Editing

CRISPR-Cas9 genome-editing is a revolutionary technology that is transforming biological research. The explosive growth and advances in CRISPR research over the last few years, coupled with the potential for clinical applications and therapeutics, is heralding a new era for genome engineering. To further support this technology platform and to provide a universal CRISPR annotation system, we introduce the grID database (http://crispr.technology), an extensive compilation of gRNA properties including sequence and variations, thermodynamic parameters, off-target analyses, and alternative PAM sites, among others. To aid in the design of optimal gRNAs, the website is integrated with other prominent databases, providing a wealth of additional resources to guide users from in silico analysis through experimental CRISPR targeting. Here, we make available all the tools, protocols, and plasmids that are needed for successful CRISPR-based genome targeting.

bioinformatics

Bayesian divergence-time estimation with genome-wide SNP data of sea catfishes (Ariidae) supports Miocene closure of the Panamanian Isthmus

The closure of the Isthmus of Panama has long been considered to be one of the best defined biogeographic calibration points for molecular divergence-time estimation. However, geological and biological evidence has recently cast doubt on the presumed timing of the initial isthmus closure around 3 Ma but has instead suggested the existence of temporary land bridges as early as the Middle or Late Miocene. The biological evidence supporting these earlier land bridges was based either on only few molecular markers or on concatenation of genome-wide sequence data, an approach that is known to result in potentially misleading branch lengths and divergence times, which could compromise the reliability of this evidence. To allow divergence-time estimation with genomic data using the more appropriate multi-species coalescent model, we here develop a new method combining the SNP-based Bayesian species-tree inference of the software SNAPP with a molecular clock model that can be calibrated with fossil or biogeographic constraints. We validate our approach with simulations and use our method to reanalyze genomic data of Neotropical army ants (Dorylinae) that previously supported divergence times of Central and South American populations before the isthmus closure around 3 Ma. Our reanalysis with the multi-species coalescent model shifts all of these divergence times to ages younger than 3 Ma, suggesting that the older estimates supporting the earlier existence of temporary land bridges were artifacts resulting at least partially from the use of concatenation. We then apply our method to a new RAD-sequencing data set of Neotropical sea catfishes (Ariidae) and calibrate their species tree with extensive information from the fossil record. We identify a series of divergences between groups of Caribbean and Pacific sea catfishes around 10 Ma, indicating that processes related to the emergence of the isthmus led to vicariant speciation already in the Late Miocene, millions of years before the final isthmus closure.

evolutionary biology

The genomic basis of adaptation to the deep water ‘twilight zone’ in Lake Malawi cichlid fishes

Deep water environments are characterized by low levels of available light at increasingly narrow spectra, great hydrostatic pressure and reduced dissolved oxygen - conditions predicted to exert highly specific selection pressures. In Lake Malawi over 800 cichlid species have evolved, and this adaptive radiation extends into the \"twilight zone\" below 100 metres. We use population-level RAD-seq data to investigate whether four endemic deep water species (Diplotaxodon spp.) have experienced divergent selection within this environment. We identify candidate genes including regulators of photoreceptor function, photopigments, lens morphology and haemoglobin, many not previously implicated in cichlid adaptive radiations. Co-localization of functionally linked genes suggests co-adapted \"supergene\" complexes. Comparisons of Diplotaxodon to the broader Lake Malawi radiation using genome resequencing data revealed functional substitutions in candidate genes. Our data provide unique insights into genomic adaptation to life at depth, and suggest genome-level specialisation for deep water habitat as an important process in cichlid radiation.

evolutionary biology

Resolution of conflict between parental genomes in a hybrid species

The development of reproductive barriers against parent species is crucial during hybrid speciation, and post-zygotic isolation can be important in this process. Genetic incompatibilities that normally isolate the parent species can become sorted in hybrids to form reproductive barriers towards either parent. However, the extent to which this sorting process is systematically biased and therefore predictable in which loci are involved and which alleles are favored is largely unknown. Theoretically, reduced fitness in hybrids due to the mixing of differentiated genomes can be resolved through rapid evolution towards allelic combinations ancestral to lineage-splitting of the parent species, as these alleles have successfully coexisted in the past. However, for each locus, this effect may be influenced by its chromosomal location, function, and interactions with other loci. We use the Italian sparrow, a homoploid hybrid species that has developed post-zygotic barriers against its parent species, to investigate this prediction. We show significant bias towards fixation of the ancestral allele among 57 nuclear intragenic SNPs, particularly those with a mitochondrial function whose ancestral allele came from the same parent species as the mitochondria. Consistent with increased pleiotropy leading to stronger fitness effects, genes with more protein-protein interactions were more biased in favor of the ancestral allele. Furthermore, the number of protein-protein interactions was especially low among candidate incompatibilities still segregating within Italian sparrows, suggesting that low pleiotropy allows steep intraspecific clines in allele frequencies to form. Finally, we report evidence for pervasive epistatic interactions within one Italian sparrow population, particularly involving loci isolating the two parent species but not hybrid and parent. However there was a lack of classic incompatibilities and no admixture linkage disequilibrium. This suggests that parental genome admixture can continue to constrain evolution and prevent genome stabilization long after incompatibilities have been purged.

evolutionary biology

LRSim: a Linked Reads Simulator generating insights for better genome partitioning

MotivationLinked reads are a form of DNA sequencing commercialized by 10X Genomics that uses highly multiplexed barcoding within microdroplets to tag short reads to progenitor molecules. The linked reads, spanning tens to hundreds of kilobases, offer an alternative to long-read sequencing for de novo assembly, haplotype phasing and other applications. However, there is no available simulator, making it difficult to measure their capability or develop new informatics tools.\n\nResultsOur analysis of 13 real linked read datasets revealed their characteristics of barcodes, molecules and partitions. Based on this, we introduce LRSim that simulates linked reads by emulating the library preparation and sequencing process with fine control of 1) the number of simulated variants; 2) the linked-read characteristics; and 3) the Illumina reads profile. We conclude from the phasing and genome assembly of multiple datasets, recommendations on coverage, fragment length, and partitioning when sequencing human and non-human genome.\n\nAvailabilityLRSIM is under MIT license and is freely available at https://github.com/aquaskyline/LRSIM\n\nContactrluo5@jhu.edu

bioinformatics