Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

De novo assembly of two Swedish genomes reveals missing segments from the human GRCh38 reference and improves variant calling of population-scale sequencing data

We have performed de novo assembly of two Swedish genomes using long-read sequencing and optical mapping, resulting in total assembly sizes of nearly 3 Gb and hybrid scaffold N50 values of over 45 Mb. A further analysis revealed over 10 Mb of sequences absent from the human GRCh38 reference in each individual. Around 6 Mb of these novel sequences (NS) are shared with a Chinese personal genome. The NS are highly repetitive, have elevated GC-content and are primarily located in centromeric or telomeric regions. A BLAST search showed that 31% of the NS are different from any sequences deposited in nucleotide databases. The remaining NS correspond to human (62%) or primate (6%) nucleotide entries, while 1% of hits show the highest similarity to other species, including mouse and a few different classes of parasitic worms. Up to 1 Mb of NS can be assigned to chromosome Y, and large segments are missing from GRCh38 also at chromosomes 14, 17 and 21. Inclusion of these novel sequences into the GRCh38 reference radically improves the alignment and variant calling of whole-genome sequencing data at several genomic loci. Through a re-analysis of 200 samples from a Swedish population-scale sequencing project, we obtained over 75,000 putative novel SNVs per individual when using a custom version of GRCh38 extended with 17.3 Mb of NS. In addition, about 10,000 false positive SNV calls per individual were removed from the GRCh38 autosomes and sex chromosomes in the re-analysis, with some of them located in protein coding regions.

genomics

RNA-Seq in 296 phased trios provides a high resolution map of genomic imprinting

Combining allelic analysis of RNA-Seq data with phased genotypes in family trios provides a powerful method to detect parent-of-origin biases in gene expression. We report findings in 296 family trios from two large studies: 165 lymphoblastoid cell lines from the 1000 Genomes Project, and 131 blood samples from the Genome of the Netherlands participants (GoNL). Based on parental haplotypes we identified >2.8 million transcribed heterozygous SNVs phased for parental origin, and developed a robust statistical framework for measuring allelic expression. We identified a total of 45 imprinted genes and one imprinted unannotated transcript, 17 of which have not previously been reported as showing parental expression bias. Multiple novel imprinted transcripts showing incomplete parental expression bias were located adjacent to known strongly imprinted genes. For example, PXDC1, a gene which lies adjacent to the paternally-expressed gene FAM50B, shows a 2:1 paternal expression bias. Other novel imprinted genes had promoter regions that coincide with sites of parentally-biased DNA methylation identified in blood from uniparental disomy (UPD) samples, thus providing independent validation of our results. Using the stranded nature of the RNA-Seq data in LCLs we identified multiple loci with overlapping sense/antisense transcripts, of which one is expressed paternally and the other maternally. Using a sliding window approach, we searched for imprinted expression across the entire genome, identifying a novel imprinted putative lncRNA in 13q21.2. Our methods and data provide a robust and high resolution map of imprinted gene expression in the human genome.

genomics

Inter- and intra-specific genomic divergence in Drosophila montana shows evidence for cold adaptation

The genomes of species that are ecological specialists will likely contain signatures of genomic adaptation to their niche. However, distinguishing genes related to their ecological specialism from other sources of selection and more random changes is a challenge. Here we describe the genome of Drosophila montana, the most extremely cold-adapted Drosophila species. We describe the genome, which is similar in size and gene content to most Drosophila species. We look for evidence of accelerated divergence from a previously sequenced relative, and do not find strong evidence for divergent selection on coding sequence variation. We use branch tests to identify genes showing accelerated divergence in contrasts between cold- and warm adapted species and identify about 250 genes that show differences, possibly driven by a lower synonymous substitution rate in cold-adapted species. Divergent genes are involved in a variety of functions, including cuticular and olfactory processes. We also re-sequenced three populations of D. montana representing its ecological and geographic range. Outlier loci were more likely to be found on the X chromosome and there was a greater than expected overlap between population outliers and those genes implicated in cold adaptation between Drosophila species, implying some continuity of selective process at these different evolutionary scales.

genomics

Large-scale gene losses underlie the genome evolution of parasitic plant Cuscuta australis

Dodders (Cuscuta spp., Convolvulaceae) are globally distributed root- and leafless parasitic plants that parasitize a wide range of hosts. The physiology, ecology, and evolution of these obligate parasites are still poorly understood. A high-quality reference genome (size 266.74 Mb and contig N50 of 3.63 Mb) of Cuscuta australis was assembled. Our analyses reveal that Cuscuta experienced accelerated evolution, and Cuscuta and the convolvulaceous morning glory (Ipomoea) shared a common whole-genome triplication event before their divergence. Importantly, C. australis genome harbors only 19805 protein-coding genes, and 11.7% of the conserved orthologs in autotrophic plants are lost in C. australis. Many of these gene loss events likely result from the plants parasitic lifestyle and large changes in its body plan. Moreover, comparison of the gene expression patterns in Cuscuta prehaustoria/haustoria and various tissues of closely related autotrophic plants suggests that Cuscuta haustorium genes largely evolved from roots. The C. australis genome provides important resources for studying the evolution of parasitism, regressive evolution, and evo-devo in plant parasites.

genomics

Effect of selection on bias and accuracy in genomic prediction of breeding values

Reference populations for genomic selection (GS) usually involve highly selected individuals, which may result in biased prediction of estimated genomic breeding values (GEBV). In the present study, bias and accuracy of GEBV were explored for various genetic models and prediction methods when using selected individuals for a reference. Data were simulated for an animal breeding program to compare Best Linear Unbiased Prediction of breeding values using pedigree based relationships (PBLUP), genomic relationships for genotyped animals only (GBLUP) and a Single Step approach (SSGBLUP), where information on genotyped individuals was used to infer a matrix H with relationships among all available genotyped and non-genotyped individuals that were linked through pedigree. In SSGBLUP, various weights (=0.95, 0.80, 0.50) for the genomic relationship matrix (G) relative to the numerator relationship matrix (A) were applied to construct H and in another version (SSGBLUP_F), inbreeding was accounted for while computing A-1. With GBLUP, accuracy of GEBV prediction increased linearly with an increase in the number of animals selected in reference. For the scenario with no-selection and random mating (RR) prediction was unbiased. For GBLUP, lower accuracy and bias observed in the scenarios with selection and random mating (SR) or selection and positive assortative mating (SA), in which prediction bias increased when a smaller and highly selected proportion genotyped. Bias disappeared when all individuals were genotyped. SSGBLUP_F showed higher accuracy compared to GBLUP and bias of prediction was negligible even with selective genotyping. However, PBLUP and SSGBLUP showed bias in SA owing to not fully accounting for allele frequency changes because of selection of quantitative trait loci (QTL) with larger effects and also due to high inbreeding rate. In genetic models with fewer QTL but each with larger effect, predictions were less accurate and more biased for selection scenarios. Results suggest that prediction accuracy and bias is affected by the genetic architecture of the trait. Selective genotyping lead to significant bias in GEBV prediction. SSGBLUP with appropriate scaling of A and G matrices can provide accurate and less biased prediction but scaling requires careful consideration in populations under selection and with high levels of inbreeding.

genomics

Genomic consequences of a recent three-way admixture in supplemented wild brown trout populations revealed by ancestry tracts

Understanding the evolutionary consequences of human-mediated introductions of domestic strains into the wild and their subsequent admixture with natural populations is of major concern in conservation biology. In the brown trout Salmo trutta, decades of stocking practices have profoundly impacted the genetic makeup of wild populations. Small local Mediterranean populations in the Orb River watershed (Southern France) have been subject to successive introductions of domestic strains derived from the Atlantic and Mediterranean lineages. However, the genomic impacts of two distinct sources of stocking (locally-derived vs divergent) on the genetic integrity of wild populations remain poorly understood. Here, we evaluate the extent of admixture from both domestic strains within three wild populations of this watershed, using 75,684 mapped SNPs obtained from double-digest restriction-site-associated DNA sequencing (dd-RADseq). Using a local ancestry inference approach, we provide a detailed picture of admixture patterns across the brown trout genome at the haplotype level. By analysing the chromosomal ancestry profiles of admixed individuals, we reveal a wider diversity of hybrid and introgressed genotypes than estimated using classical methods for inferring ancestry and hybrid pedigree. In addition, the length distribution of introgressed tracts retained different timings of introgression between the two domestic strains. We finally reveal opposite consequences of admixture on the level of polymorphism of the recipient populations between domestic strains. Our study illustrates the potential of using the information contained in the genomic mosaic of ancestry tracts in combination with classical methods based on allele frequencies for analysing multiple-way admixture with population genomic data.

genomics

Distinctive types of postzygotic single-nucleotide mosaicisms in healthy individuals revealed by genome-wide profiling of multiple organs

Postzygotic single-nucleotide mosaicisms (pSNMs) have been extensively studied in tumors and are known to play critical roles in tumorigenesis. However, the patterns and origin of pSNMs in normal organs of healthy humans remain largely unknown. Using whole-genome sequencing and ultra-deep amplicon re-sequencing, we identified and validated 164 pSNMs from 27 postmortem organ samples obtained from five healthy donors. The mutant allele fractions ranged from 1.0% to 29.7%. Inter- and intra-organ comparison revealed two distinctive types of pSNMs, with about half originating during early embryogenesis (embryonic pSNMs) and the remaining more likely to result from clonal expansion events that had occurred more recently (clonal expansion pSNMs). Compared to clonal expansion pSNMs, embryonic pSNMs had higher proportion of C>T mutations with elevated mutation rate at CpG sites. We observed differences in replication timing between these two types of pSNMs, with embryonic and clonal expansion pSNMs enriched in early- and late-replicating regions, respectively. An increased number of embryonic pSNMs were located in open chromatin states and topologically associating domains that transcribed embryonically. Our findings provide new insights into the origin and spatial distribution of postzygotic mosaicism during normal human development.\n\nAuthor SummaryGenomic mosaicism led by postzygotic mutation is the major cause of cancers and many non-cancer developmental disorders. Theoretically, postzygotic mutations should be accumulated during the developmental process of healthy individuals, but the genome-wide characterization of postzygotic mosaicisms across many organ types of the same individual remained limited. In this study, we identified and validated two types of postzygotic mosaicism from the whole-genomes of 27 organs obtained from five healthy donors. We further found that the postzygotic mosaicisms arising during early embryogenesis and later clonal expansion events show distinct genomic patterns in mutation spectrum, replication timing, and chromatin status.

genomics

Genome Sequence of Indian Peacock Reveals the Peculiar Case of a Glittering Bird

The unique ornamental features and extreme sexual traits of Peacock have always intrigued the scientists. However, the genomic evidence to explain its phenotype are yet unknown. Thus, we report the first genome sequence and comparative analysis of peacock with the available high-quality genomes of chicken, turkey, duck, flycatcher and zebra finch. The candidate genes involved in early developmental pathways including TGF-{beta}, BMP, and Wnt signaling pathway, which are also involved in feather patterning, bone morphogenesis, and skeletal muscle development, showed signs of adaptive evolution and provided useful clues on the phenotype of peacock. The innate and adaptive immune components such as complement system and T-cell response also showed signs of adaptive evolution in peacock suggesting their possible role in building a robust immune system which is consistent with the between species predictions of Hamilton-Zuk hypothesis. This study provides novel genomic and evolutionary insights into the molecular understanding towards the phenotypic evolution of Indian peacock.

genomics

Utilizing random regression models for genomic prediction of a longitudinal trait derived from high-throughput phenotyping

The accessibility of high-throughput phenotyping platforms in both the greenhouse and field, as well as the relatively low cost of unmanned aerial vehicles, have provided researchers with an effective means to characterize large populations throughout the growing season. These longitudinal phenotypes can provide important insight into plant development and responses to the environment. Despite the growing use of these new phenotyping approaches in plant breeding, the use of genomic prediction models for longitudinal phenotypes is limited in major crop species. The objective of this study is to demonstrate the utility of random regression (RR) models using Legendre polynomials for genomic prediction of shoot growth trajectories in rice (Oryza sativa). An estimate of shoot biomass, projected shoot area (PSA), was recored over a period of 20 days for a panel of 357 diverse rice accessions using an image-based greenhouse phenotyping platform. A RR that included a fixed second-order Legendre polynomial, a random second-order Legendre polynomial for the additive genetic effect, a first-order Legendre polynomial for the environmental effect, and heterogeneous residual variances was used to model PSA trajectories. The utility of the RR model over a single time point (TP) approach, where PSA is fit at each time point independently, is shown through four prediction scenarios. In the first scenario, the RR and TP approaches were used to predict PSA for a set of lines lacking phenotypic data. The RR approach showed a 11.6% increase in prediction accuracy over the TP approach. Much of this improvement could be attributed to the greater additive genetic variance captured by the RR approach. The remaining scenarios focused forecasting future phenotypes using a subset of early time points for known lines with phenotypic data, as well new lines lacking phenotypic data. In all cases, PSA could be predicted with high accuracy (r: 0.79 to 0.89 and 0.55 to 0.58 for known and unknown lines, respectively). This study provides the first application of RR models for genomic prediction of a longitudinal trait in rice, and demonstrates that RR models can be effectively used to improve the accuracy of genomic prediction for complex traits compared to a TP approach.

genomics

Genome-wide discovery of epistatic loci affecting antibiotic resistance using evolutionary couplings

The analysis of whole genome sequencing data should, in theory, allow the discovery of interdependent loci that cause antibiotic resistance. In practice, however, identifying this epistasis remains a challenge as the vast number of possible interactions erodes statistical power. To solve this problem, we extend a method that has been successfully used to identify epistatic residues in proteins to infer genomic loci that are strongly coupled and associated with antibiotic resistance. Our method reduces the number of tests required for an epistatic genome-wide association study and increases the likelihood of identifying causal epistasis. We discovered 38 loci and 250 epistatic pairs that influence the dose needed to inhibit growth for five different antibiotics in 1,102 isolates of Neisseria gonorrhoeae that were confirmed in an independent dataset of 495 isolates. Many known resistance-affecting loci were recovered; however, the majority of loci occurred in unreported genes, including murE which was associated with cefixime. About half of the novel epistasis we report involved at least one locus previously associated with antibiotic resistance, including interactions between gyrA and parC associated with ciprofloxacin. Still, many combinations involved unreported loci and genes. Our work provides a systematic identification of epistasis pairs affecting antibiotic resistance in N. gonorrhoeae and a generalizable method for epistatic genome-wide association studies.

genomics

Candidatus Ornithobacterium hominis sp. nov.: insights gained from draft genomes obtained from nasopharyngeal swabs

Candidatus Ornithobacterium hominis sp. nov. represents a new member of the Flavobacteriaceae detected in 16S rRNA gene surveys from Southeast Asia, Africa and Australia. It frequently colonises the infant nasopharynx at high proportional abundance, and we demonstrate its presence in 42% of nasopharyngeal swabs from 12 month old children in the Maela refugee camp in Thailand. The species, a Gram negative bacillus, has not yet been cultured but the cells can be identified in mixed samples by fluorescent hybridisation. Here we report seven genomes assembled from metagenomic data, two to improved draft standard. The genomes are approximately 1.9Mb, sharing 62% average amino acid identity with the only other member of the genus, the bird pathogen Ornithobacterium rhinotracheale. The draft genomes encode multiple antibiotic resistance genes, competition factors, Flavobacterium johnsoniae-like gliding motility genes and a homolog of the Pasteurella multocida mitogenic toxin. Intra- and inter-host genome comparison suggests that colonisation with this bacterium is both persistent and strain exclusive.

genomics

FALCON-Phase: Integrating PacBio and Hi-C data for phased diploid genomes

Haplotype-resolved genome assemblies are important for understanding how combinations of variants impact phenotypes. These assemblies can be created in various ways, such as use of tissues that contain single-haplotype (haploid) genomes, or by co-sequencing of parental genomes, but these approaches can be impractical in many situations. We present FALCON-Phase, which integrates long-read sequencing data and ultra-long-range Hi-C chromatin interaction data of a diploid individual to create high-quality, phased diploid genome assemblies. The method was evaluated by application to three datasets, including human, cattle, and zebra finch, for which high-quality, fully haplotype resolved assemblies were available for benchmarking. Phasing algorithm accuracy was affected by heterozygosity of the individual sequenced, with higher accuracy for cattle and zebra finch (>97%) compared to human (82%). In addition, scaffolding with the same Hi-C chromatin contact data resulted in phased chromosome-scale scaffolds.

genomics

Whole genome linkage disequilibrium and effective population size in a coho salmon (Oncorhynchus kisutch) breeding population

The estimation of linkage disequilibrium between molecular markers within a population is critical when establishing the minimum number of markers required for association studies, genomic selection and for inferring historical events influencing different populations. This work aimed to evaluate the extent and decay of linkage disequilibrium in a coho salmon breeding population using ddRAD genomic markers.\n\nLinkage disequilibrium was estimated between a total of 7,505 SNPs found in 62 individuals (33 dams and 29 sires) from the breeding population. The makers encompass all 30 coho salmon chromosomes and comprise 1,655.19 Mb of the genome. The average density of markers per chromosome ranged from 3.45 to 6.11 per 1 Mbp. The minor allele frequency averaged 0.20 (with a range from 0.08 to 0.50). The overall average linkage disequilibrium among SNPs pairs measured as r2 was 0.054. The Average r2 value decreased with increasing physical distance, with values ranging from 0.37 to 0.054 at distances lower than 1 kb and up to 10 Mb, respectively. An r2 threshold of 0.1 was reached at distance of approximately 1.3 Mb. Chromosomes Okis05, Okis15 and Okis28 showed high levels of linkage disequilibrium (> 0.20 at distances lower than 1 Mb). Average r2 values were lower than 0.1 for all chromosomes at distances greater than 4 Mb. Linkage disequilibrium values suggest that whole genome association and selection studies could be performed using about 75,000 SNPs in aquaculture populations (depending on the trait under investigation). From the identified SNPs, an effective population size of 100 was estimated for the population 10 generation ago, and 1,000, for 139 generations ago.\n\nBased on the extent of r2 decay, we suggest that at least 75,000 SNPs would be necessary for an association mapping study. Over 100,000 SNPs would be necessary for a high power study, in the current coho salmon population.

genomics

Oxford Nanopore MinION genome sequencer: performance characteristics, optimised analysis workflow, phylogenetic analysis and prediction of antimicrobial resistance in Neisseria gonorrhoeae

Antimicrobial resistant (AMR) Neisseria gonorrhoeae strains are common and compromise gonorrhoea treatment internationally. Rapid identification and characterisation of AMR gonococcal strains could ensure appropriate and even personalised treatment, and support identification and investigation of gonorrhoea outbreaks in nearly real-time. Whole-genome sequencing is ideal for investigation of the emergence and dissemination of AMR determinants that predict AMR in the gonococcal population and spread of AMR strains in the human population. The novel, rapid and revolutionary long-read sequencer MinION is a small hand-held device that can generate bacterial genomes within one day. However, the accuracy of MinION reads has been suboptimal for many objectives and the MinION has not been evaluated for gonococci. In this first MinION study for gonococci, we show that MinION-derived sequences analysed with existing open-access, web-based sequence analysis tools are not sufficiently accurate to identify key gonococcal AMR determinants. Nevertheless, using an in house-developed CLC Genomics Workbench, we show that ONT-derived sequences can be used for accurate prediction of decreased susceptibility or resistance to recommended therapeutic antimicrobials. We also show that the ONT-derived sequences can be useful for rapid phylogenomic-based molecular epidemiological investigations, and, in hybrid assemblies with Illumina sequences, for producing contiguous assemblies and finished reference genomes.

genomics

genomeview - an extensible python-based genomics visualization engine

Visual inspection and analysis is integral to quality control, hypothesis generation, methods development and validation of genomic data. The richness and complexity of genomic data necessitates customized visualizations highlighting specific features of interest while hiding the often vast tide of irrelevant attributes. However, the majority of genome-visualization occurs either in general-purpose tools such as IGV (Robinson et al, 2011) or the UCSC Genome Browser (Kent et al, 2002) - which offer many options to adjust visualization parameters, but very little in the way of extensibility - or narrowly-focused tools aiming to solve a single visualization problem. Here, we present genomeview, a python-based visualization engine which is easy to extend and simple to integrate into existing analysis pipelines.

genomics

Efficient whole genome haplotyping and high-throughput single molecule phasing with barcode-linked reads

The future of human genomics is one that seeks to resolve the entirety of genetic variation through sequencing. The prospect of utilizing genomics for medical purposes require cost-efficient and accurate base calling, long-range haplotyping capability, and reliable calling of structural variants. Short read sequencing has lead the development towards such a future but has struggled to meet the latter two of these needs1. To address this limitation, we developed a technology that preserves the molecular origin of short sequencing reads, with an insignificant increase to sequencing costs. We demonstrate a novel library preparation method which enables whole genome haplotyping, long-range phasing of single DNA molecules, and de novo genome assembly through barcode-linked reads (BLR). Millions of random barcodes are used to reconstruct megabase-scale phase blocks and call structural variants. We also highlight the versatility of our technology by generating libraries from different organisms using only picograms to nanograms of input material.

genomics

Up, down, and out: next generation libraries for genome-wide CRISPRa, CRISPRi, and CRISPR-Cas9 knockout genetic screens

Advances in CRISPR-Cas9 technology have enabled the flexible modulation of gene expression at large scale. In particular, the creation of genome-wide libraries for CRISPR knockout (CRISPRko), CRISPR interference (CRISPRi), and CRISPR activation (CRISPRa) has allowed gene function to be systematically interrogated. Here, we evaluate numerous CRISPRko libraries and show that our recently-described CRISPRko library (Brunello) is more effective than previously published libraries at distinguishing essential and non-essential genes, providing approximately the same perturbation-level performance improvement over GeCKO libraries as GeCKO provided over RNAi. Additionally, we developed genome-wide libraries for CRISPRi (Dolcetto) and CRISPRa (Calabrese). Negative selection screens showed that Dolcetto substantially outperforms existing CRISPRi libraries with fewer sgRNAs per gene and achieves comparable performance to CRISPRko in the detection of gold-standard essential genes. We also conducted positive selection CRISPRa screens and show that Calabrese outperforms the SAM library approach at detecting vemurafenib resistance genes. We further compare CRISPRa to genome-scale libraries of open reading frames (ORFs). Together, these libraries represent a suite of genome-wide tools to efficiently interrogate gene function with multiple modalities.tracr

genomics

Genome analysis of the unicellular eukaryote Euplotes vannus provides insights into mating type determination and tolerance to environmental stresses

As a model organism in studies of cell and environmental biology, the free-living and cosmopolitan ciliated protist Euplotes vannus has more than ten mating types (sexes) and shows strong resistance to environmental stresses. However, the molecular basis of its sex determination mechanism and how the cell responds to stress remain largely unknown. Here we report a combined analysis of de novo assembled high-quality macronucleus (MAC; i.e. somatic) genome and partial micronucleus (MIC; i.e. germline) genome of Euplotes vannus. Furthermore, MAC genomic and transcriptomic data from several mating types of E. vannus were investigated and gene expression levels were profiled under different environmental stresses, including nutrient scarcity, extreme temperature, salinity and the presence of free ammonia. We found that E. vannus, which possesses gene-sized nanochromosomes in its MAC, shares a similar pattern on frameshifting and stop codon usage as Euplotes octocarinatus and may be undergoing incipient sympatric speciation with Euplotes crassus. Somatic pheromone loci of E. vannus are generated from programmed DNA rearrangements of multiple germline macronuclear destined sequences (MDS) and the mating types of E. vannus are distinguished by the different combinations of pheromone loci instead of possessing mating type-specific genes. Lastly, we linked the resilience to environmental temperature change to the evolved loss of temperature stress-sensitive regulatory regions of HSP70 gene in E. vannus. Together, the genome resources generated in this study, which are available online at Euplotes vannus DB (http://evan.ciliate.org), provide new evidence for sex determination mechanism in eukaryotes and common pheromone-mediated cell-cell signaling and cross-mating.

genomics