Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 937 records · Page 52Linked to original sources

The Population Genomics Of Archaeological Transition In West Iberia

We analyse new genomic data (0.05-2.95x) from 14 ancient individuals from Portugal distributed from the Middle Neolithic (4200-3500 BC) to the Middle Bronze Age (1740-1430 BC) and impute genomewide diploid genotypes in these together with published ancient Eurasians. While discontinuity is evident in the transition to agriculture across the region, sensitive haplotype-based analyses suggest a significant degree of local hunter-gatherer contribution to later Iberian Neolithic populations. A more subtle genetic influx is also apparent in the Bronze Age, detectable from analyses including haplotype sharing with both ancient and modern genomes, D-statistics and Y-chromosome lineages. However, the limited nature of this introgression contrasts with the major Steppe migration turnovers within third Millennium northern Europe and echoes the survival of non-Indo-European language in Iberia. Changes in genomic estimates of individual height across Europe are also associated with these major cultural transitions, and ancestral components continue to correlate with modern differences in stature.\n\nAuthor SummaryRecent ancient DNA work has demonstrated the significant genetic impact of mass migrations from the Steppe into Central and Northern Europe during the transition from the Neolithic to the Bronze Age. In Iberia, archaeological change at the level of material culture and funerary rituals has been reported during this period, however, the genetic impact associated with this cultural transformation has not yet been estimated. In order to investigate this, we sequence Neolithic and Bronze Age samples from Portugal, which we compare to other ancient and present-day individuals. Genome-wide imputation of a large dataset of ancient samples enabled sensitive methods for detecting population structure and selection in ancient samples. We revealed subtle genetic differentiation between the Portuguese Neolithic and Bronze Age samples suggesting a markedly reduced influx in Iberia compared to other European regions. Furthermore, we predict individual height in ancients, suggesting that stature was reduced in the Neolithic and affected by subsequent admixtures. Lastly, we examine signatures of strong selection in important traits and the timing of their origins.

genomics

Mapping The Malaria Parasite Drug-Able Genome Using In Vitro Evolution And Chemogenomics

Chemogenetic characterization through in vitro evolution combined with whole genome analysis is a powerful tool to discover novel antimalarial drug targets and identify drug resistance genes. Our comprehensive genome analysis of 262 Plasmodium falciparum parasites treated with 37 diverse compounds reveals how the parasite evolves to evade the action of small molecule growth inhibitors. This detailed data set revealed 159 gene amplifications and 148 nonsynonymous changes in 83 genes which developed during resistance acquisition. Using a new algorithm, we show that gene amplifications contribute to 1/3 of drug resistance acquisition events. In addition to confirming known multidrug resistance mechanisms, we discovered novel multidrug resistance genes. Furthermore, we identified promising new drug target-inhibitor pairs to advance the malaria elimination campaign, including: thymidylate synthase and a benzoquinazolinone, farnesyltransferase and a pyrimidinedione, and a dipeptidylpeptidase and an arylurea. This deep exploration of the P. falciparum resistome and drug-able genome will guide future drug discovery and structural biology efforts, while also advancing our understanding of resistance mechanisms of the deadliest malaria parasite.\n\nOne Sentence SummaryWhole genome sequencing reveals how Plasmodium falciparum evolves resistance to diverse compounds and identifies new antimalarial drug targets.

genomics

A High HIV-1 Strain Variability in London, UK, Revealed by Full-Genome Analysis: Results from the ICONIC Project

Background & MethodsThe ICONIC project has developed an automated high-throughput pipeline to generate HIV nearly full-length genomes (NFLG, i.e. from gag to nef) from next-generation sequencing (NGS) data. The pipeline was applied to 420 HIV samples collected at University College London Hospital and Barts Health NHS Trust (London) and sequenced using an Illumina MiSeq at the Wellcome Trust Sanger Institute (Cambridge). Consensus genomes were generated and subtyped using COMET, and unique recombinants were studied with jpHMM and SimPlot. Maximum-likelihood phylogenetic trees were constructed using RAxML to identify transmission networks using the Cluster Picker.\n\nResultsThe pipeline generated sequences of at least 1Kb of length (median=7.4Kb) for 375 out of the 420 samples (89%), with 174 (46.4%) being NFLG. A total of 365 sequences (169 of them NFLG) corresponded to unique subjects and were included in the down-stream analyses. The most frequent HIV subtypes were B (n=149, 40.8%) and C (n=77, 21.1%) and the circulating recombinant form CRF02_AG (n=32, 8.8%). We found 14 different CRFs (n=66, 18.1%) and multiple URFs (n=32, 8.8%) that involved recombination between 12 different subtypes/CRFs. The most frequent URFs were B/CRF01_AE (4 cases) and A1/D, B/C, and B/CRF02_AG (3 cases each). Most URFs (19/26, 73%) lacked breakpoints in the PR+RT pol region, rendering them undetectable if only that was sequenced. Twelve (37.5%) of the URFs could have emerged within the UK, whereas the rest were probably imported from sub-Saharan Africa, South East Asia and South America. For 2 URFs we found highly similar pol sequences circulating in the UK. We detected 31 phylogenetic clusters using the full dataset: 25 pairs (mostly subtypes B and C), 4 triplets and 2 quadruplets. Some of these were not consistent across different genes due to inter- and intra-subtype recombination. Clusters involved 70 sequences, 19.2% of the dataset.\n\nConclusionsThe initial analysis of genome sequences detected substantial hidden variability in the London HIV epidemic. Analysing full genome sequences, as opposed to only PR+RT, identified previously undetected recombinants. It provided a more reliable description of CRFs (that would be otherwise misclassified) and transmission clusters.

genomics

Genome Sequencing Links Persistent Outbreak Of Legionellosis In Sydney To An Emerging Clone Of Legionella pneumophila ST211

The city of Sydney, Australia, experienced a persistent outbreak of Legionella pneumophila serogroup 1 (Lp1) pneumonia in 2016. To elucidate the source and bring the outbreak to a close we examined the genomes of clinical and environmental Lp1 isolates recovered over 7 weeks. A total of 48 isolates from patients and cooling towers were sequenced and compared using SNP-based, core-genome MLST and pangenome approaches. All three methods confirmed phylogenetic relatedness between isolates associated with outbreaks in the Central Business District (March and May) and Suburb 1. These isolates were designated \"Main cluster\" and consisted of isolates from two patients from the CBD March outbreak, one patient and one tower isolate from Suburb 1 and isolates from two cooling towers and three patients from the CDB May outbreak. All main cluster isolates were sequence type ST211 which has only ever been reported in Canada. Significantly, pangenome analysis identified mobile genetic elements containing a unique T4ASS that was specific to the main cluster and co-circulating clinical strains, suggesting a potential mechanism for increased fitness and persistence of the outbreak clone. Genome sequencing was key in deciphering the environmental sources of infection among the spatially and temporally coinciding cases of legionellosis in this highly populated urban setting. Further, the discovery of a unique T4ASS emphasises the potential contribution of genome recombination in the emergence of successful Lp1 clones.

genomics

Discovery Of The First Genome-Wide Significant Risk Loci For ADHD

Attention-Deficit/Hyperactivity Disorder (ADHD) is a highly heritable childhood behavioral disorder affecting 5% of school-age children and 2.5% of adults. Common genetic variants contribute substantially to ADHD susceptibility, but no individual variants have been robustly associated with ADHD. We report a genome-wide association meta-analysis of 20,183 ADHD cases and 35,191 controls that identifies variants surpassing genome-wide significance in 12 independent loci, revealing new and important information on the underlying biology of ADHD. Associations are enriched in evolutionarily constrained genomic regions and loss-of-function intolerant genes, as well as around brain-expressed regulatory marks. These findings, based on clinical interviews and/or medical records are supported by additional analyses of a self-reported ADHD sample and a study of quantitative measures of ADHD symptoms in the population. Meta-analyzing these data with our primary scan yielded a total of 16 genome-wide significant loci. The results support the hypothesis that clinical diagnosis of ADHD is an extreme expression of one or more continuous heritable traits.

genetics

Indexcov: fast coverage quality control for whole-genome sequencing

The BAM1 and CRAM2 formats provide a supplementary linear index that facilitates rapid access to sequence alignments in arbitrary genomic regions. Comparing consecutive entries in a BAM or CRAM index allows one to infer the number of alignment records per genomic region for use as an effective proxy of sequence depth in each genomic region. Based on these properties, we have developed indexcov, an efficient estimator of whole-genome sequencing coverage to rapidly identify samples with aberrant coverage profiles, reveal large scale chromosomal anomalies, recognize potential batch effects, and infer the sex of a sample. Indexcov is available at: https://github.com/brentp/goleft under the MIT license.

genomics

Draft genome of the Reindeer (Rangifer tarandus)

AbstractO_ST_ABSBackgroundC_ST_ABSReindeer (Rangifer tarandus) is the only fully domesticated species in the Cervidae family, and is the only cervid with a circumpolar distribution. Unlike all other cervids, female reindeer regularly grow cranial appendages (antlers, the defining characteristics of cervids), as well as males. Moreover, reindeer milk contains more protein and less lactose than bovids milk. A high quality reference genome of this specie will assist efforts to elucidate these and other important features in the reindeer.\n\nFindingsWe obtained 723.2 Gb (Gigabase) of raw reads by an Illumina Hiseq 4000 platform, and a 2.64 Gb final assembly, representing 95.7% of the estimated genome (2.76 Gb according to k-mer analysis), including 92.6% of expected genes according to BUSCO analysis. The contig N50 and scaffold N50 sizes were 89.7 kilo base (kb) and 0.94 mega base (Mb), respectively. We annotated 21,555 protein-coding genes and 1.07 Gb of repetitive sequences by de novo and homology-based prediction. Homology-based searches detected 159 rRNA, 547 miRNA, 1,339 snRNA and 863 tRNA sequences in the genome of R. tarandus. The divergence time between R. tarandus, and ancestors of Bos taurus and Capra hircus, is estimated to be 29.55 million years ago (Mya).\n\nConclusionsOur results provide the first high-quality reference genome for the reindeer, and a valuable resource for studying evolution, domestication and other unusual characteristics of the reindeer.

genomics

High-quality genome assemblies uncover caste-specific long non-coding RNAs in ants

Ants are an emerging model system for neuroepigenetics, as embryos with virtually identical genomes develop into different adult castes that display strikingly different physiology, morphology, and behavior. Although a number of ant genomes have been sequenced to date, their draft quality is an obstacle to sophisticated analyses of epigenetic gene regulation. Using long reads generated with Pacific Biosystem single molecule real time sequencing, we have reassembled de novo high-quality genomes for two ant species: Camponotus floridanus and Harpegnathos saltator. The long reads allowed us to span large repetitive regions and join sequences previously found in separate scaffolds, leading to comprehensive and accurate protein-coding annotations that facilitated the identification of a Gp-9-like gene as differentially expressed in Harpegnathos castes. The new assemblies also enabled us to annotate long non-coding RNAs for the first time in ants, revealing several that were specifically expressed during Harpegnathos development and in the brains of different castes. These upgraded genomes, along with the new coding and non-coding annotations, will aid future efforts to identify epigenetic mechanisms of phenotypic and behavioral plasticity in ants.

genomics

Genomic footprints of activated telomere maintenance mechanisms in cancer

Cancers require telomere maintenance mechanisms for unlimited replicative potential. We dissected whole-genome sequencing data of over 2,500 matched tumor-control samples from 36 different tumor types to characterize the genomic footprints of these mechanisms. While the telomere content of tumors with ATRX or DAXX mutations (ATRX/DAXXtrunc) was increased, tumors with TERT modifications showed a moderate decrease of telomere content. One quarter of all tumor samples contained somatic integrations of telomeric sequences into non-telomeric DNA. With 80% prevalence, ATRX/DAXXtrunc tumors display a 3-fold enrichment of telomere insertions. A systematic analysis of telomere composition identified aberrant telomere variant repeat (TVR) distribution as a genomic marker of ATRX/DAXXtrunc tumors. In this clinically relevant subgroup, singleton TTCGGG and TTTGGG TVRs (previously undescribed) were significantly enriched or depleted, respectively. Overall, our findings provide new insight into the recurrent genomic alterations that are associated with the establishment of different telomere maintenance mechanisms in cancer.

genomics

Assembly of hundreds of microbial genomes from the cow rumen reveals novel microbial species encoding enzymes with roles in carbohydrate metabolism

The cow rumen is a specialised organ adapted for the efficient breakdown of plant material into energy and nutrients, and it is the rumen microbiome that encodes the enzymes responsible. Many of these enzymes are of huge industrial interest. Despite this, rumen microbes are under-represented in the public databases. Here we present 220 high quality bacterial and archaeal genomes assembled directly from 768 gigabases of rumen metagenomic sequence data. Comparative analysis with current publicly available genomes reveals that the majority of these represent previously unsequenced strains and species of bacteria and archaea. The genomes contain over 13,000 proteins predicted to be involved in carbohydrate metabolism, over 90% of which do not have a good match in the public databases. Inclusion of the 220 genomes presented here improves metagenomic read classification by 2-3-fold, both in our data and in other publicly available rumen datasets. This release improves the coverage of rumen microbes in the public databases, and represents a hugely valuable resource for biomass-degrading enzyme discovery and studies of the rumen microbiome

genomics

Spatial and temporal distribution of genome divergence among California populations of Aedes aegypti.

In the summer of 2013, Aedes aegypti Linnaeus was first detected in three cities in central California (Clovis, Madera and Menlo Park). It has now been detected in multiple locations in central and southern CA as far south as San Diego and Imperial Counties. A number of published reports suggest that CA populations have been established from multiple independent introductions. Here we report the first population genomics analyses of Ae. aegypti based on individual, field collected whole genome sequences. We analyzed 46 Ae. aegypti genomes to establish genetic relationships among populations from sites in California, Florida and South Africa. We identified 3 major genetic clusters within California; one that includes all sample sites in the southern part of the state (South of Tehachapi mountain range) plus the town of Exeter in central California and two additional clusters in central California. A lack of concordance between mitochondrial and nuclear genealogies suggests that the three founding populations were polymorphic for two main mitochondrial haplotypes prior to being introduced to California. One of these has been lost in the Clovis populations, possibly by a founder effect. Genome-wide comparisons indicate extensive differentiation between genetic clusters. Our observations support recent introductions of Ae. aegypti into California from multiple, genetically diverged source populations. Our data reveal signs of hybridization among diverged populations within CA. Genetic markers identified in this study will be of great value in pursuing classical population genetic studies which require larger sample sizes.

genomics

CTCF mediated genome architecture regulates the dosage of mitotically stable mono-allelic expression of autosomal genes

Mammalian genomes exhibit widespread mono-allelic expression of autosomal genes. However, the mechanistic insight that allows specific expression of one allele remains enigmatic. Here, we present evidence that the linear and the three dimensional architectures of the genome ascribe the appropriate framework that guides the mono-allelic expression of genes. We show that: 1) mono-allelically expressed genes are assorted into genomic domains that are insulated from domains of bi-allelically expressed genes through CTCF mediated chromatin loops; 2) evolutionary and cell-type specific gain and loss of mono-allelic expression coincide respectively with the gain and loss of chromatin insulator sites; 3) dosage of mono- allelically expressed genes is more sensitive to loss of chromatin insulationn associated with CTCF depletion as compared to bi-allelically expressed genes; 4) distinct susceptibility of mono- and bi-allelically expressed genes to CTCF depletion can be attributed to distinct functional roles of CTCF around these genes. Altogether, our observations highlight a general topological framework for the mono-allelic expression of genes, wherein the alleles are insulated from the spatial interference of chromatin and transcriptional states from neighbouring bi-allelic domains via CTCF mediated chromatin loops. The study also suggests that the three-dimensional genome organization might have evolved under the constraint to mitigate the fluctuations in the dosage of mono-allelically expressed genes, which otherwise are dosage sensitive.

genomics

Insights into regeneration from the genome, transcriptome and metagenome analysis of Eisenia fetida

Earthworms show a wide spectrum of regenerative potential with certain species like Eisenia fetida capable of regenerating more than two-thirds of their body while other closely related species, such as Paranais litoralis seem to have lost this ability. Earthworms belong to the phylum annelida, in which the genomes of the marine oligochaete Capitella telata, and the freshwater leech Helobdella robusta have been sequenced and studied. The terrestrial annelids, in spite of their ecological relevance and unique biochemical repertoire, are represented by a single rough genome draft of Eisenia fetida (North American isolate), which suggested that extensive duplications have led to a large number of HOX genes in this annelid. Herein, we report the draft genome sequence of Eisenia fetida (Indian isolate), a terrestrial redworm widely used for vermicomposting assembled using short reads and mate-pair reads. An in-depth analysis of the miRNome of the worm, showed that many miRNA gene families have also undergone extensive duplications. Genes for several important proteins such as sialidases and neurotrophins were identified by RNA sequencing of tissue samples. We also used de novo assembled RNA-Seq data to identify genes that are differentially expressed during regeneration, both in the newly regenerating cells and in the adjacent tissue. Sox4, a master regulator of TGF-beta induced epithelial-mesenchymal transition was induced in the newly regenerated tissue. The regeneration of the ventral nerve cord was also accompanied by the induction of nerve growth factor and neurofilament genes. The metagenome of the worm, characterized using 16S rRNA sequencing, revealed the identity of several bacterial species that reside in the nephridia of the worm. Comparison of the bodywall and cocoon metagenomes showed exclusion of hereditary symbionts in the regenerated tissue. In summary, we present extensive genome, transcriptome and metagenome data to establish the transcriptome and metagenome dynamics during regeneration.

genomics

RNA-seq highlights parallel and contrasting patterns in the evolution of the nuclear genome of holo-mycoheterotrophic plants

* While photosynthesis is the most notable trait of plants, several lineages of plants (so-called holo-heterotrophs) have adapted to obtain organic compounds from other sources. The switch to heterotrophy leads to profound changes at the morphological, physiological and genomic levels.\n\n* Here, we characterize the transcriptomes of three species representing two lineages of mycoheterotrophic plants: orchids (Epipogium aphyllum and Epipogium roseum) and Ericaceae (Hypopitys monotropa). Comparative analysis is used to highlight the parallelism between distantly related holo-heterotrophic plants.\n\n* In both lineages, we observed genome-wide elimination of nuclear genes that encode proteins related to photosynthesis, while systems associated with protein import to plastids as well as plastid transcription and translation remain active. Genes encoding components of plastid ribosomes that have been lost from the plastid genomes have not been transferred to the nuclear genomes; instead, some of the encoded proteins have been substituted by homologs. The nuclear genes of both Epipogium species accumulated mutations twice as rapidly as their photosynthetic relatives; in contrast, no increase in the substitution rate was observed in H.monotropa.\n\n* Holo-heterotrophy leads to profound changes in nuclear gene content. The observed increase in the rate of nucleotide substitutions is lineage specific, rather than a universal phenomenon among non-photosynthetic plants.

genomics

Project Dhaka: Variational Autoencoder for Unmasking Tumor Heterogeneity from Single Cell Genomic Data

Intra-tumor heterogeneity is one of the key confounding factors in deciphering tumor evolution. Malignant cells will have variations in their gene expression, copy numbers, and mutation even when coming from a single tumor. Single cell sequencing of tumor cells is of paramount importance for unmasking the underlying the tumor heterogeneity. However extracting features from the single cell genomic data coherent with the underlying biology is computationally challenging, given the extremely noisy and sparse nature of the data. Here we are proposing Dhaka a variational autoencoder based single cell analysis tool to transform genomic data to a latent encoded feature space that is more efficient in differentiating between the hidden tumor subpopulations. This technique is generalized across different types of genomic data such as copy number variation from DNA sequencing and gene expression data from RNA sequencing. We have tested the method on two gene expression datasets having 4K to 6K tumor cells and two copy number variation datasets having 250 to 260 tumor cells. Analysis of the encoded feature space revealed sub-populations of cells bearing distinct genomic signatures and the evolutionary relationship between them, which other existing feature transformation methods like t-SNE and PCA fail to do.

genomics

The genome of Trichoplusia ni, an agricultural pest and novel model for small RNA biology

The cabbage looper, Trichoplusia ni (Lepidoptera: Noctuidae), is a destructive insect pest that feeds on a wide range of plants. The High Five cell line (Hi5), originally derived from T. ni ovaries, is often used for efficient expression of recombinant proteins. Here, we report a draft assembly of the 368.2 Mb T. ni genome, with 90.6% of all bases assigned to one of its 28 chromosomes and predicted 14,037 predicted protein-coding genes. Manual curation of gene families involved in chemoreception and detoxification reveals T. ni-specific gene expansions that may explain its widespread distribution and rapid adaptation to insecticides. Using male and female genome sequences, we define Z-linked and repeat-rich W-linked sequences. Transcriptome and small RNA data from T. ni thorax, ovary, testis, and Hi5 cells reveal distinct expression profiles for 295 microRNA- and >393 piRNA-producing loci, as well as 39 genes encoding core small RNA pathway proteins. siRNAs target both endogenous transposons and the exogenous TNCL virus. Surprisingly, T. ni siRNAs are not 2{acute}-O-methylated. Five piRNA-producing loci account for 34.9% piRNAs in the ovary, 49.3% piRNAs in the testis, and 44.0% piRNAs in Hi5 cells. Nearly all of the W chromosome is devoted to piRNA production: >76.0% of bases in the assembled W produce piRNAs in ovary. To enable use of the T. ni germline-derived Hi5 cell line as a model system, we have established efficient genome editing and single-cell cloning protocols. Taken together, the T. ni genome provides insights into pest control and allows Hi5 cells to become a new tool for studying small RNAs ex vivo.

genomics

Comparative analysis of the genomes of Stylophora pistillata and Acropora digitifera provides evidence for extensive differences between species of corals

Stony corals form the foundation of coral reef ecosystems. Their phylogeny is characterized by a deep evolutionary divergence that separates corals into a robust and complex clade dating back to at least 245 mya. However, the genomic consequences and clade-specific evolution remain unexplored. In this study we have produced the genome of a robust coral, Stylophora pistillata, and compared it to the available genome of a complex coral, Acropora digitifera. We conducted a fine-scale gene-based analysis focusing on ortholog groups. Among the core set of conserved proteins, we found an emphasis on processes related to the cnidarian-dinoflagellate symbiosis. Similarly, genes associated with the algal symbiosis were also independently expanded in both species, but both corals diverged on the identity of ortholog groups expanded, and we found uneven expansions in genes associated with innate immunity and stress response. Our analyses demonstrate that coral genomes can be surprisingly disparate. Importantly, if the patterns elucidated here are representative of differences between corals from the robust and complex clade, the ability of a coral to respond to climate change may be dependent on its clade association.

genomics

Population-based analysis of ocular Chlamydia trachomatis in trachoma-endemic West African communities identifies genomic markers of disease severity

Chlamydia trachomatis (Ct) is the most common infectious cause of blindness and bacterial sexually transmitted infection worldwide. Using Ct whole genome sequences obtained directly from conjunctival swabs, we studied Ct genomic diversity and associations between Ct genetic polymorphisms with ocular localization and disease severity in a treatment-naive trachoma-endemic population in Guinea Bissau, West Africa. All sequences fall within the T2 ocular clade phylogenetically. This is consistent with the presence of the characteristic deletion in trpA resulting in a truncated non-functional protein and the ocular tyrosine repeat regions present in tarP associated with ocular tissue localization. We have identified twenty-one Ct non-synonymous single nucleotide polymorphisms (SNPs) associated with ocular localization, including SNPs within pmpD (OR=4.07, p*=0.001) and tarP (OR=0.34, p*=0.009). Eight SNPs associated with disease severity were found in yjfH (rlmB) (OR=0.13, p*=0.037), CTA0273 (OR=0.12, p*=0.027), trmD (OR=0.12, p*=0.032), CTA0744 (OR=0.12, p*=0.041), glgA (OR=0.10, p*=0.026), alaS (OR=0.10, p*=0.032), pmpE (OR=0.08, p*=0.001) and the intergenic region CTA0744-CTA0745 (OR=0.13, p*=0.043). This study demonstrates the extent of genomic diversity within a naturally circulating population of ocular Ct, and the first to describe novel genomic associations with disease severity. These findings direct investigation of host-pathogen interactions that may be important in ocular Ct pathogenesis and disease transmission.

genomics