Search bioRxivSearch

SEARCH · Search bioRxiv

Results for “Genomics”

Search indexed bioRxiv preprints in genomics, neuroscience, cell biology and bioinformatics. Read source abstracts and check manuscript versions; preprints are not peer reviewed.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,189 records · Page 66Linked to original sources

Medaka population genome structure and demographic history unveiled via Genotyping-by-Sequencing

Medaka is a model organism in medicine, genetics, developmental biology and population genetics. Lab stocks composed of more than 100 local wild populations are available for research in these fields. Thus, medaka represents a potentially excellent bioresource for screening disease-risk- and adaptation-related genes in genome-wide association studies. Although the genetic population structure should be known before performing such an analysis, a comprehensive study on the genome-wide diversity of wild medaka populations has not been performed. Here, we performed genotyping-by-sequencing (GBS) for 81 and 12 medakas captured from a bioresource and the wild, respectively. Based on the GBS data, we evaluated the genetic population structure and estimated the demographic parameters using an approximate Bayesian computation (ABC) framework. The autosomal data confirmed that there were substantial differences between local populations and supported our previously proposed hypothesis on medaka dispersal based on mitochondrial genome (mtDNA) data. A new finding was that a local group that was thought to be a hybrid between the northern and the southern Japanese groups was actually a sister group of the northern Japanese group. Thus, this paper presents the first population-genomic study of medaka and reveals its population structure and history based on autosomal diversity.

evolutionary biology

Real-time search of all bacterial and viral genomic data

Genome sequencing of pathogens is now ubiquitous in microbiology, and the sequence archives are effectively no longer searchable for arbitrary sequences. Furthermore, the exponential increase of these archives is likely to be further spurred by automated diagnostics. To unlock their use for scientific research and real-time surveillance we have combined knowledge about bacterial genetic variation with ideas used in web-search, to build a DNA search engine for microbial data that can grow incrementally. We indexed the complete global corpus of bacterial and viral whole genome sequence data (447,833 genomes), using four orders of magnitude less storage than previous methods. The method allows future scaling to millions of genomes. This renders the global archive accessible to sequence search, which we demonstrate with three applications: ultra-fast search for resistance genes MCR1-3, analysis of host-range for 2827 plasmids, and quantification of the rise of antibiotic resistance prevalence in the sequence archives.

bioinformatics

Cultivation and genomic analysis of Candidatus Nitrosocaldus islandicus, a novel obligately thermophilic ammonia-oxidizing Thaumarchaeon

Ammonia-oxidizing archaea (AOA) within the phylum Thaumarchaea are the only known aerobic ammonia oxidizers in geothermal environments. Although molecular data indicate the presence of phylogenetically diverse AOA from the Nitrosocaldus clade, group 1.1b and group 1.1a Thaumarchaea in terrestrial high-temperature habitats, only one enrichment culture of an AOA thriving above 50 {degrees}C has been reported and functionally analyzed. In this study, we physiologically and genomically characterized a novel Thaumarchaeon from the deep-branching Nitrosocaldaceae family of which we have obtained a high ([~]85 %) enrichment from biofilm of an Icelandic hot spring (73 {degrees}C). This AOA, which we provisionally refer to as \"Candidatus Nitrosocaldus islandicus\", is an obligately thermophilic, aerobic chemolithoautotrophic ammonia oxidizer, which stoichiometrically converts ammonia to nitrite at temperatures between 50 {degrees}C and 70 {degrees}C. Ca. N. islandicus encodes the expected repertoire of enzymes proposed to be required for archaeal ammonia oxidation, but unexpectedly lacks a nirK gene and also possesses no identifiable other enzyme for nitric oxide (NO) generation. Nevertheless, ammonia oxidation by this AOA appears to be NO-dependent as Ca. N. islandicus is, like all other tested AOA, inhibited by the addition of an NO scavenger. Furthermore, comparative genomics revealed that Ca. N. islandicus has the potential for aromatic amino acid fermentation as its genome encodes an indolepyruvate oxidoreductase (iorAB) as well as a type 3b hydrogenase, which are not present in any other sequenced AOA. A further surprising genomic feature of this thermophilic ammonia oxidizer is the absence of DNA polymerase D genes - one of the predominant replicative DNA polymerases in all other ammonia-oxidizing Thaumarchaea. Collectively, our findings suggest that metabolic versatility and DNA replication might differ substantially between obligately thermophilic and other AOA.

microbiology

The structure of the influenza A virus genome

Influenza A viruses (IAVs) are segmented single-stranded negative sense RNA viruses that constitute a major threat to human health. The IAV genome consists of eight RNA segments contained in separate viral ribonucleoprotein complexes (vRNPs) that are packaged together into a single virus particle1,2. While IAVs are generally considered to have an unstructured single-stranded genome, it has also been suggested that secondary RNA structures are required for selective packaging of the eight vRNPs into each virus particle3,4. Here, we employ high-throughput sequencing approaches to map both the intra and intersegment RNA interactions inside influenza virions. Our data demonstrate that a redundant network of RNA-RNA interactions is required for vRNP packaging and virus growth. Furthermore, the data demonstrate that IAVs have a much more structured genome than previously thought and the redundancy of RNA interactions between the different vRNPs explains how IAVs maintain the potential for reassortment between different strains, while also retaining packaging selectivity. Our study establishes a framework towards further work into IAV RNA structure and vRNP packaging, which will lead to better models for predicting the emergence of new pandemic influenza strains and will facilitate the development of antivirals specifically targeting genome assembly.

microbiology

The essential genome of Escherichia coli K-12

Transposon-Directed Insertion-site Sequencing (TraDIS) is a high-throughput method coupling transposon mutagenesis with short-fragment DNA sequencing. It is commonly used to identify essential genes. Single gene deletion libraries are considered the gold standard for identifying essential genes. Currently, the TraDIS method has not been benchmarked against such libraries and therefore it remains unclear whether the two methodologies are comparable. To address this, a high density transposon library was constructed in Escherichia coli K-12. Essential genes predicted from sequencing of this library were compared to existing essential gene databases. To decrease false positive identification of essential gene candidates, statistical data analysis included corrections for both gene length and genome length. Through this analysis new essential genes and genes previously incorrectly designated as essential were identified. We show that manual analysis of TraDIS data reveals novel features that would not have been detected by statistical analysis alone. Examples include short essential regions within genes, orientation-dependent effects and fine resolution identification of genome and protein features. Recognition of these insertion profiles in transposon mutagenesis datasets will assist genome annotation of less well characterized genomes and provides new insights into bacterial physiology and biochemistry.\n\nIMPORTANCEIncentives to define lists of genes that are essential for bacterial survival include the identification of potential targets for antibacterial drug development, genes required for rapid growth for exploitation in biotechnology, and discovery of new biochemical pathways. To identify essential genes in E. coli, we constructed a very high density transposon mutant library. Initial automated analysis of the resulting data revealed many discrepancies when compared to the literature. We now report more extensive statistical analysis supported by both literature searches and detailed inspection of high density TraDIS sequencing data for each putative essential gene for the model laboratory organism, Escherichia coli. This paper is important because it provides a better understanding of the essential genes of E. coli, reveals the limitations of relying on automated analysis alone and a provides new standard for the analysis of TraDIS data.

microbiology

Alignment-Free Approaches Predict Novel Nuclear Mitochondrial Segments (NUMTs) in the Human Genome

The nuclear human genome harbors sequences of mitochondrial origin, indicating an ancestral transfer of DNA from the mitogenome. Several Nuclear Mitochondrial Segments (NUMTs) have been detected by alignment-based sequence similarity search, as implemented in the Basic Local Alignment Search Tool (BLAST). Identifying NUMTs is important for the comprehensive annotation and understanding of the human genome. Here we explore the possibility of detecting NUMTs in the human genome by alignment-free sequence similarity search, such as k-mers (k-tuples, k-grams, oligos of length k) distributions. We find that when k=6 or larger, the k-mer approach and BLAST search produce almost identical results, e.g., detect the same set of NUMTs longer than 3kb. However, when k=5 or k=4, certain signals are only detected by the alignment-free approach, and these may indicate yet unrecognized, and potentially more ancestral NUMTs. We introduce a \"Manhattan plot\" style representation of NUMT predictions across the genome, which are calculated based on the reciprocal of the Jensen-Shannon divergence between the nuclear and mitochondrial k-mer frequencies. The further inspection of the k-mer-based NUMT predictions however shows that most of them contain long-terminal-repeat (LTR) annotations, whereas BLAST-based NUMT predictions do not. Thus, similarity of the mitogenome to LTR sequences is recognized, which we validate by finding the mitochondrial k-mer distribution closer to those for transposable sequences and specifically, close to some types of LTR.

bioinformatics

Population structure of the Brachypodium species complex and genome wide association of agronomic traits in response to climate.

The development of model systems requires a detailed assessment of standing genetic variation across natural populations. The Brachypodium species complex has been promoted as a plant model for grass genomics with translational to small grain and biomass crops. To capture the genetic diversity within this species complex, thousands of Brachypodium accessions from around the globe were collected and sequenced using genotyping by sequencing (GBS). Overall, 1,897 samples were classified into two diploid or allopolyploid species and then further grouped into distinct inbred genotypes. A core set of diverse B. distachyon diploid lines were selected for whole genome sequencing and high resolution phenotyping. Genome-wide association studies across simulated seasonal environments was used to identify candidate genes and pathways tied to key life history and agronomic traits under current and future climatic conditions. A total of 8, 22 and 47 QTLs were identified for flowering time, early vigour and energy traits, respectively. Overall, the results highlight the genomic structure of the Brachypodium species complex and allow powerful complex trait dissection within this new grass model species.

plant biology

Genomic risk prediction of coronary artery disease in nearly 500,000 adults: implications for early screening and primary prevention

BackgroundCoronary artery disease (CAD) has substantial heritability and a polygenic architecture; however, genomic risk scores have not yet leveraged the totality of genetic information available nor been externally tested at population-scale to show potential utility in primary prevention.\n\nMethodsUsing a meta-analytic approach to combine large-scale genome-wide and targeted genetic association data, we developed a new genomic risk score for CAD (metaGRS), consisting of 1.7 million genetic variants. We externally tested metaGRS, individually and in combination with available conventional risk factors, in 22,242 CAD cases and 460,387 non-cases from UK Biobank.\n\nFindingsIn UK Biobank, a standard deviation increase in metaGRS had a hazard ratio (HR) of 1.71 (95% CI 1.68-1.73) for CAD, greater than any other externally tested genetic risk score. Individuals in the top 20% of the metaGRS distribution had a HR of 4.17 (95% CI 3.97-4.38) compared with those in the bottom 20%. The metaGRS had higher C-index (C=0.623, 95% CI 0.615-0.631) for incident CAD than any of four conventional factors (smoking, diabetes, hypertension, and body mass index), and addition of the metaGRS to a model of conventional risk factors increased C-index by 3.7%. In individuals on lipid-lowering or anti-hypertensive medications at recruitment, metaGRS hazard for incident CAD was significantly but only partially attenuated with HR of 2.83 (95% CI 2.61- 3.07) between the top and bottom 20% of the metaGRS distribution.\n\nInterpretationRecent genetic association studies have yielded enough information to meaningfully stratify individuals using the metaGRS for CAD risk in both early and later life, thus enabling targeted primary intervention in combination with conventional risk factors. The metaGRS effect was partially attenuated by lipid and blood pressure-lowering medication, however other prevention strategies will be required to fully benefit from earlier genomic risk stratification.\n\nFundingNational Health and Medical Research Council of Australia, British Heart Foundation, Australian Heart Foundation.

genetics

Multiple large-scale gene and genome duplications during the evolution of hexapods

Polyploidy or whole genome duplication (WGD) is a major contributor to genome evolution and diversity. Although polyploidy is recognized as an important component of plant evolution, it is generally considered to play a relatively minor role in animal evolution. Ancient polyploidy is found in the ancestry of some animals, especially fishes, but there is little evidence for ancient WGDs in other metazoan lineages. Here we use recently published transcriptomes and genomes from more than 150 species across the insect phylogeny to investigate whether ancient WGDs occurred during the evolution of Hexapoda, the most diverse clade of animals. Using gene age distributions and phylogenomics, we found evidence for 18 ancient WGDs and six other large-scale bursts of gene duplication during insect evolution. These bursts of gene duplication occurred in the history of lineages such as the Lepidoptera, Trichoptera, and Odonata. To further corroborate the nature of these duplications, we evaluated the pattern of gene retention from putative WGDs observed in the gene age distributions. We found a relatively strong signal of convergent gene retention across many of the putative insect WGDs. Considering the phylogenetic breadth and depth of the insect phylogeny, this observation is consistent with polyploidy as we expect dosage-balance to drive the parallel retention of genes. Together with recent research on plant evolution, our hexapod results suggest that genome duplications contributed to the evolution of two of the most diverse lineages of eukaryotes on Earth.

evolutionary biology

Deconvolution and phylogeny inference of structural variations in tumor genomic samples

Phylogenetic reconstruction of tumor evolution has emerged as a crucial tool for making sense of the complexity of emerging cancer genomic data sets. Despite the growing use of phylogenetics in cancer studies, though, the field has only slowly adapted to many ways that tumor evolution differs from classic species evolution. One crucial question in that regard is how to handle inference of structural variations (SVs), which are a major mechanism of evolution in cancers but have been largely neglected in tumor phylogenetics to date, in part due to the challenges of reliably detecting and typing SVs and interpreting them phylogenetically. We present a novel method for reconstructing evolutionary trajectories of SVs from bulk whole-genome sequence data via joint deconvolution and phylogenetics, to infer clonal subpopulations and reconstruct their ancestry. We establish a novel likelihood model for joint deconvolution and phylogenetic inference on bulk SV data and formulate an associated optimization algorithm. We demonstrate the approach to be efficient and accurate for realistic scenarios of SV mutation on simulated data. Application to breast cancer genomic data from The Cancer Genome Atlas (TCGA) shows it to be practical and effective at reconstructing features of SV-driven evolution in single tumors. All code can be found at https://github.com/jaebird123/tusv

cancer biology

First nuclear genome assembly of an extinct moa species, the little bush moa (Anomalopteryx didiformis)

AO_SCPLOWBSTRACTC_SCPLOWHigh throughput sequencing (HTS) has revolutionized the field of ancient DNA (aDNA) by facilitating recovery of nuclear DNA for greater inference of evolutionary processes in extinct species than is possible from mitochondrial DNA alone. We used HTS to obtain ancient DNA from the little bush moa (Anomalopteryx didiformis), one of the iconic species of large, flightless birds that became extinct following human settlement of New Zealand in the 13 th century. In addition to a complete mitochondrial genome at 249.9X depth of coverage, we recover almost 900 Mb of the moa nuclear genome by mapping reads to a high quality reference genome for the emu (Dromaius novaehollandiae). This first nuclear genome assembly for moa covers approximately 75% of the 1.2 Gb emu reference with sequence contiguity sufficient to identify more than 85% of bird universal single-copy orthologs. From this assembly, we isolate 40 polymorphic microsatellites to serve as a community resource for future population-level studies in moa. We also compile data for a suite of candidate genes associated with vertebrate limb development and show that the wingless moa phenotype is likely not attributable to gene loss or pseudogenization among this candidate set. We also identify potential function-altering coding sequence variants in moa for future experimental assays.

evolutionary biology

A genome-wide miRNA screen identifies regulators of tetraploid cell proliferation

Tetraploid cells, which are most commonly generated by errors in cell division, are genomically unstable and have been shown to promote tumorigenesis. Recent genomic studies have estimated that [~]40% of all solid tumors have undergone a genome-doubling event during their evolution, suggesting a significant role for tetraploidy in driving the development of human cancers. To safeguard against the deleterious effects of tetraploidy, non-transformed cells that fail mitosis and become tetraploid activate both the Hippo and p53 tumor suppressor pathways to restrain further proliferation. Tetraploid cells must therefore overcome these anti-proliferative barriers to ultimately drive tumor development. However, the genetic routes through which spontaneously arising tetraploid cells adapt to regain proliferative capacity remain poorly characterized. Here, we conducted a comprehensive, gain-of-function genome-wide screen to identify miRNAs that are sufficient to promote the proliferation of tetraploid cells. Our screen identified 23 miRNAs whose overexpression significantly promotes tetraploid proliferation. The vast majority of these miRNAs facilitate tetraploid growth by enhancing mitogenic signaling pathways (e.g. miR-191-3p); however, we also identified several miRNAs that impair the p53/p21 pathway (e.g. miR-523-3p), and a single miRNA (miR-24-3p) that potently inactivates the Hippo pathway via downregulation of the tumor suppressor gene NF2. Collectively, our data reveal several avenues through which tetraploid cells may regain the proliferative capacity necessary to drive tumorigenesis.

cancer biology

RAD sequencing and a hybrid Antarctic fur seal genome assembly reveal rapidly decaying linkage disequilibrium, global population structure and evidence for inbreeding

Recent advances in high throughput sequencing have transformed the study of wild organisms by facilitating the generation of high quality genome assemblies and dense genetic marker datasets. These resources have the potential to significantly advance our understanding of diverse phenomena at the level of species, populations and individuals, ranging from patterns of synteny through rates of linkage disequilibrium (LD) decay and population structure to individual inbreeding. Consequently, we used PacBio sequencing to refine an existing Antarctic fur seal (Arctocephalus gazella) genome assembly and genotyped 83 individuals from six populations using restriction site associated DNA (RAD) sequencing. The resulting hybrid genome comprised 6,169 scaffolds with an N50 of 6.21 Mb and provided clear evidence for the conservation of large chromosomal segments between the fur seal and dog (Canis lupus familiaris). Focusing on the most extensively sampled population of South Georgia, we found that LD decayed rapidly, reaching the background level of r2 = 0.09 by around 26 kb, consistent with other vertebrates but at odds with the notion that fur seals experienced a strong historical bottleneck. We also found evidence for population structuring, with four main Antarctic island groups being resolved. Finally, appreciable variance in individual inbreeding could be detected, reflecting the strong polygyny and site fidelity of the species. Overall, our study contributes important resources for future genomic studies of fur seals and other pinnipeds while also providing a clear example of how high throughput sequencing can generate diverse biological insights at multiple levels of organisation.

evolutionary biology

Subset-based genomic prediction provides insights into the genetic architecture of free amino acid levels in dry Arabidopsis thaliana seeds

Plant growth, development, and nutritional quality depends upon amino acid homeostasis, especially in seeds. However, our understanding of the underlying genetics influencing amino acid content and composition remains limited, with only a few candidate genes and quantitative trait loci identified to date. Improved knowledge of the genetics and biological processes that determine amino acid levels will enable researchers to use this information for plant breeding and biological discovery. Towards this goal, we used genomic prediction to identify biological processes that are associated with, and therefore potentially influence, free amino acid (FAA) composition in seeds of the model plant Arabidopsis thaliana. Markers were split into categories based on metabolic pathway annotations and fit using a genomic partitioning model to evaluate the influence of each pathway on heritability explained, model fit, and predictive ability. Selected pathways included processes known to influence FAA composition, albeit to an unknown degree, and spanned four categories: amino acid, core, specialized, and protein metabolism. Using this approach, we identified associations for pathways containing known variants for FAA traits, in addition to finding new trait-pathway associations. Markers related to amino acid metabolism, which are directly involved in the FAA regulation, improved predictive ability for branched chain amino acids and histidine. The use of genomic partitioning also revealed patterns across biochemical families, in which serine-derived FAAs were associated with protein related annotations and aromatic FAAs were associated with specialized metabolic pathways. Taken together, these findings provide evidence that genomic partitioning is a viable strategy to uncover the relative contributions of biological processes to FAA traits in seeds, offering a promising framework to guide hypothesis testing and narrow the search space for candidate genes.

genetics

In vivo CRISPR-Cas gene editing with no detectable genome-wide off-target mutations

CRISPR-Cas genome-editing nucleases hold substantial promise for human therapeutics1-5 but identifying unwanted off-target mutations remains an important requirement for clinical translation6, 7. For ex vivo therapeutic applications, previously published cell-based genome-wide methods provide potentially useful strategies to identify and quantify these off-target mutation sites8-12. However, a well-validated method that can reliably identify off-targets in vivo has not been described to date, leaving the question of whether and how frequently these types of mutations occur. Here we describe Verification of In Vivo Off-targets (VIVO), a highly sensitive, unbiased, and generalizable strategy that we show can robustly identify genome-wide CRISPR-Cas nuclease off-target effects in vivo. To our knowledge, these studies provide the first demonstration that CRISPR-Cas nucleases can induce substantial off-target mutations in vivo, a result we obtained using a deliberately promiscuous guide RNA (gRNA). More importantly, we used VIVO to show that appropriately designed gRNAs can direct efficient in vivo editing without inducing detectable off-target mutations. Our findings provide strong support for and should encourage further development of in vivo genome editing therapeutic strategies.

molecular biology

A dense linkage map of Lake Victoria cichlids improved the Pundamilia genome assembly and revealed a major QTL for sex-determination

Genetic linkage maps are essential for comparative genomics, high quality genome sequence assembly and fine scale quantitative trait locus (QTL) mapping. In the present study we identified and genotyped markers via restriction-site associated DNA (RAD) sequencing and constructed a genetic linkage map based on 1,597 SNP markers of an interspecific F2 cross of two closely related Lake Victoria cichlids (Pundamilia pundamilia and P. sp. \"red head\"). The SNP markers were distributed on 22 linkage groups and the total map size was 1,594 cM with an average marker distance of 1.01 cM. This high-resolution genetic linkage map was used to anchor the scaffolds of the Pundamilia genome and estimate recombination rates along the genome. Via QTL mapping we identified a major QTL for sex in a [~]1.9 Mb region on Pun-LG10, which is homologous to Oreochromis niloticus LG 23 (Ore-LG23) and includes a well-known vertebrate sex-determination gene (amh).

evolutionary biology

VIGA: a sensitive, precise and automatic de novo VIral Genome Annotator.

Viral (meta)genomics is a rapidly growing field of study that is hampered by an inability to annotate the majority of viral sequences; therefore, the development of new bioinformatic approaches is very important. Here, we present a new automatic de novo genome annotation pipeline, called VIGA, to annotate prokaryotic and eukaryotic viral sequences from (meta)genomic studies. VIGA was benchmarked on a database of known viral genomes and a viral metagenomics case study. VIGA generated the most accurate outputs according to the number of coding sequences and their coordinates, outputs also had a lower number of non-informative annotations compared to other programs.

bioinformatics

Ancient Genomics Reveals Four Prehistoric Migration Waves into Southeast Asia

Two distinct population models have been put forward to explain present-day human diversity in Southeast Asia. The first model proposes long-term continuity (Regional Continuity model) while the other suggests two waves of dispersal (Two Layer model). Here, we use whole-genome capture in combination with shotgun sequencing to generate 25 ancient human genome sequences from mainland and island Southeast Asia, and directly test the two competing hypotheses. We find that early genomes from Hoabinhian hunter-gatherer contexts in Laos and Malaysia have genetic affinities with the Onge hunter-gatherers from the Andaman Islands, while Southeast Asian Neolithic farmers have a distinct East Asian genomic ancestry related to present-day Austroasiatic-speaking populations. We also identify two further migratory events, consistent with the expansion of speakers of Austronesian languages into Island Southeast Asia ca. 4 kya, and the expansion by East Asians into northern Vietnam ca. 2 kya. These findings support the Two Layer model for the early peopling of Southeast Asia and highlight the complexities of dispersal patterns from East Asia.

evolutionary biology