Search bioRxiv⌕ Search

Biology subjects

Mamerto, A.

Publications and source records attributed to Mamerto, A..

3 recordsLinked to original sources

Domesticated cannabinoid synthases amid a wild mosaic cannabis pangenome

Cannabis sativa is a globally significant seed-oil, fiber, and drug-producing plant species. However, a century of prohibition has severely restricted legal breeding and germplasm resource development, leaving potential hemp-based nutritional and fiber applications unrealized. Existing cultivars are highly heterozygous and lack competitiveness in the overall fiber and grain markets, relegating hemp to less than 200,000 hectares globally1. The relaxation of drug laws in recent decades has generated widespread interest in expanding and reincorporating cannabis into agricultural systems, but progress has been impeded by the limited understanding of genomics and breeding potential. No studies to date have examined the genomic diversity and evolution of cannabis populations using haplotype-resolved, chromosome-scale assemblies from publicly available germplasm. Here we present a cannabis pangenome, constructed with 181 new and 12 previously released genomes from a total of 156 biological samples from both male (XY) and female (XX) plants, including 42 trio phased and 36 haplotype-resolved, chromosome-scale assemblies. We discovered widespread regions of the cannabis pangenome that are surprisingly diverse for a single species, with high levels of genetic and structural variation, and propose a novel population structure and hybridization history. Conversely, the cannabinoid synthase genes contain very low levels of diversity, despite being embedded within a variable region containing multiple pseudogenized paralogs and distinct transposable element arrangements. Additionally, we identified variants of acyl-lipid thioesterase (ALT) genes2 that are associated with fatty acid chain length variation and the production of the rare cannabinoids, tetrahydrocannabinol varin (THCV) and cannabidiol varin (CBDV). We conclude the Cannabis sativa gene pool has only been partially characterized, and that the existence of wild relatives in Asia remains likely, while its potential as a crop species remains largely unrealized.

plant biology↗

Telomere Length in Plants Estimated with Long Read Sequencing

Telomeres play an important role in chromosome stability and their length is thought to be related to an organisms lifestyle and lifespan. Telomere length is variable across plant species and between cultivars of the same species, possibly conferring adaptive advantage. However, it is not known whether telomere length is related to lifestyle or life span across a diverse array of plant species due to the lack of information on telomere length in plants. Here we leverage genomes assembled with long read sequencing data to estimate telomere length by chromosome. We find that long read assemblies based on Oxford Nanopore Technologies (ONT) accurately predict telomere length in the two model plant species Arabidopsis thaliana and Oryza sativa matching lab-based length estimates. We then estimate telomere length across an array of plant species with different lifestyles and lifespans and find that in general gymnosperms have shorter telomeres compared to eudicots and monocots. Crop species frequently have longer telomeres than their wild relatives, and species that have been maintained clonally such as hemp have long telomeres possibly reflecting that this lifestyle requires long term chromosomal stability.

plant biology↗

PanKmer: k-mer based and reference-free pangenome analysis

SummaryPangenomes are replacing single reference genomes as the definitive representation of DNA sequence within a species or clade. Pangenome analysis predominantly leverages graph-based methods that require computationally intensive multiple genome alignments, do not scale to highly complex eukaryotic genomes, limit their scope to identifying structural variants (SVs), or incur bias by relying on a reference genome. Here, we present PanKmer, a toolkit designed for reference-free analysis of pangenome datasets consisting of dozens to thou-sands of individual genomes. PanKmer decomposes a set of input genomes into a table of observed k-mers and their presence-absence values in each genome. These are stored in an efficient k-mer index data format that encodes SNPs, INDELs, and SVs. It also includes functions for downstream analysis of the k-mer index, such as calculating sequence similarity statistics between individuals at whole-genome or local scales. For example, k-mers can be "anchored" in any individual genome to quantify sequence variability or conservation at a specific locus. This facilitates workflows with various biological applications, e.g. identifying cases of hybridization between plant species. PanKmer provides researchers with a valuable and convenient means to explore the full scope of genetic variation in a population, without reference bias. Availability and implementationPanKmer is implemented as a Python package with components written in Rust, released under a BSD license. The source code is available from the Python Package Index (PyPI) at https://pypi.org/project/pankmer/ as well as Gitlab at https://gitlab.com/salk-tm/pankmer. Full documentation is available at https://salk-tm.gitlab.io/pankmer/. Supplementary informationSupplementary data are available online

bioinformatics↗