Search bioRxiv⌕ Search

Biology subjects

Goodstein, D. M.

Publications and source records attributed to Goodstein, D. M..

4 recordsLinked to original sources

A view of the pan-genome of domesticated cowpea (Vigna unguiculata Walp.)

Cowpea, Vigna unguiculata L. Walp., is a diploid warm-season legume of critical importance as both food and fodder in sub-Saharan Africa. This species is also grown in Northern Africa, Europe, Latin America, North America, and East to Southeast Asia. To capture the genomic diversity of domesticates of this important legume, de novo genome assemblies were produced for representatives of six sub-populations of cultivated cowpea identified previously from genotyping of several hundred diverse accessions. In the most complete assembly (IT97K-499-35), 26,026 core and 4,963 noncore genes were identified, with 35,436 pan genes when considering all seven accessions. GO-terms associated with response to stress and defense response were highly enriched among the noncore genes, while core genes were enriched in terms related to transcription factor activity, and transport and metabolic processes. Over 5 million SNPs relative to each assembly and over 40 structural variants >1 Mb in size were identified by comparing genomes. Vu10 was the chromosome with the highest frequency of SNPs, and Vu04 had the most structural variants. Noncore genes harbor a larger proportion of potentially disruptive variants than core genes, including missense, stop gain, and frameshift mutations; this suggests that noncore genes substantially contribute to diversity within domesticated cowpea. Article SummaryThis study reports annotated genome assemblies of six cowpea accessions. Together with the previously reported annotated genome of IT97K-499-35, these constitute a pan-genome resource representing six subpopulations of domesticated cowpea. Annotations include genes, variant calls for SNPs and short indels, larger presence or absence variants, and inversions. Noncore genes are enriched for loci involved in stress response and harbor many genic variants with potential effects on coding sequence.

genomics↗

The Chlamydomonas Genome Project, version 6: reference assemblies for mating type plus and minus strains reveal extensive structural mutation in the laboratory

Five versions of the Chlamydomonas reinhardtii reference genome have been produced over the last two decades. Here we present version 6, bringing significant advances in assembly quality and structural annotations. PacBio-based chromosome-level assemblies for two laboratory strains, CC-503 and CC-4532, provide resources for the plus and minus mating type alleles. We corrected major misassemblies in previous versions and validated our assemblies via linkage analyses. Contiguity increased over ten-fold and >80% of filled gaps are within genes. We used Iso-Seq and deep RNA-seq datasets to improve structural annotations, and updated gene symbols and textual annotation of functionally characterized genes via extensive curation. We discovered that the cell wall-less classical reference strain CC-503 exhibits genomic instability potentially caused by deletion of RECQ3 helicase, with major structural mutations identified that affect >100 genes. We therefore present the CC-4532 assembly as the primary reference, although this strain also carries unique structural mutations and is experiencing rapid proliferation of a Gypsy retrotransposon. We expect all laboratory strains to harbor gene-disrupting mutations, which should be considered when interpreting and comparing experimental results across laboratories and over time. Collectively, the resources presented here herald a new era of Chlamydomonas genomics and will provide the foundation for continued research in this important reference.

genomics↗

GENESPACE: syntenic pan-genome annotations for eukaryotes

The development of multiple high-quality reference genome sequences in many taxonomic groups has yielded a high-resolution view of the patterns and processes of molecular evolution. Nonetheless, leveraging information across multiple reference haplotypes remains a significant challenge in nearly all eukaryotic systems. These challenges range from studying the evolution of chromosome structure, to finding candidate genes for quantitative trait loci, to testing hypotheses about speciation and adaptation in nature. Here, we address these challenges through the concept of a pan-genome annotation, where conserved gene order is used to restrict gene families and define the expected physical position of all genes that share a common ancestor among multiple genome annotations. By leveraging pan-genome annotations and exploring the underlying syntenic relationships among genomes, we dissect presence-absence and structural variation at four levels of biological organization: among three tetraploid cotton species, across 300 million years of vertebrate sex chromosome evolution, across the diversity of the Poaceae (grass) plant family, and among 26 maize cultivars. The methods to build and visualize syntenic pan-genome annotations in the GENESPACE R package offer a significant addition to existing gene family and synteny programs, especially in polyploid, outbred and other complex genomes.

genomics↗

ePlant in 2021: New Species, Viewers, Data Sets, and Widgets

ePlant was introduced in 2017 for exploring large Arabidopsis thaliana data sets from the kilometre to nanometre scales. In the past four years we have used the ePlant framework to develop ePlants for 15 agronomically-important species: maize, poplar, tomato, Camelina sativa, soybean, potato, barley, Medicago truncatula, eucalyptus, rice, willow, sunflower, Cannabis sativa, wheat and sugarcane. We also updated the interface to improve performance and accessibility, and added two new views to the Arabidopsis ePlant - the Navigator and Pathways viewers. The former shows phylogenetic relationships between homologs in other species and their expression pattern similarities, with links to view data for those genes in the respective ePlants. The latter shows Plant Reactome metabolic reactions. We also describe new Arabidopsis data sets including single cell RNA-seq data from roots, and how to embed ePlant eFP expression pictographs into any web page.

bioinformatics↗