Search bioRxiv⌕ Search

Biology subjects

Jenike, K. M.

Publications and source records attributed to Jenike, K. M..

6 recordsLinked to original sources

MEM-based pangenome indexing for k-mer queries

Pangenomes are growing in number and size, thanks to the prevalence of high-quality long-read assemblies. However, current methods for studying sequence composition and conservation within pangenomes have limitations. Methods based on graph pangenomes require a computationally expensive multiple-alignment step, which can leave out some variation. Indexes based on k-mers and de Bruijn graphs are limited to answering questions at a specific substring length k. We present Maximal Exact Match Ordered (MEMO), a pangenome indexing method based on maximal exact matches (MEMs) between sequences. A single MEMO index can handle arbitrary-length queries over pangenomic windows. MEMO enables both queries that test k-mer presence/absence (membership queries) and that count the number of genomes containing k-mers in a window (conservation queries). MEMOs index for a pangenome of 89 human autosomal haplotypes fits in 2.04 GB, 8.8x smaller than a comparable KMC3 index and 11.4x smaller than a PanKmer index. MEMO indexes can be made smaller by sacrificing some counting resolution, with our decile-resolution HPRC index reaching 0.67 GB. MEMO can conduct a conservation query for 31-mers over the human leukocyte antigen locus in 13.89 seconds, 2.5x faster than other approaches. MEMOs small index size, lack of k-mer length dependence, and efficient queries make it a flexible tool for studying and visualizing substring conservation in pangenomes.

bioinformatics↗

Gapless assembly of complete human and plant chromosomes using only nanopore sequencing

The combination of ultra-long Oxford Nanopore (ONT) sequencing reads with long, accurate PacBio HiFi reads has enabled the completion of a human genome and spurred similar efforts to complete the genomes of many other species. However, this approach for complete, "telomere-to-telomere" genome assembly relies on multiple sequencing platforms, limiting its accessibility. ONT "Duplex" sequencing reads, where both strands of the DNA are read to improve quality, promise high per-base accuracy. To evaluate this new data type, we generated ONT Duplex data for three widely-studied genomes: human HG002, Solanum lycopersicum Heinz 1706 (tomato), and Zea mays B73 (maize). For the diploid, heterozygous HG002 genome, we also used "Pore-C chromatin contact mapping to completely phase the haplotypes. We found the accuracy of Duplex data to be similar to HiFi sequencing, but with read lengths tens of kilobases longer, and the Pore-C data to be compatible with existing diploid assembly algorithms. This combination of read length and accuracy enables the construction of a high-quality initial assembly, which can then be further resolved using the ultra-long reads, and finally phased into chromosome-scale haplotypes with Pore-C. The resulting assemblies have a base accuracy exceeding 99.999% (Q50) and near-perfect continuity, with most chromosomes assembled as single contigs. We conclude that ONT sequencing is a viable alternative to HiFi sequencing for de novo genome assembly, and has the potential to provide a single-instrument solution for the reconstruction of complete genomes.

bioinformatics↗

Uncalled4 improves nanopore DNA and RNA modification detection via fast and accurate signal alignment

Nanopore signal analysis enables detection of nucleotide modifications from native DNA and RNA sequencing, providing both accurate genetic/transcriptomic and epigenetic information without additional library preparation. Presently, only a limited set of modifications can be directly basecalled (e.g. 5-methylcytosine), while most others require exploratory methods that often begin with alignment of nanopore signal to a nucleotide reference. We present Uncalled4, a toolkit for nanopore signal alignment, analysis, and visualization. Uncalled4 features an efficient banded signal alignment algorithm, BAM signal alignment file format, statistics for comparing signal alignment methods, and a reproducible de novo training method for k-mer-based pore models, revealing potential errors in ONTs state-of-the-art DNA model. We apply Uncalled4 to RNA 6-methyladenine (m6A) detection in seven human cell lines, identifying 26% more modifications than Nanopolish using m6Anet, including in several genes where m6A has known implications in cancer. Uncalled4 is available open-source at github.com/skovaka/uncalled4.

bioinformatics↗

Convergent evolution of plant prickles is drivenby repeated gene co-option over deep time

An enduring question in evolutionary biology concerns the degree to which episodes of convergent trait evolution depend on the same genetic programs, particularly over long timescales. Here we genetically dissected repeated origins and losses of prickles, sharp epidermal projections, that convergently evolved in numerous plant lineages. Mutations in a cytokinin hormone biosynthetic gene caused at least 16 independent losses of prickles in eggplants and wild relatives in the genus Solanum. Strikingly, homologs promote prickle formation across angiosperms that collectively diverged over 150 million years ago. By developing new Solanum genetic systems, we leveraged this discovery to eliminate prickles in a wild species and an indigenously foraged berry. Our findings implicate a shared hormone-activation genetic program underlying evolutionarily widespread and recurrent instances of plant morphological innovation.

evolutionary biology↗

Direct observation of the evolution of cell-type specific microRNA expression signatures supports the hematopoietic origin model of endothelial cells

The evolution of specialized cell-types is a long-standing interest of biologists, but given the deep time-scales very difficult to reconstruct or observe. microRNAs have been linked to the evolution of cellular complexity and may inform on specialization. The endothelium is a vertebrate specific specialization of the circulatory system that enabled a critical new level of vasoregulation. The evolutionary origin of these endothelial cells is unclear. We hypothesized that Mir-126, an endothelial cell-specific microRNA may be informative. We here reconstruct the evolutionary history of Mir-126. Mir-126 likely appeared in the last common ancestor of vertebrates and tunicates, a species without an endothelium, within an intron of the evolutionary much older EGF Like Domain Multiple (Egfl) locus. Mir-126 has a complex evolutionary history due to duplications and losses of both the host gene and the microRNA. Taking advantage of the strong evolutionary conservation of the microRNA among Olfactores, and using RNA in situ hybridization (RISH), we localized Mir-126 in the tunicate Ciona robusta. We found exclusive expression of the mature Mir-126 in granular amebocytes, supporting a long-proposed scenario that endothelial cells arose from hemoblasts, a type of proto-endothelial amoebocyte found throughout invertebrates. This observed change of expression of Mir-126 from proto-endothelial amoebocytes in the tunicate to endothelial cells in vertebrates is the first direct observation of the evolution of a cell-type in relation to microRNA expression indicating that microRNAs can be a prerequisite of cell-type evolution. Research HighlightsO_LIdirect observation of cell-type evolution C_LIO_LIhigh conservation of sequence enables for simple RISH experiment of expression C_LIO_LIMir-126 follows the evolution of hematopoetic cells to endothelial cells C_LI

cell biology↗

Expression Microdissection for use in qPCR based analysis of miRNA in a single cell type

Cell-specific microRNA (miRNA) expression estimates are important in characterizing the localization of miRNA signaling within tissues. Much of this data is obtained from cultured cells, a process known to significantly alter miRNA expression levels. Thus, our knowledge of in vivo cell miRNA expression estimates is poor. We previously demonstrated expression microdissection-miRNA-sequencing (xMD-miRNA-seq) as a means to acquire in vivo estimates, directly from formalin fixed tissues, albeit with limited yield. Here we optimized each step of the xMD process including tissue retrieval, tissue transfer, film preparation, and RNA isolation to increase RNA yields and ultimately show strong enrichment for in vivo miRNA expression by qPCR array. These method improvements, including the development of a non-crosslinked ethylene vinyl acetate (EVA) membrane, resulted in a 23-45 fold increase in miRNA yield, depending on cell type. By qPCR, miR-200a was increased 14-fold in xMD-derived small intestine epithelial cells, with a concurrent 336-fold reduction in miR-143, relative to the matched non-dissected duodenal tissue. xMD is now an optimized method to obtain robust in vivo miRNA expression estimates from cells.

genomics↗