Search bioRxiv⌕ Search

Biology subjects

Beeloo, R.

Publications and source records attributed to Beeloo, R..

5 recordsLinked to original sources

Barbell Resolves Demultiplexing and Trimming Issues in Nanopore Data

BackgroundOxford Nanopore sequencing enables long-read sequencing across diverse applications, yet the experimental artifacts introduced by Nanopore barcoding are not well characterized. These artifacts can affect demultiplexing accuracy and downstream analyses. ResultsWe performed a rapid barcoding experiment on 66 diagnostic samples and found that 83% of reads carried the expected single-barcode pattern, while 17% contained multiple barcodes or other artifacts. Current demultiplexers, including the widely used Dorado, fail to correctly handle these complex cases, leaving approximately 7% of reads partially trimmed and contaminated with adapter fragments. Additional issues include the presence of two barcodes at the same read end--either identical, originating from the same sample, or different, introduced after pooling. The latter can lead to barcode bleeding when the outer barcode is incorrectly selected. To address these challenges, we developed Barbell, a pattern-aware demultiplexer that detects all barcode configurations. Barbell reduces trimming errors by three orders of magnitude, minimizes barcode bleeding, and supports custom experimental setups such as shorter barcodes, dual-end barcodes, and custom flank sequences. ConclusionsOur results highlight the impact of complex barcode attachments in Nanopore sequencing and demonstrate that Barbell drastically reduces their effects on downstream analyses. Barbell is open source and available at https://github.com/rickbeeloo/barbell.

bioinformatics↗

Bacterial community adaptation after freshwater and seawater coalescence

Microbial community coalescence, the merging of entire microbial communities, is common across ecosystems, particularly in estuaries where freshwater and seawater mix. The complexity of these habitats makes the in situ study of community dynamics after coalescence challenging, highlighting the need for controlled experiments to unravel the factors influencing the coalescence in the estuary. To study these processes, we combined natural freshwater and seawater bacterial communities at five different mixing ratios and incubated them in parallel microcosms containing freshwater or seawater incubation media. Forty mixed communities were tracked over six passages using Nanopore full-length 16S rRNA gene sequencing. In the original field samples, freshwater hosted more diverse communities than seawater. While the communities were structurally distinct, shared bacterial families accounted for approximately 95% of total reads. Many low-abundance taxa were lost upon laboratory incubation, while potentially faster-growing ones were enriched. We found that the coalescence outcome was strongly shaped by the incubation media, whereas the mixing ratio had a minor influence. Mixed communities converged toward the source community native to the incubation media, with increasing similarity at a higher source community proportion. In freshwater, a 25% inoculum of the freshwater community was sufficient to re-establish a near-native freshwater community, whereas in seawater, similarity to the native seawater community depended on the inoculation ratio. Network analysis showed a tightly connected module of seawater families, reflecting their shared habitat preference or cooperation, whereas freshwater families were more loosely connected. We also observed that most families were unaffected by mixing ratios or temporal dynamics, with only a few showing mixing ratio dependence. For instance, the freshwater family Comamonadaceae and seawater families Marinomonadaceae and Pseudoalteromonadaceae were dominant in their respective native environments, and increased in mixed communities in proportion to their initial source proportion. Overall, environmental filtering had a stronger impact than the mixing ratio of source communities on coalesced communities, with habitat-specific taxa further modulating the outcome. These findings advanced our understanding of microbial responses to coalescence and provided insights into microbial community assembly in dynamic estuarine systems. HighlightsO_LIEnvironmental filtering outweighs the microbial source community ratio in shaping coalescence outcomes. C_LIO_LIAsymmetric resilience: freshwater communities require a lower inoculum to re-establish than seawater communities. C_LIO_LIModularity of the seawater source community was observed during coalescence. C_LIO_LIFull-length 16S rRNA gene profiling with a custom dual-barcoding Nanopore protocol enables cost-effective bacterial community tracking. C_LIO_LIControlled coalescence experiments offer mechanistic insights into estuarine microbial community transitions. C_LI

ecology↗

Sassy: Searching Short DNA Strings in the 2020s

MotivationApproximate string matching (ASM) is the problem of finding all occurrences of a pattern in a text while allowing up to k errors. Many modern methods use seed-chain-extend, which is fast in practice, but does not guarantee finding all matches with [≤] k errors. However, applications such as CRISPR off-target detection require exhaustive results. MethodsWe introduce Sassy, a library and tool for ASM of short patterns in long texts. Sassy splits the text into 4 parts that are searched in parallel, and uses bitvectors in the text direction rather than the pattern direction. This has compexity O(k{lceil}n/W {rciel}) when searching a random text of length n, where W = 256 is the SIMD width, and provides significant speedups for small k. Separately, we allow matches of the pattern to extend beyond the text for an overhang cost of e.g. = 0.5 per character, to find matches near contig or read ends. ResultsSassy is 4x to 15x faster than Edlib for patterns [≤] 1000bp, and can search text with a throughput near 2 Gbp/s. Likewise, Sassy is over 100x faster than parasail. We apply Sassy to CRISPR off-target detection by searching 61 guide sequences in a human genome. Sassy is 100x faster than SWOffinder and only slightly slower (for k [≤] 3) than CHOPOFF, for which building its index takes 20 minutes. Sassy also scales well to larger k, unlike CHOPOFF whose index took over 10 hours to build for k = 5. AvailibilitySassy is available as library and binary at https://github.com/RagnarGrootKoerkamp/sassy, and archived at swh:1:dir:e884758dce5777a441bc2799dc8824e563c5f97b.

bioinformatics↗

Jaeger: an accurate and fast deep-learning tool to detect bacteriophage sequences

Viruses are integral to every biome on Earth, yet we still need a more comprehensive picture of their identity and global distribution. Global metagenomics sequencing efforts revealed the genomic content of tens of thousands of environmental samples, however identifying the viral sequences in these datasets remains challenging due to their vast genomic diversity. Here, we address identifying bacteriophage sequences in unlabeled sequencing data. In a recent benchmarking paper, we observed that existing deep-learning tools show a high true positive rate, but may also produce many false positives when confronted with divergent sequences. To tackle this challenge, we introduce Jaeger, a novel deep-learning method designed specifically for identifying bacteriophage genome fragments. Extensive benchmarking on the IMG/VR database and real-world metagenomes reveals Jaegers consistent high sensitivity (0.87) and precision (0.92). Applying Jaeger to over 16,000 metagenomic assemblies from the MGnify database yielded over five million putative phage contigs. On average, Jaeger is around 20 times faster than the other state-of-the-art methods. Jaeger is available at https://github.com/MGXlab/Jaeger.

bioinformatics↗

Graphite: painting genomes using a colored De Bruijn graph

The recent growth of microbial sequence data allows comparisons at unprecedented scales, enabling tracking of strains, mobile genetic elements, or genes. Querying a genome against a large reference database can easily yield thousands of matches that are tedious to interpret and pose computational challenges. We developed Graphite that uses a colored De Bruijn graph (cDBG) to paint query genomes, selecting the local best matches along the full query length. By focusing on the closest genomic match of each query region, Graphite reduces the number of matches while providing promising leads for genomic forensics. When applied to hundreds of Campylobacter genomes we found extensive gene sharing, including a previously undetected C. coli plasmid that matched a C. jejuni chromosome. Together, genome painting using cDBGs as enabled by Graphite, can reveal new biological phenomena by mitigating computational hurdles. Graphite is implemented in Julia, available at https://github.com/MGXlab/Graphite.

bioinformatics↗