Search bioRxiv⌕ Search

Biology subjects

Uno, F.

Publications and source records attributed to Uno, F..

4 recordsLinked to original sources

Digital Kennison: A bioinformatics pipeline for rapid mapping of sequences to the Drosophila melanogaster Y chromosome

The Drosophila melanogaster Y chromosome is currently known to contain 13 single-copy protein-coding genes, six of which are essential for male fertility, as well as several non-coding genes and abundant repetitive DNA. Localization of Y-linked sequences has traditionally relied on labor-intensive crosses using Kennisons translocation strains, which map Y-linked loci by generating flies deficient for each of the six Y-chromosome fertility regions (ks-1, ks-2, kl-1, kl-2, kl-3, and kl-5). Here we present Digital Kennison, a computational pipeline that recasts this classical mapping strategy as a sequence-based analysis. The pipeline queries eight genomic databases derived from Kennisons strains using BLAST and read coverage, assigning sequences to fertility regions with a calibrated confidence score. We benchmarked the method on 60 Y-linked sequences spanning all six regions, including single-copy protein-coding genes, Mst77Y family members, non-coding RNAs, and the centromere. Digital Kennison achieved 97% precision while resolving challenging cases, including boundary-spanning genes (PRY and Ppr-Y), fragmented Mst77Y copies, and FDY, which has a closely related autosomal paralog. Beyond validating known localizations, the pipeline localized the unmapped gene CG41561 to the kl-1region and reassigned the transcript CR40629-RC from the kl-2 region to kl-5. It also localized 7 of 16 recently transferred Y-linked sequences described by Tobler et al. (2017), including 4 with high confidence. Applied to 904 small R6 scaffolds, Digital Kennison assigned 75% to fertility regions, including five currently annotated as autosomal-pericentromeric. Digital Kennison reduces sequence localization from weeks of genetic crosses to minutes of computation while preserving the power of classical translocation mapping. Article summaryThe Drosophila melanogaster Y chromosome is difficult to study because it consists largely of repetitive, non-recombining DNA. Researchers have traditionally mapped Y-linked genes using slow, labor-intensive genetic crosses. Here we introduce Digital Kennison, a computational pipeline that replicates this classical mapping strategy using DNA sequence data instead of live flies. By comparing a query sequence against genomic databases built from fly strains, the pipeline assigns it to one of six Y-chromosome regions and reports a confidence score. Tested on 60 known sequences, it achieved 97% precision, corrected an annotation error, and mapped previously unplaced sequences in minutes rather than weeks.

genomics↗

Triplex DNA and inverted repeats cause long-read sequencing bias against simple satellite DNA

We recently showed that Oxford Nanopore Technologies (ONT) and Pacific Biosciences (PacBio) have very strong sequencing bias against simple satellites, probably caused by single-stranded DNA folding into non-canonical (non-B) structures during sequencing. Here we extend these observations by computational and experimental approaches in the Drosophila and human genomes. We found that (i) only a small subset of simple satellites cause sequencing bias; many satellites (e.g., (ACTGGG)n) are benign and easily sequenced. (ii) The biases most likely are caused by two distinct non-B DNA structures: triplex DNA formed by some, but not all, AG-rich satellites (only those predicted to form strong mirror repeats), and hairpins formed by some, but not all, AT-rich satellites (only those predicted to form fairly strong inverted repeats). (iii) The correlation between the predicted stability of these non-B structures, and the strength of sequencing bias indicates that non-B DNA is indeed the culprit. (iv) The likely source of these non-B structures is single-stranded DNA formed during ONT and PacBio sequencing, and hence its removal might solve the bias. We tested this by adding single-strand binding protein to ONT sequencing, and found that it irreversibly kills the flow cells. (v) A recent sequencing effort in Drosophila melanogaster using very high depth Ultra-Long ONT sequencing (967x) still failed to assemble many genes located near satellite blocks. Brute-force will not solve the problem; instead, further investment is needed by the sequencing companies to achieve truly unbiased sequencing.

genomics↗

Anopheles (Kerteszia) cruzii, the main malaria vector in the Brazilian Atlantic Forest, is a complex of at least five cryptic species

Malaria, a tropical disease caused by Plasmodium and transmitted by Anopheles, remains a public health concern in Brazil. While most cases occur in the Amazon, transmission persists in the Atlantic Forest, where Anopheles mosquitoes of the Kerteszia subgenus are the primary vectors of human and simian malaria. Previous studies using cytogenetics, isoenzymes, and molecular markers have suggested cryptic species within Anopheles (Kerteszia) cruzii and Anopheles (Kerteszia) bellator. We sequenced 55 genomes: 35 An. cruzii s.l. (four with Nanopore and 31 with Illumina), 12 An. bellator s.l., and eight An. homunculus, the latter two with Illumina. Phylogenomic analysis revealed at least five cryptic species within An. cruzii s.l., labelled A-E, with evidence of sympatry in some locations. Anopheles bellator s.l. also forms a species complex, comprising at least three distinct lineages. These cryptic species showed high genetic differentiation (FST range: 0.4-0.7), typical of interspecific comparisons. In contrast, An. homunculus populations showed low differentiation (FST [~] 0.2), suggesting a single widespread species. Our analysis confirms cryptic speciation in An. cruzii and An. bellator, but not in An. homunculus. These findings are important for understanding malaria transmission in the Atlantic Forest, given that vector competence may differ among cryptic species.

evolutionary biology↗

Strong sequencing bias in Nanopore and PacBio prevents assembly of Drosophila melanogaster Y-linked genes

Nanopore and PacBio are generally considered free from sequence composition bias, a key factor - alongside read length - that explains their success in producing high quality genome assemblies. However, our study reveals a systematic failure of both technologies to sequence and assemble specific exons of Drosophila melanogaster genes, indicating an overlooked limitation. Namely, multiple Y-linked exons are nearly or completely absent from raw reads produced by deep sequencing with state-of-the-art Nanopore (10.4 flow cells, 200x coverage) and PacBio (HiFi 50x). The same exons are accurately assembled using Illumina 65x coverage. We found that these missing exons are consistently located near simple satellite sequences, where sequencing fails at multiple levels: read initiation (very few reads start within satellite regions), read elongation (satellite-containing reads are shorter on average), and base-calling (quality scores drop as sequencing enters a satellite sequence). These findings challenge the assumption that long-read technologies is unbiased and reveal a critical barrier to assembling sequences near repetitive regions. As large-scale sequencing projects move towards telomere-to-telomere assemblies in a wide range of organisms, recognizing and addressing these biases will be important to achieving truly complete and accurate genomes. Additionally, the underrepresented Y-linked exons provides a valuable benchmark for refining those sequencing technologies while improving the assembly of the highly heterochromatic and often neglected Drosophila Y chromosome.

genomics↗