Search bioRxiv⌕ Search

Biology subjects

Darby, C.

Publications and source records attributed to Darby, C..

2 recordsLinked to original sources

De novo assembly and annotation of the Patagonian toothfish (Dissostichus eleginoides) genome

Patagonian toothfish (Dissostichus eleginoides) is an economically and ecologically important fish species in the family Nototheniidae, found at depths between 70 and 2,500 meters on the southern shelves and slopes around the sub-Antarctic islands of the Southern Ocean. Genomic sequence data for this species is limited. Here, we report a high-quality assembly and annotation of the D. eleginoides genome, generated using a combination of Illumina, PacBio and Omni-C sequencing technologies. To aid the genome annotation, the transcriptome derived from a variety of toothfish tissues was also generated using both short and long read sequencing methods. The final genome assembly was 797.8 Mb with a N50 scaffold length of 3.5 Mb. Approximately 31.7% of the genome consisted of repetitive elements. A total of 35,543 putative protein-coding regions were identified, of which 50% have been functionally annotated. Transcriptomics analysis showed that approximately 64% of the predicted genes (22,617 genes) were found to be expressed in the tissues sampled. Comparative genomics analysis revealed that the anti-freeze glycoprotein (AFGP) locus of D. eleginoides does not contain any AFGP proteins compared to the same locus in the Antarctic toothfish (Dissostichus mawsoni). This is in agreement with previously published results looking at hybridization signals and confirms that Patagonian toothfish do not possess AFGP coding sequences in their genome. The high-quality genome assembly of the Patagonian toothfish will provide a valuable genetic resource for ecological and evolutionary studies on this and other closely related species.

genomics↗

Integrated analysis of multimodal single-cell data

The simultaneous measurement of multiple modalities, known as multimodal analysis, represents an exciting frontier for single-cell genomics and necessitates new computational methods that can define cellular states based on multiple data types. Here, we introduce weighted-nearest neighbor analysis, an unsupervised framework to learn the relative utility of each data type in each cell, enabling an integrative analysis of multiple modalities. We apply our procedure to a CITE-seq dataset of hundreds of thousands of human white blood cells alongside a panel of 228 antibodies to construct a multimodal reference atlas of the circulating immune system. We demonstrate that integrative analysis substantially improves our ability to resolve cell states and validate the presence of previously unreported lymphoid subpopulations. Moreover, we demonstrate how to leverage this reference to rapidly map new datasets, and to interpret immune responses to vaccination and COVID-19. Our approach represents a broadly applicable strategy to analyze single-cell multimodal datasets, including paired measurements of RNA and chromatin state, and to look beyond the transcriptome towards a unified and multimodal definition of cellular identity. AvailabilityInstallation instructions, documentation, tutorials, and CITE-seq datasets are available at http://www.satijalab.org/seurat

genomics↗