Search bioRxivSearch

Biology subjects

Nam, J.-W.

Publications and source records attributed to Nam, J.-W..

3 recordsLinked to original sources

Whole genome hybrid assembly and protein-coding gene annotation of the entirely black native Korean chicken breed Yeonsan Ogye

Yeonsan Ogye (YO), an indigenous Korean chicken breed (gallus gallus domesticus), has entirely black external features and internal organs. In this study, the draft genome of YO was assembled using a hybrid de novo assembly method that takes advantage of high-depth Illumina short-reads (232.2X) and low-depth PacBio long-reads (11.5X). Although the contig and scaffold N50s (defined as the shortest contig or scaffold length at 50% of the entire assembly) of the initial de novo assembly were 53.6Kbp and 10.7Mbp, respectively, additional and pseudo-reference-assisted assemblies extended the assembly to 504.8Kbp for contig N50 (pseudo-contig) and 21.2Mbp for scaffold N50, which included 551 structural variations including the Fibromelanosis (FM) locus duplication, compared to galGal4 and 5. The completeness (97.6%) of the draft genome (Ogye_1) was evaluated with single copy orthologous genes using BUSCO, and found to be comparable to the current chicken reference genome (galGal5; 97.4%), which was assembled with a long read-only method, and superior to other avian genomes (92~93%), assembled with short read-only and hybrid methods. To comprehensively reconstruct transcriptome maps, RNA sequencing (RNA-seq) and representation bisulfite sequencing (RRBS) data were analyzed from twenty different tissues, including black tissues. The maps included 15,766 protein-coding and 6,900 long non-coding RNA genes, many of which were expressed in the tissue-specific manner, closely related with the DNA methylation pattern in the promoter regions.

genomics

Non-coding Transcriptome Maps Across Twenty Tissues of the Korean Black Chicken, Yeonsan Ogye

The Yeonsan Ogye (Ogye) is a rare Korean domestic chicken breed, the entire body of which, including its feathers and skin, has a unique black coloring. Although some protein-coding genes related to this unique feature have been examined, non-coding elements have not been globally investigated. In this study, high-throughput RNA sequencing and DNA methylation sequencing were performed to dissect the expression landscape of 14,264 Ogye protein-coding and 6900 long non-coding RNA (lncRNA) genes along with DNA methylation landscape in twenty different Ogye tissues. About 75% of Ogye lncRNAs showed tissue-specific expression whereas about 45% of protein-coding genes did. For some genes, the tissue-specific expression levels were inversely correlated with DNA methylation levels in their promoters. About 39% of the tissue-specific lncRNAs displayed functional association with proximal or distal protein-coding genes. In particular, heat shock transcription factor 2 (HSF2)-associated lncRNAs were discovered to be functionally linked to protein-coding genes that are specifically expressed in black skin tissues, tended to be more syntenically conserved in mammals, and were differentially expressed in black tissues relative to white tissues. Our results not only facilitate understanding how the non-coding genome regulates unique phenotypes but also should be of use for future genomic breeding of chickens.

bioinformatics

High-confidence Coding and Noncoding Transcriptome Maps

The advent of high-throughput RNA-sequencing (RNA-seq) has led to the discovery of unprecedentedly immense transcriptomes encoded by eukaryotic genomes. However, the transcriptome maps are still incomplete partly because they were mostly reconstructed based on RNA-seq reads that lack their orientations (known as unstranded reads) and certain boundary information. Methods to expand the usability of unstranded RNA-seq data by predetermining the orientation of the reads and precisely determining the boundaries of assembled transcripts could significantly benefit the quality of the resulting transcriptome maps. Here, we present a high-performing transcriptome assembly pipeline, called CAFE, that significantly improves the original assemblies, respectively assembled with stranded and/or unstranded RNA-seq data, by orienting unstranded reads using the maximum likelihood estimation and by integrating information about transcription start sites and cleavage and polyadenylation sites. Applying large-scale transcriptomic data comprising ninety-nine billion RNAs-seq reads from the ENCODE, human BodyMap projects, The Cancer Genome Atlas, and GTEx, CAFE enabled us to predict the directions of about eighty-nine billion unstranded reads, which led to the construction of more accurate transcriptome maps, comparable to the manually curated map, and a comprehensive lncRNA catalogue that includes thousands of novel lncRNAs. Our pipeline should not only help to build comprehensive, precise transcriptome maps from complex genomes but also to expand the universe of non-coding genomes.

bioinformatics