Search bioRxivSearch

Biology subjects

Green, R. E.

Publications and source records attributed to Green, R. E..

6 recordsLinked to original sources

R2C2: Improving nanopore read accuracy enables the sequencing of highly-multiplexed full-length single-cell cDNA

High-throughput short-read sequencing has revolutionized how transcriptomes are quantified and annotated. However, while Illumina short-read sequencers can be used to analyze entire transcriptomes down to the level of individual splicing events with great accuracy, they fall short of analyzing how these individual events are combined into complete RNA transcript isoforms. Because of this shortfall, long-read sequencing is required to complement short-read sequencing to analyze transcriptomes on the level of full-length RNA transcript isoforms. However, there are issues with both Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT) long-read sequencing technologies that prevent their widespread adoption. Briefly, PacBio sequencers produce low numbers of reads with high accuracy, while ONT sequencers produce higher numbers of reads with lower accuracy. Here we introduce and validate a new long-read ONT based sequencing method. At the same cost, our Rolling Circle Amplification to Concatemeric Consensus (R2C2) method generates more accurate reads of full-length RNA transcript isoforms than any other available long-read sequencing method. These reads can then be used to generate isoform-level transcriptomes for both genome annotation and differential expression analysis in bulk or single cell samples.\n\nSignificance StatementSubtle changes in RNA transcript isoform expression can have dramatic effects on cellular behaviors in both health and disease. As such, comprehensive and quantitative analysis of isoform-level transcriptomes would open an entirely new window into cellular diversity in fields ranging from developmental to cancer biology. The R2C2 method we are presenting here is the first method with sufficient throughput and accuracy to make the comprehensive and quantitative analysis of RNA transcript isoforms in bulk and single cell samples economically feasible.

genomics

EquCab3, an Updated Reference Genome for the Domestic Horse

EquCab2, a high-quality reference genome for the domestic horse, was released in 2007. Since then, it has served as the foundation for nearly all genomic work done in equids. Recent advances in genomic sequencing technology and computational assembly methods have allowed scientists to improve reference assemblies of large animal and plant genomes in terms of contiguity and composition. In 2014, the equine genomics research community began a project to improve the reference sequence for the horse, building upon the solid foundation of EquCab2 and incorporating new short-read data, long-read data, and proximity ligation data. The result, EquCab3, is presented here. The count of non-N bases in the incorporated chromosomes is improved from 2.33Gb in EquCab2 to 2.41Gb from EquCab3. Contiguity has also been improved nearly 40-fold with a contig N50 of 4.5Mb and scaffold contiguity enhanced to where all but one of the 32 chromosomes is comprised of a single scaffold.

genomics

Structural variation detection by proximity ligation from FFPE tumor tissue

The clinical management and therapy of many solid tumor malignancies is dependent on detection of medically actionable or diagnostically relevant genetic variation. However, a principal challenge for genetic assays from tumors is the fragmented and chemically damaged state of DNA in formalin-fixed paraffin-embedded (FFPE) samples. From highly fragmented DNA and RNA there is no current technology for generating long-range DNA sequence data as is required to detect genomic structural variation or long-range genotype phasing. We have developed a high-throughput chromosome conformation capture approach for FFPE samples that we call \"Fix-C\", which is similar in concept to Hi-C. Fix-C enables structural variation detection from fresh and archival FFPE samples. We applied this method to 15 clinical adenocarcinoma and sarcoma specimens spanning a broad range of tumor purities. In this panel, Fix-C analysis achieves a 90% concordance rate with FISH assays - the current clinical gold standard. Additionally, we are able to identify novel structural variation undetected by other methods and recover long-range chromatin configuration information from these FFPE samples harboring highly degraded DNA. This powerful approach will enable detailed resolution of global genome rearrangement events during cancer progression from FFPE material, and inform the development of targeted molecular diagnostic assays for patient care.

genomics

The genomic false shuffle: epigenetic maintenance of topological domains in the rearranged gibbon genome

The relationship between evolutionary genome remodeling and the three-dimensional structure of the genome remain largely unexplored. Here we use the heavily rearranged gibbon genome to examine how evolutionary chromosomal rearrangements impact genome-wide chromatin interactions, topologically associating domains (TADs), and their epigenetic landscape. We use high-resolution maps of gibbon-human breaks of synteny (BOS), apply Hi-C in gibbon, measure an array of epigenetic features, and perform cross-species comparisons. We find that gibbon rearrangements occur at TAD boundaries, independent of the parameters used to identify TADs. This overlap is supported by a remarkable genetic and epigenetic similarity between BOS and TAD boundaries, namely presence of CpG islands and SINE elements, and enrichment in CTCF and H3K4me3 binding. Cross-species comparisons reveal that regions orthologous to BOS also correspond with boundaries of large (400-600kb) TADs in human and other mammalian species. The co-localization of rearrangement breakpoints and TAD boundaries may be due to higher chromatin fragility at these locations and/or increased selective pressure against rearrangements that disrupt TAD integrity. We also examine the small portion of BOS that did not overlap with TAD boundaries and gave rise to novel TADs in the gibbon genome. We postulate that these new TADs generally lack deleterious consequences. Lastly, we show that limited epigenetic homogenization occurs across breakpoints, irrespective of their time of occurrence in the gibbon lineage. Overall, our findings demonstrate remarkable conservation of chromatin interactions and epigenetic landscape in gibbons, in spite of extensive genomic shuffling.

genomics

Genomic evidence of globally widespread admixture from polar bears into brown bears during the last ice age

Recent genomic analyses have provided substantial evidence for past periods of gene flow from polar bears (Ursus maritimus) into Alaskan brown bears (Ursus arctos), with some analyses suggesting a link between climate change and genomic introgression. However, because it has only been possible to sample bears from the present day, the timing, frequency, and evolutionary significance of this admixture remains unknown. Here, we analyze genomic DNA from three additional and geographically distinct brown bear populations, including two that lived temporally close to the peak of the last ice age. We find evidence of admixture in all three populations, suggesting that admixture between these species has been common in their recent evolutionary history. In addition, analyses of ten fossil bears from the now-extinct Irish population indicate that admixture peaked during the last ice age, when brown bear and polar bear ranges overlapped. Following this peak, the proportion of polar bear ancestry in Irish brown bears declined rapidly until their extinction. Our results support a model in which ice age climate change created geographically widespread conditions conducive to admixture between polar bears and brown bears, as is again occurring today. We postulate that this model will be informative for many admixing species pairs impacted by climate change. Our results highlight the power of paleogenomes to reveal patterns of evolutionary change that are otherwise masked with only contemporary data.

evolutionary biology

Natural selection shaped the rise and fall of passenger pigeon genomic diversity

The extinct passenger pigeon was once the most abundant bird in North America, and possibly the world. While theory predicts that large populations will be more genetically diverse and respond more efficiently to selection, passenger pigeon genetic diversity was surprisingly low. To investigate this we analysed 41 mitochondrial and 4 nuclear genomes from passenger pigeons, and 2 genomes from band-tailed pigeons, passenger pigeons closest living relatives. We find that passenger pigeons large population size allowed for faster adaptive evolution and removal of harmful mutations, but that this drove a huge loss in neutral genetic diversity. These results demonstrate how great an impact selection can have on a vertebrate genome, and invalidate previous results that suggested population instability contributed to this species surprisingly rapid extinction.

genomics