Search bioRxiv⌕ Search

Biology subjects

Heinz, J.

Publications and source records attributed to Heinz, J..

4 recordsLinked to original sources

CHESS 3: an improved, comprehensive catalog of human genes and transcripts based on large-scale expression data, phylogenetic analysis, and protein structure

The original CHESS database of human genes was assembled from nearly 10,000 RNA sequencing experiments in 53 human body sites produced by the Genotype-Tissue Expression (GTEx) project, and then augmented with genes from other databases to yield a comprehensive collection of protein-coding and noncoding transcripts. The construction of the new CHESS 3 database employed improved transcript assembly algorithms, a new machine learning classifier, and protein structure predictions to identify genes and transcripts likely to be functional and to eliminate those that appeared more likely to represent noise. The new catalog contains 41,356 genes on the GRCh38 reference human genome, of which 19,839 are protein-coding, and a total of 158,377 transcripts. These include 14,863 novel protein-coding transcripts. The total number of transcripts is substantially smaller than earlier versions due to improved transcriptome assembly methods and to a stricter protocol for filtering out noisy transcripts. Notably, CHESS 3 contains all of the transcripts in the MANE database, and at least one transcript corresponding to the vast majority of protein-coding genes in the RefSeq and GENCODE databases. CHESS 3 has also been mapped onto the complete CHM13 human genome, which gives a more-complete gene count of 43,773 genes and 19,968 protein-coding genes. The CHESS database is available at http://ccb.jhu.edu/chess.

genomics↗

The complete sequence of a human Y chromosome

The human Y chromosome has been notoriously difficult to sequence and assemble because of its complex repeat structure including long palindromes, tandem repeats, and segmental duplications1-3. As a result, more than half of the Y chromosome is missing from the GRCh38 reference sequence and it remains the last human chromosome to be finished4, 5. Here, the Telomere-to-Telomere (T2T) consortium presents the complete 62,460,029 base pair sequence of a human Y chromosome from the HG002 genome (T2T-Y) that corrects multiple errors in GRCh38-Y and adds over 30 million base pairs of sequence to the reference, revealing the complete ampliconic structures of TSPY, DAZ, and RBMY gene families; 41 additional protein-coding genes, mostly from the TSPY family; and an alternating pattern of human satellite 1 and 3 blocks in the heterochromatic Yq12 region. We have combined T2T-Y with a prior assembly of the CHM13 genome4 and mapped available population variation, clinical variants, and functional genomics data to produce a complete and comprehensive reference sequence for all 24 human chromosomes.

genomics↗

Perchlorate-Specific Proteomic Stress Responses of Debaryomyces hansenii Could Enable Microbial Survival in Martian Brines

If life exists on Mars, it would face several challenges including the presence of perchlorates, which destabilize biomacromolecules by inducing chaotropic stress. However, little is known about perchlorate toxicity for microorganism on the cellular level. Here we present the first proteomic investigation on the perchlorate-specific stress responses of the halotolerant yeast Debaryomyces hansenii and compare these to generally known salt stress adaptations. We found that the responses to NaCl and NaClO4-induced stresses share many common metabolic features, e.g., signaling pathways, elevated energy metabolism, or osmolyte biosynthesis. However, several new perchlorate-specific stress responses could be identified, such as protein glycosylation and cell wall remodulations, presumably in order to stabilize protein structures and the cell envelope. These stress responses would also be relevant for life on Mars, which - given the environmental conditions - likely developed chaotropic defense strategies such as stabilized confirmations of biomacromolecules and the formation of cell clusters.

cell biology↗

A reference-quality, fully annotated genome from a Puerto Rican individual

Until 2019, the human genome was available in only one fully-annotated version, GRCh38, which was the result of 18 years of continuous improvement and revision. Despite dramatic improvements in sequencing technology, no other genome was available as an annotated reference until 2019, when the genome of an Ashkenazi individual, Ash1, was released. In this study, we describe the assembly and annotation of a second individual genome, from a Puerto Rican individual whose DNA was collected as part of the Human Pangenome project. The new genome, called PR1, is the first true reference genome created from an individual of African descent. Due to recent improvements in both sequencing and assembly technology, and particularly to the use of the recently completed CHM13 human genome as a guide to assembly, PR1 is more complete and more contiguous than either GRCh38 or Ash1. Annotation revealed 37,755 genes (of which 19,999 are protein-coding), including 12 additional gene copies that are present in PR1 and missing from CHM13. 57 genes have fewer copies in PR1 than in CHM13, 9 map only partially, and 3 genes (all non-coding) from CHM13 are entirely missing from PR1.

genomics↗