Search bioRxivSearch

Biology subjects

Delgado, J.

Publications and source records attributed to Delgado, J..

2 recordsLinked to original sources

An introduction to MPEG-G, the new ISO standard for genomic information representation

The MPEG-G standardization initiative is a coordinated international effort to specify a compressed data format that enables large scale genomic data to be processed, transported and shared. The standard consists of a set of specifications (i.e., a book) describing: i) a nor-mative format syntax, and ii) a normative decoding process to retrieve the information coded in a compliant file or bitstream. Such decoding process enables the use of leading-edge com-pression technologies that have exhibited significant compression gains over currently used formats for storage of unaligned and aligned sequencing reads. Additionally, the standard provides a wealth of much needed functionality, such as selective access, data aggregation, ap-plication programming interfaces to the compressed data, standard interfaces to support data protection mechanisms, support for streaming and a procedure to assess the conformance of implementations. ISO/IEC is engaged in supporting the maintenance and availability of the standard specification, which guarantees the perenniality of applications using MPEG-G. Fi-nally, the standard ensures interoperability and integration with existing genomic information processing pipelines by providing support for conversion from the FASTQ/SAM/BAM file formats.\n\nIn this paper we provide an overview of the MPEG-G specification, with particular focus on the main advantages and novel functionality it offers. As the standard only specifies the decoding process, encoding performance, both in terms of speed and compression ratio, can vary depending on specific encoder implementations, and will likely improve during the lifetime of MPEG-G. Hence, the performance statistics provided here are only indicative baseline examples of the technologies included in the standard.

bioinformatics

VarQ: a tool for the structural analysis of Human Protein Variants

Understanding the functional effect of Single Amino acid Substitutions (SAS), derived from the occurrence of single nucleotide variants (SNVs), and their relation to disease development is a major issue in clinical genomics. Even though there are several bioinformatic algorithms and servers that predict if a SAS can be pathogenic or not they give little or non-information on the actual effect on the protein function. Moreover, many of these algorithms are able to predict an effect that no necessarily translates directly into pathogenicity. VarQ Web Server is an online tool that given an UniProt id automatically analyzes known and user provided SAS for their effect on protein activity, folding, aggregation and protein interactions among others. VarQ assessment was performed over a set of previously manually curated variants, showing its ability to correctly predict the phenotypic outcome and its underlying cause. This resource is available online at http://varq.qb.fcen.uba.ar/.\n\nContact: lradusky@qb.fcen.uba.ar\n\nSupporting Information & Tutorials may be found in the webpage of the tool.

bioinformatics