Search bioRxivSearch

Biology subjects

Hsu, W.-L.

Publications and source records attributed to Hsu, W.-L..

2 recordsLinked to original sources

UniLoc: A universal protein localization site predictor for eukaryotes and prokaryotes

There is a growing gap between protein subcellular localization (PSL) data and protein sequence data, raising the need for computation methods to rapidly determine subcellular localizations for uncharacterized proteins. Currently, the most efficient computation method involves finding sequence-similar proteins (hereafter referred to as similar proteins) in the annotated database and transferring their annotations to the target protein. When a sequence-similarity search fails to find similar proteins, many PSL predictors adopt machine learning methods for the prediction of localization sites. We proposed a universal protein localization site predictor - UniLoc - to take advantage of implicit similarity among proteins through sequence analysis alone. The notion of related protein words is introduced to explore the localization site assignment of uncharacterized proteins. UniLoc is found to identify useful template proteins and produce reliable predictions when similar proteins were not available.

bioinformatics

The accuracy and bias of single-step genomic prediction for populationsunder selection

In single-step analyses, missing genotypes are explicitly or implicitly imputed, and this requires centering the observed genotypes, ideally using the mean of the unselected founders. If genotypes are only available on selected individuals, centering on the unselected founder mean is impossible. Here, computer simulation is used to study an alternative analysis that does not require centering genotypes but fits the mean {micro}g of unselected individuals as a fixed effect. To improve numerical properties of the analysis, centering the entire matrix of observed and imputed genotypes, using their sample means can be done in addition to fitting {micro}g. Starting with observed diplotypes from 721 cattle, a 5 generation population was simulated with sire selection to produce 40,000 individuals with phenotypes of which the 1,000 sires had genotypes. The next generation of 8,000 genotyped individuals was used for validation. Evaluations were undertaken: with (J) or without (N) {micro}g when marker covariates were not centered; and with (JC) or without (C) {micro}g when all marker covariates were centered. A pedigree based evaluation was less accurate than genomic analyses. Centering did not influence accuracy of genomic prediction, but fitting {micro}g did. Accuracies were improved when the panel comprised only QTL, models JC and J had accuracies of 99.2%; and models C and N had accuracies of 85.6%. When only markers were in the panel, the 4 models had accuracies of 63.9%. In panels that included causal variants, fitting {micro}g in the model improved accuracy, but had little impact when the panel contained only markers.

genomics