Search bioRxivSearch

Biology subjects

Nagamatsu, S.

Publications and source records attributed to Nagamatsu, S..

2 recordsLinked to original sources

League of Brazilian Bioinformatics: a competition framework to promote scientific training

Backgroundthe scientific training to become a bioinformatician includes multidisciplinary abilities, which increase the challenges to professional development. Competition frameworkin order to improve and promote the ongoing training of the Brazilian bioinformatics community, we organize a national competition, with the main goal to develop human resources and abilities in Computational Biology at the national level. The competition framework was designed in three phases: 1) a one-day challenge composed of 60 multiple-choice questions covering Biology, Computer Science, and Bioinformatics knowledge; 2) five Computational Biology challenges to be solved in three days; and 3) development of an original project evaluated during the 15th X-meeting. Resultsthe first edition of the League of Brazilian Bioinformatics (LBB) counted 168 competitors and 59 groups, distributed into undergraduate students (14.4%), graduate students (12.6% master and 16.8%, Ph.D.), and other professional fields. The first phase selected 46 teams to proceed in the competition, while the second phase selected the three top-performing teams. Conclusionduring the competition, we were able to stimulate teamwork in the main areas of Bioinformatics, with the engagement of all research-level competitors. Furthermore, we identified opportunities to deliver and offer better training to the community and we intend to apply the acquired experience in the second edition of the LBB, which will occur in 2021. Supplementary informationSupplementary data are available at Bioinformatics

scientific communication and education

Expected Genotype Quality and Diploidized Marker Data from Genotyping-by-Sequencing of Urochloa spp. Tetraploids

Although genotyping-by-sequencing (GBS) is a well-established marker technology in diploids, the development of best practices for tetraploid species is a topic of current research. We determined the theoretical relationship between read depth and expected genotype quality (EGQ) for tetraploid vs. diploidized genotype calls. If the GBS method has 1% error, then 17 reads are needed to classify tetraploid samples as heterozygous vs. homozygous with 95% accuracy, compared with 63 reads to determine allele dosage. We developed an R script to convert tetraploid GBS data in Variant Call Format (VCF) into diploidized genotype calls and applied it to 267 interspecific hybrids of the tetraploid forage grass Urochloa (syn. Brachiaria). When reads were aligned to a mock reference genome created from GBS data of the U. brizantha cultivar Marandu, 25,678 bi-allelic SNPs were discovered, compared to approximately 3000 SNPs when aligning to the closest true reference genomes, Setaria viridis and S. italica. Crossvalidation revealed that missing genotypes were imputed by the Random Forest method with a median accuracy of 0.85, regardless of heterozygote frequency. Using the Urochloa spp. hybrids, we illustrated how filtering samples based only on GQ creates genotype bias; a depth threshold with corresponding EGQ equal to the GQ threshold is also needed, regardless of whether genotypes are called using a diploidized or allele dosage model.

genomics