Search bioRxiv⌕ Search

Biology subjects

Huang, Y.-N.

Publications and source records attributed to Huang, Y.-N..

3 recordsLinked to original sources

Data availability of open T-cell receptor repertoire data, a systematic assessment

The improvement of next-generation sequencing technologies has promoted the field of immunogenetics and produced numerous immunogenomics data. Modern data-driven research has the power to promote novel biomedical discoveries through secondary analysis of such data. Therefore, it is important to ensure data-driven research with great reproducibility and robustness for promoting a precise and accurate secondary analysis of the immunogenomics data. In scientific research, rigorous conduct in designing and conducting experiments is needed, specifically in scientific and articulate writing, reporting and interpreting results. It is also crucial to make raw data available, discoverable, and well described or annotated in order to promote future re-analysis of the data. In order to assess the data availability of published T cell receptor (TCR) repertoire data, we examined 11,918 TCR-Seq samples corresponding to 134 TCR-Seq studies ranging from 2006 to 2022. Among the 134 studies, only 38.1% had publicly available raw TCR-Seq data shared in public repositories. We also found a statistically significant association between the presence of data availability statements and the increase in raw data availability (p=0.014). Yet, 46.8% of studies with data availability statements failed to share the raw TCR-Seq data. There is a pressing need for the biomedical community to increase awareness of the importance of promoting raw data availability in scientific research and take immediate action to improve its raw data availability enabling cost-effective secondary analysis of existing immunogenomics data by the larger scientific community.

scientific communication and education↗

The systematic assessment of completeness of public metadata accompanying omics studies

Recent advances in high-throughput sequencing technologies have made it possible to collect and share a massive amount of omics data, along with its associated metadata. Enhancing metadata availability is critical to ensure data reusability and reproducibility and to facilitate novel biomedical discoveries through effective data reuse. Yet, incomplete metadata accompanying public omics data may hinder reproducibility and reusability by reducing sample interpretability and limiting secondary analyses. In this study, we performed a comprehensive assessment of metadata completeness shared in both scientific publications and/or public repositories by analyzing over 253 studies encompassing over 164 thousands samples, including both human and non-human mammalian studies. We observed that studies often omit over a quarter of important phenotypes, with an average of only 74.8% of them shared either in the text of publication or the corresponding repository. Notably, public repositories alone contained 62% of the metadata, surpassing the textual content of publications by 3.5%. Only 11.5% of studies completely shared all phenotypes, while 37.9% shared less than 40% of the phenotypes. Studies involving non-human samples were more likely to share metadata than studies involving human samples. We observed similar results on the extended dataset spanning 2.1 million samples across over 61,000 studies from the Gene Expression Omnibus repository. The limited availability of metadata reported in our study emphasizes the necessity for improved metadata sharing practices and standardized reporting. Finally, we discuss the numerous benefits of improving the availability and quality of metadata to the scientific community and beyond, supporting data-driven decision-making and policy development in the field of biomedical research. This work provides a scalable framework for evaluating metadata availability and may help guide future policy and infrastructure development.

bioinformatics↗

Virtual meetings promise to eliminate the geographical and administrative barriers and increase accessibility, diversity, and inclusivity

The COVID-19 pandemic brought a new set of unprecedented challenges not only for healthcare, education, and everyday jobs but also in terms of academic conferences. In this study, we investigate the effect of the broad adoption of virtual platforms for academic conferences as a response to COVID-19 restrictions. We show that virtual platforms enable higher participation from underrepresented minority groups, increased inclusion, and broader geographic distribution. We also discuss emerging challenges associated with the virtual conference format resulting in a decreased engagement of social activities, limited possibilities of cross-fertilization between participants, and reduced peer-to-peer interactions. Lastly, we conclude that a novel comprehensive approach needs to be adopted by the conference organizers to ensure increased accessibility, diversity, and inclusivity of post-pandemic conferences. Our findings provide evidence favoring a hybrid format for future conferences, marrying the strength of both in-person and virtual platforms.

scientific communication and education↗