Search bioRxiv⌕ Search

bioRxiv · 10.1101/2021.03.06.434213

Nurturing Diversity and Inclusion in AI in Biomedicine through a Virtual Summer Program for High School Students

Abstract

Artificial Intelligence (AI) has the power to improve our lives through a wide variety of applications, many of which fall into the healthcare space; however, a lack of diversity is contributing to flawed systems that perpetuate gender and racial biases, and limit how broadly AI can help people. The UCSF AI4ALL program was established in 2019 to address this issue by promoting diversity and inclusion in AI. The program targets high school students from underrepresented backgrounds in AI and gives them a chance to learn about AI with a focus on biomedicine. In 2020, the UCSF AI4ALL three-week program was held entirely online due to the COVID-19 pandemic. Thus students participated virtually to gain experience with AI, interact with diverse role models in AI, and learn about advancing health through AI. Specifically, they attended lectures in coding and AI, received an in-depth research experience through hands-on projects exploring COVID-19, and engaged in mentoring and personal development sessions with faculty, researchers, industry professionals, and undergraduate and graduate students, many of whom were women and from underrepresented racial and ethnic backgrounds. At the conclusion of the program, the students presented the results of their research projects at our final symposium. Comparison of pre- and post-program survey responses from students demonstrated that after the program, significantly more students were familiar with how to work with data and to evaluate and apply machine learning algorithms. There was also a nominally significant increase in the students knowing people in AI from historically underrepresented groups, feeling confident in discussing AI, and being aware of careers in AI. We found that we were able to engage young students in AI via our online training program and nurture greater inclusion in AI.

Source connections

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Oskotsky, T., Bajaj, R., Burchard, J., Cavazos, T. B., Chen, I., Connell, W., Eaneff, S., Grant, T., Kanungo, I., Lindquist, K., Myers-Turnbull, D. J., Naing, Z. Z. C., Tang, A., Vora, B., Wang, J. X., Karim, I., Swadling, C., Yang, J., AI4ALL Student Cohort 2020,, Sirota, M.. 2021-03-08. Nurturing Diversity and Inclusion in AI in Biomedicine through a Virtual Summer Program for High School Students. https://doi.org/10.1101/2021.03.06.434213

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related preprints

Evaluating Large Language Models as Tools to Navigate Researchers in Rapidly Evolving Research Landscapes: A Case Study in Cancer Drug Response Prediction

Large Language Models (LLMs) have emerged as promising tools for assisting researchers in automating and accelerating the synthesis of literature reviews. However, their reliability is a significant concern due to issues like factual inaccuracies and hallucinations. The key question is whether LLMs can reliably provide comprehensive, up-to-date overviews and analyses. This study evaluates the performance of three leading LLMs (OpenAI's ChatGPT, Google's Gemini, and DeepSeek) on the complex task of generating a comprehensive survey paper on deep learning for cancer Drug Response Prediction (DRP). By testing both standard and Deep Research (DR) / Deep Think (DT) modes of LLMs with prompts of varying detail, this paper assesses key academic dimensions, including reference management, content quality, and analytical depth. Key findings reveal that while DR modes of LLMs significantly improve reliability by eliminating hallucinations, performance variations exist across models and prompts. A trade-off between reference quantity and integration quality was observed, and even the best-performing models lacked the analytical depth of human experts, often requiring extensive human supervision. The study concludes that LLMs currently serve as powerful assistive tools but still cannot replace the critical validation and synthesis provided by human researchers. Choosing the best LLM to use depends on the task in hand, while several strategies can be implemented to improve the produced output.

scientific communication and education↗

SynBio in 3D: the first synthetic genetic circuit as a 3D-printed STEM educational resource

Synthetic biology is a new area of science that operates at the intersection of engineering and biology and aims to design and synthesize living organisms and systems to perform new or improved functions. This novel biological area is considered of extreme relevance for the development of solutions to global problems. However, its teaching is often inaccessible to students, since many educational resources and methodological procedures are not available for understanding of complex molecular processes. On the other hand, digital fabrication tools, which allow the creation of 3D objects, are increasingly used for educational purposes, and several computational structures of molecular components commonly used in synthetic biology processes are deposited in open databases. Therefore, we hypothesize that the creation of biomolecular structures models by handling 3D physical objects using computer-assisted design (CAD) and additive fabrication based on 3D printing could help professors in synthetic biology teaching. In this sense, the present work describes the design and 3D print of the molecular models of the first synthetic genetic circuit, the toggle switch, which can be freely downloaded and used by teachers to facilitate the training of STEM students in synthetic biology.

scientific communication and education↗

Of Problems and Opportunities - How to Treat and How to not Treat Crystallographic Fragment-Screening Data

In their recent commentary in Protein Science, Jaskolski et al. analyze three randomly picked diffraction data sets from fragment-screening group depositions from the PDB and, based on that, claim that such data are principally problematic. We demonstrate here that if such data are treated properly, none of the proclaimed criticisms persist.

scientific communication and education↗