Search bioRxiv⌕ Search

Biology subjects

Wong, B. S.-H.

Publications and source records attributed to Wong, B. S.-H..

2 recordsLinked to original sources

Few-shot in-context learning with large language models for antibody characterization

Large language models (LLMs) can learn new tasks by in-context learning (ICL), but it is unknown whether this ability reliably transfers to biological sequence classification. Here, we systematically evaluate how demonstration selection, shot count, and prompting strategies affect performance across 20 general-purpose LLMs. Using antibody characterization as a representative test case, we compare zero-shot, few-shot, and chain-of-thought (CoT) ICL on three classification tasks: humanness, antigen specificity, and isotype. Our results reveal a clear performance hierarchy: while zero-shot prompting performs near chance, few-shot prompting with randomly selected demonstrations improves performance, showing that LLMs can perform ICL using biological sequences from minimal supervision. However, matching protein-language model (pLM)-based classifier accuracy is only achieved when using label-diverse demonstrations drawn from antibodies similar to the query sequence. To leverage this insight, we introduce Sim-ICL, a framework that automatically retrieves such demonstrations. Using only 32-shot prompting, Sim-ICL achieves performance competitive with pLM-based classifiers, matching or outperforming them in two of the three tasks. Furthermore, reasoning-oriented prompts yield marginal gains and often produce fluent but biologically incorrect rationales, suggesting that current CoT explanations function as after-the-fact rationalizations rather than capturing mechanistic determinants of antibody properties. From these experiments, we derive practical design principles for ICL on biological sequences: use similarity-based, label-diverse demonstrations and modest shot counts, and treat reasoning prompts primarily as post hoc narratives rather than drivers of performance. Sim-ICL implements these principles in a streamlined, prompt-based framework for antibody sequence classification and, in principle, could be adapted to other biological sequence tasks.

bioinformatics↗

Multi-Omic Analysis of Tyrophagus putrescentiae Reveals Insights into the Allergen Complexity of Storage Mites

BackgroundThe storage mite Tyrophagus putrescentiae is one of the major mites causing allergies in Chinese and Korean populations, but its allergen profile in incomplete when compared with that of house dust mites. Multiple genome-based methods have been introduced into the allergen study of mites and have enabled a better understanding of these medically important organisms. ObjectiveWe sought to reveal a comprehensive allergen profile of Tyrophagus putrescentiae and advance the allergen study of storage mites. MethodsBased on a high-quality assembled and annotated genome, an in silico analysis was performed by searching reference sequences to identify allergens. Immunoassay ELISA assessed the allergenicities of recombinant proteins. MALDI-TOF mass spectrometry identified the IgE-binding proteins. Comparative genomics analysis was employed for the important allergen gene families. ResultsA complete allergen profile of Tyrophagus putrescentiae was revealed, including thirty-seven allergen groups (up to Tyr p 42). Among them, five novel allergens were verified using the sera of allergy patients. Massive allergen homologs were identified as the result of gene duplications in genome evolution. Proteomic identification again revealed a wide range of allergen homologs. In the NPC2 family and GSTs, comparative analysis shed light on the expansion and diversification of the allergen groups. ConclusionUsing multi-omic approaches, the comprehensive allergen profile including massive homologs was disclosed in Tyrophagus putrescentiae, which revealed the allergen complexity of the storage mite and could ultimately facilitate the component-resolved diagnosis.

immunology↗