OpenCompass biodata (biology-instruction)

OpenCompass suite of 21 biology tasks on opencompass/biology-instruction, scored with MCC, correlation, R², AUC, accuracy and EC-number Fmax.

Also known as: biology-instruction

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryDNA, RNA, protein and multi-sequence property prediction
Page statusunknown
Metrictask-specific (MCC, PCC, Spearman, R², AUC, accuracy, EC Fmax, Mixed)
Directionhigher_is_better
PublisherOpenCompass

What it measures

OpenCompass biodata is a generation benchmark over biological sequence problems. Each item is a prompt that asks a model to predict a property of DNA, RNA, protein, or a multi-sequence interaction, then to put the answer in \\boxed{}. Tasks include DNA classification (cpd, emp, pd, transcription-factor binding), enhancer activity regression, RNA isoform and ribosome-loading regression, RNA modification labels, protein solubility, fluorescence, stability, thermostability, enzyme commission (EC) numbers, antibody–antigen and RNA–protein interaction, and siRNA efficiency. The intended skill is biological prediction from sequence context, not general reading comprehension.

Task format

Zero-shot generation with a biology-expert system prompt. Two OpenCompass configs differ only in prompt template class (PromptTemplate vs RawPromptTemplate). Answers are parsed from \\boxed{} or from a JSON object for dict-valued labels.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub