2,264 expert-validated questions across 11 task types simulating gene retrieval, gene-function analysis and variety breeding for rice, scored by accuracy, macro-F1 or ROUGE-L per task.
unassessed
| Category | domain |
|---|---|
| Subcategory | seed science and rice breeding decision support |
| Page status | active |
| Metric | accuracy / macro-F1 / ROUGE-L (task-dependent) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 2264 |
| Dataset licence | GPL-3.0 (code repository; a separate data licence was not confirmed) |
| Publisher | Shanghai AI Laboratory (InternScience / open-sciencelab) |
SeedBench tests whether a model can support the decision-making stages a seed breeder works through: retrieving gene information, analysing gene function and regulation, and reasoning about variety breeding outcomes. Content is built from a corpus of roughly 308,727 breeding publications distilled to about 1.1 billion tokens, initially scoped to rice, with maize, soybean and wheat planned as future extensions. Tasks span three families: question answering (multiple choice, multiple answer, fill-in-the-blank, open generation), summarisation (plain summary and key-information extraction) and reading comprehension (multiple choice, multiple answer, fill-in-the-blank, generation and subcategory classification).
Mixed by task type: single- and multi-answer multiple choice, cloze-style fill-in-the-blank, free-text generation, extractive/abstractive summarisation, and classification. OpenCompass runs it as the `seedbench_gen` config, reading `instruction` and `question` fields against an `answer` field, applying different postprocessors per subcategory (1-1 through 3-5).
No model card in ModelSpec reports this benchmark yet.