SeedBench

2,264 expert-validated questions across 11 task types simulating gene retrieval, gene-function analysis and variety breeding for rice, scored by accuracy, macro-F1 or ROUGE-L per task.

Also known as: SeedBench: A Multi-task Benchmark for Evaluating Large Language Models in Seed Science

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryseed science and rice breeding decision support
Page statusactive
Metricaccuracy / macro-F1 / ROUGE-L (task-dependent)
Directionhigher_is_better
Unit%
Dataset size2264
Dataset licenceGPL-3.0 (code repository; a separate data licence was not confirmed)
PublisherShanghai AI Laboratory (InternScience / open-sciencelab)

What it measures

SeedBench tests whether a model can support the decision-making stages a seed breeder works through: retrieving gene information, analysing gene function and regulation, and reasoning about variety breeding outcomes. Content is built from a corpus of roughly 308,727 breeding publications distilled to about 1.1 billion tokens, initially scoped to rice, with maize, soybean and wheat planned as future extensions. Tasks span three families: question answering (multiple choice, multiple answer, fill-in-the-blank, open generation), summarisation (plain summary and key-information extraction) and reading comprehension (multiple choice, multiple answer, fill-in-the-blank, generation and subcategory classification).

Task format

Mixed by task type: single- and multi-answer multiple choice, cloze-style fill-in-the-blank, free-text generation, extractive/abstractive summarisation, and classification. OpenCompass runs it as the `seedbench_gen` config, reading `instruction` and `question` fields against an `answer` field, applying different postprocessors per subcategory (1-1 through 3-5).

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub