SuperCLUE-Agent

Chinese-native agent eval of tool use, planning and memory across ten tasks; GPT-4 led the 2023 table at 80.56, with no published item count or licence.

Also known as: SuperCLUE Agent

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
SubcategoryChinese-native agent skills: tool use, task planning, and long/short-term memory
Page statusunknown
Metricpublished total score plus three capability scores and ten task scores
Directionhigher_is_better
Unit%
PublisherCLUE / CLUEbenchmark

What it measures

SuperCLUE-Agent tests whether a Chinese LLM can act as an agent on native Chinese tasks rather than translated English agent suites. The publisher groups ten tasks into three skills: tool use (call, retrieve and plan APIs, plus general tools such as search, browsing, files and databases), task planning (decomposition, self-reflection and chain-of-thought), and long/short-term memory (multi-document QA, long-turn dialogue and in-context example learning). Prompts are Chinese user requests that require API choice, multi-step plans or recall across documents and dialogue turns.

Task format

Open-ended Chinese agent prompts spanning the ten tasks above. The official page and GitHub README show worked examples as screenshots; they do not publish a machine-readable item schema, shot count or tool-sandbox specification.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub