Chinese-native agent eval of tool use, planning and memory across ten tasks; GPT-4 led the 2023 table at 80.56, with no published item count or licence.
unassessed
| Category | agentic |
|---|---|
| Subcategory | Chinese-native agent skills: tool use, task planning, and long/short-term memory |
| Page status | unknown |
| Metric | published total score plus three capability scores and ten task scores |
| Direction | higher_is_better |
| Unit | % |
| Publisher | CLUE / CLUEbenchmark |
SuperCLUE-Agent tests whether a Chinese LLM can act as an agent on native Chinese tasks rather than translated English agent suites. The publisher groups ten tasks into three skills: tool use (call, retrieve and plan APIs, plus general tools such as search, browsing, files and databases), task planning (decomposition, self-reflection and chain-of-thought), and long/short-term memory (multi-document QA, long-turn dialogue and in-context example learning). Prompts are Chinese user requests that require API choice, multi-step plans or recall across documents and dialogue turns.
Open-ended Chinese agent prompts spanning the ten tasks above. The official page and GitHub README show worked examples as screenshots; they do not publish a machine-readable item schema, shot count or tool-sandbox specification.
No model card in ModelSpec reports this benchmark yet.