SkillsBench

SkillsBench evaluates how well agent skills work and how effectively agents use them across practical tasks.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryagent skill use
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
PublisherBenchFlow AI

What it measures

Agent task success with and without selected skills, under the benchmark's task and tool-use protocol.

Task format

Agent interaction with task environments, tools and skill instructions.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub