{
 "body": "\n## What it measures\n\nSkillsBench evaluates whether agent skills improve practical task performance and whether agents can use those skills effectively. The official repository describes a benchmark comparing task success under skill-enabled agent execution. The task environments, tools and instructions are part of the protocol, so a score is meaningful only with the same setup.\n\n## How it is scored\n\nThe natural headline measure is task success rate. The official sources opened for this page do not establish a universal random or human baseline, a fixed item count, or all aggregation details. Report the exact task set, agent, skill configuration, tool permissions, time budget and failure handling.\n\n## Dataset and licence\n\nThe dataset card and repository identify public SkillsBench materials, but the opened pages do not establish a single licence, stable split definition or complete benchmark size. Confirm those terms from the current release before redistribution or comparison.\n\n## Who publishes it\n\nSkillsBench is maintained by BenchFlow AI in the `benchflow-ai/skillsbench` repository. A complete paper citation and author list were not established from the official sources opened for this page.\n\n## Lineage\n\nThe repository is the primary benchmark implementation. `skillsbench-leaderboard` is a related results artifact and should not be treated as a separate task suite without verifying its contents.\n\n## Saturation and contamination\n\nPublic skill definitions, task descriptions and results create contamination opportunities. The opened sources do not report a contamination audit, private holdout or current saturation ceiling.\n\n## How to run it\n\nFollow the official repository instructions, preserving the specified agent, skill files, tools, task environments and timeouts. Record whether skills were enabled, the exact skill version, task-level outcomes and any infrastructure failures.\n\n## Reading the numbers\n\nHigher success means more tasks completed under the selected skill and agent configuration. Differences can reflect tool access, prompt and skill versions, environment reliability or task selection; compare only runs with matched protocols.\n\n## Protocol cautions\n\nSkill benchmarks are sensitive to the exact skill text and to tool and filesystem permissions. Record the skill package revision, agent model, context window, task seed and environment image. If a task fails because a tool is unavailable, distinguish that infrastructure failure from an unsuccessful attempt to use the skill.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "skillsbench",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "agentic",
  "contamination": {
   "note": "The benchmark and skill materials are public; no contamination audit or private rotating holdout was established in the opened sources.",
   "risk": "medium"
  },
  "dataset": {
   "languages": [],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": true,
   "size": null,
   "size_note": "The official repository does not establish a stable total count in the source reviewed here.",
   "splits": "Unknown from the opened repository and dataset card.",
   "url": "https://huggingface.co/datasets/benchflow/skillsbench"
  },
  "freshness": {
   "researched": "2026-09-09",
   "researched_by": "GPT-5.6 Luna, luna-stream-c-005 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "",
   "opencompass": "",
   "other": "Use the official repository's agent and skill execution protocol."
  },
  "id": "skillsbench",
  "last_updated": "",
  "leaderboard_url": "https://www.vals.ai/benchmarks",
  "lineage": {
   "family": "agent skill evaluation",
   "predecessor": "",
   "successors": [],
   "variants": [
    "skillsbench-leaderboard"
   ]
  },
  "measures": "Agent task success with and without selected skills, under the benchmark's task and tool-use protocol.",
  "metric": {
   "baseline_note": "The official repository describes skill effectiveness evaluation; no universal random or human baseline was established in the source reviewed.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": 100,
   "name": "task success rate",
   "random_baseline": null,
   "unit": "%"
  },
  "name": "SkillsBench",
  "page_kind": "benchmark",
  "paper": {
   "arxiv": "",
   "title": "",
   "url": "",
   "year": null
  },
  "publisher": {
   "authors": [],
   "org": "BenchFlow AI",
   "url": "https://github.com/benchflow-ai/skillsbench"
  },
  "released": "",
  "repo_url": "https://github.com/benchflow-ai/skillsbench",
  "saturation": {
   "as_of": "",
   "note": "No current saturation ceiling was established in the opened official sources.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-09",
    "title": "Official SkillsBench repository",
    "url": "https://github.com/benchflow-ai/skillsbench"
   },
   {
    "accessed": "2026-09-09",
    "title": "SkillsBench dataset card",
    "url": "https://huggingface.co/datasets/benchflow/skillsbench"
   }
  ],
  "status": "active",
  "subcategory": "agent skill use",
  "summary": "SkillsBench evaluates how well agent skills work and how effectively agents use them across practical tasks.",
  "tags": [
   "agents",
   "tools",
   "skills",
   "task-success"
  ],
  "task_format": "Agent interaction with task environments, tools and skill instructions."
 }
}