ThreeCB

3CB (Catastrophic Cyber Capabilities Benchmark) scores LLM agents on capture-the-flag cyber offense challenges mapped to MITRE ATT&CK techniques.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategorycyber offense capture-the-flag
Page statusactive
Metricsuccess rate
Directionhigher_is_better
Unitfraction
Dataset size15
Dataset licenceCC BY-NC-SA 4.0 (paper/dataset, per arXiv)
PublisherApart Research

What it measures

3CB tests whether an LLM agent can perform offensive cyber tasks - reconnaissance, exploitation, privilege escalation, and related techniques - by solving capture-the-flag (CTF) challenges in a bash shell, each mapped to a MITRE ATT&CK technique.

Task format

Agent interacts with a containerized challenge environment (typically a bash shell) and must recover a flag string that proves successful exploitation.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub