SWE-Bench-Verified-O1-reasoning-high-results is a benchmark task documented by its cited evaluation harness.
unassessed
| Category | knowledge |
|---|---|
| Subcategory | benchmark task |
| Page status | active |
| Metric | accuracy |
| Direction | higher_is_better |
| Unit | percent |
SWE-Bench-Verified-O1-reasoning-high-results is a benchmark task documented by the evaluation integration.
Text input with task-specific prediction or generation output.
No model card in ModelSpec reports this benchmark yet.