SWE-Together is a benchmark task documented by its cited evaluation harness.
unverified
This page is not in the default catalogue. Evidence required by the catalogue contract is missing or was not approved by a reviewer. That is a statement about the evidence we hold, not a claim that the benchmark is stale or illegitimate.
Recorded reasons:
model glm-5.2 is absent from reference set for domain agentic
model gpt-5.5 is absent from reference set for domain agentic
no qualifying current frontier/open coverage from different organizations
review is not approved
usefulness is not established
usefulness is unknown
Category
knowledge
Subcategory
benchmark task
Page status
active
Metric
accuracy
Direction
higher_is_better
Unit
percent
What it measures
SWE-Together is a benchmark task documented by the evaluation integration.
Task format
Text input with task-specific prediction or generation output.
Models reporting this benchmark
No model card in ModelSpec reports this benchmark yet.