Benchmarking the Residual proposes a horizon residual for separating ordinary stage errors from degradation caused by long-horizon execution.
unassessed
| Category | long-context |
|---|---|
| Metric | horizon residual |
| Direction | higher_is_better |
| Unit | log-ratio |
| Publisher | Benchmarking the Residual authors |
The work evaluates long-horizon agent tasks against matched short-task stages. It asks whether full-task success differs from a baseline prediction formed by composing short-stage success, while keeping the agent configuration fixed.
Agent trajectories with staged checkpoints, short-task controls, and full-task rollouts.
No model card in ModelSpec reports this benchmark yet.