Benchmarking the Residual

Benchmarking the Residual proposes a horizon residual for separating ordinary stage errors from degradation caused by long-horizon execution.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Metrichorizon residual
Directionhigher_is_better
Unitlog-ratio
PublisherBenchmarking the Residual authors

What it measures

The work evaluates long-horizon agent tasks against matched short-task stages. It asks whether full-task success differs from a baseline prediction formed by composing short-stage success, while keeping the agent configuration fixed.

Task format

Agent trajectories with staged checkpoints, short-task controls, and full-task rollouts.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub