Bench-CoE

Bench-CoE evaluates routing and collaboration among specialist language and multimodal experts using benchmark-derived training data.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Metrictask performance
Directionhigher_is_better
Unitscore
PublisherBench-CoE authors

What it measures

Bench-CoE studies whether a router can assign each query to suitable expert models and improve aggregate performance through collaboration. It covers language and multimodal tasks under query-level and subject-level routing settings.

Task format

A mixture of language and multimodal benchmark queries routed to specialist experts.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub