EQ-Bench (Spanish)

Spanish translation and cultural adaptation of EQ-Bench version 2's 168-171 question emotion-intensity rating task, released by the Barcelona Supercomputing Center.

Also known as: EQ-bench_es

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryreasoning
Subcategoryemotional and social intelligence (dialogue emotion-intensity prediction), Spanish translation
Page statusactive
MetricEQ-Bench score (distance from reference ratings)
Directionhigher_is_better
Unitpoints
Dataset size168
Dataset licenceCC BY 4.0
PublisherBarcelona Supercomputing Center (BSC), Language Technologies Unit

What it measures

eq_bench_es shows a model a short dialogue translated and culturally adapted into Spanish, then asks it to rate the intensity (0-10) of four named emotions one character is likely feeling at the end of the scene -- the same task format as English EQ-Bench. The Barcelona Supercomputing Center's Language Technologies Unit produced the adaptation: converting adjectival emotion labels to nominal forms to avoid Spanish grammatical-gender ambiguity (e.g. "proud" to "orgullo"), replacing Anglo-Saxon character names with Spanish ones, and unifying emotion labels that were equivalent but differently inflected in the English original.

Task format

Given a Spanish dialogue and four named emotions, output an intensity rating from 0 to 10 for each in a fixed format; scored by the same distance-from-reference formula as EQ-Bench v2 (see the family page). lm-evaluation-harness runs it as task `eqbench_es`, generating greedily at temperature 0.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub