Spanish translation and cultural adaptation of EQ-Bench version 2's 168-171 question emotion-intensity rating task, released by the Barcelona Supercomputing Center.
unassessed
| Category | reasoning |
|---|---|
| Subcategory | emotional and social intelligence (dialogue emotion-intensity prediction), Spanish translation |
| Page status | active |
| Metric | EQ-Bench score (distance from reference ratings) |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 168 |
| Dataset licence | CC BY 4.0 |
| Publisher | Barcelona Supercomputing Center (BSC), Language Technologies Unit |
eq_bench_es shows a model a short dialogue translated and culturally adapted into Spanish, then asks it to rate the intensity (0-10) of four named emotions one character is likely feeling at the end of the scene -- the same task format as English EQ-Bench. The Barcelona Supercomputing Center's Language Technologies Unit produced the adaptation: converting adjectival emotion labels to nominal forms to avoid Spanish grammatical-gender ambiguity (e.g. "proud" to "orgullo"), replacing Anglo-Saxon character names with Spanish ones, and unifying emotion labels that were equivalent but differently inflected in the English original.
Given a Spanish dialogue and four named emotions, output an intensity rating from 0 to 10 for each in a fixed format; scored by the same distance-from-reference formula as EQ-Bench v2 (see the family page). lm-evaluation-harness runs it as task `eqbench_es`, generating greedily at temperature 0.
No model card in ModelSpec reports this benchmark yet.