Swiss-Bench 003

808 Swiss-adapted items that extend HAAS with self-graded D7 reliability and judged D8 security scores aimed at FINMA and nDSG deployment questions.

Also known as: Swiss-Bench SBP-003, SBP-003

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
SubcategorySwiss-adapted reliability proxy (D7) and adversarial security (D8) under HAAS v2
Page statusunknown
MetricD7 self-graded composite and D8 judged composite (PII-Scope 60%, leakage 40%)
Directionhigher_is_better
Unit%
Dataset size808
PublisherFatih Uenal, University of Colorado Boulder

What it measures

Swiss-Bench 003 (SBP-003) adds two HAAS dimensions that SBP-002 did not report in its legal C% table. D7 is a self-graded reliability proxy on Swiss-adapted TruthfulQA, IFEval, SimpleQA and Needle-in-a-Haystack items. D8 is adversarial security: Swiss PII-Scope, system-prompt leakage, and a Swiss-German dialect comprehension probe. Items are written against FINMA Guidance 08/2024, the revised nDSG and OWASP LLM risks, in German, French, Italian and English. This is not a rerun of SBP-002's legal-advice tasks.

Task format

Zero-shot prompts via OpenRouter at provider-default decoding. D7 uses Inspect AI model_graded_fact with the evaluand as its own judge. D8 uses task-specific rubrics and Qwen3-235B as an external judge, with regex pre-checks for AHV/IBAN and prompt-substring leaks.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub