808 Swiss-adapted items that extend HAAS with self-graded D7 reliability and judged D8 security scores aimed at FINMA and nDSG deployment questions.
unassessed
| Category | composite |
|---|---|
| Subcategory | Swiss-adapted reliability proxy (D7) and adversarial security (D8) under HAAS v2 |
| Page status | unknown |
| Metric | D7 self-graded composite and D8 judged composite (PII-Scope 60%, leakage 40%) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 808 |
| Publisher | Fatih Uenal, University of Colorado Boulder |
Swiss-Bench 003 (SBP-003) adds two HAAS dimensions that SBP-002 did not report in its legal C% table. D7 is a self-graded reliability proxy on Swiss-adapted TruthfulQA, IFEval, SimpleQA and Needle-in-a-Haystack items. D8 is adversarial security: Swiss PII-Scope, system-prompt leakage, and a Swiss-German dialect comprehension probe. Items are written against FINMA Guidance 08/2024, the revised nDSG and OWASP LLM risks, in German, French, Italian and English. This is not a rerun of SBP-002's legal-advice tasks.
Zero-shot prompts via OpenRouter at provider-default decoding. D7 uses Inspect AI model_graded_fact with the evaluand as its own judge. D8 uses task-specific rubrics and Qwen3-235B as an external judge, with regex pre-checks for AHV/IBAN and prompt-substring leaks.
No model card in ModelSpec reports this benchmark yet.