Swiss-Bench SBP-002

Trilingual Swiss regulatory-compliance eval of 395 expert items; a three-judge panel found even the top model only 38.2% correct under zero retrieval.

Also known as: Swiss-Bench 002, SBP-002

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryapplied Swiss regulatory compliance (FINMA, Legal-CH, EFK) in DE/FR/IT
Page statusunknown
Metriccorrect rate (C%) from majority-vote C/P/I grades
Directionhigher_is_better
Unit%
Dataset size395
Dataset licenceCC-BY-NC-SA-4.0
PublisherFatih Uenal, University of Colorado Boulder

What it measures

Swiss-Bench SBP-002 tests whether a frontier LLM can give usable Swiss regulatory advice from parameters alone. Items cover three domains: FINMA financial supervision, Swiss federal law (Legal-CH, including nDSG and Code of Obligations plus EU AI Act impact on Swiss firms), and Swiss Federal Audit Office (EFK) control scenarios. Seven task types mix regulatory Q&A, hallucination detection, gap analysis, jurisdiction discrimination, statutory interpretation, case analysis and legal translation. The author contrasts this applied-compliance setting with Swiss exam recall (LEXam) and Swiss legal translation (SwiLTra-Bench).

Task format

Single-turn user prompt, no system prompt and no retrieval, temperature 0, max 4,096 output tokens. Languages are German, French and Italian. Judges return three numeric dimensions; a grade is computed from those numbers.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub