Keiran Paster's late-2023 test of language models against that year's Hungarian national high-school mathematics finals, evaluated by hand because the exam was published after every tested model's training cutoff.
unassessed
| Category | math |
|---|---|
| Subcategory | held-out contamination check: the 2023 Hungarian national high-school mathematics finals, hand-graded |
| Page status | active |
| Metric | percentage of exam points earned, hand-graded against the official rubric |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 33 |
| Publisher | Independent project (Keiran Paster); source exam published by Hungary's national exam authority (Oktatási Hivatal, distributed via educatio.hu) |
This benchmark is the 2023 Hungarian national high-school mathematics finals (the "matematika érettségi," standard/"közép" level), given to language models as a free-response test rather than reused from a standard training corpus. Each of 33 problems asks for a worked solution -- domain restrictions, combinatorics, percentage change, vector geometry, number bases, inequalities and similar secondary-school topics -- translated into English from the original Hungarian. The point was never the mathematics itself, which is no harder than GSM8K or MATH; it is that the exam was sat and published in May 2023, after the training-data cutoff of every model in the original comparison, so a model's score could not be inflated by having memorised the answers.
Single-turn, free-response generation: given one problem, the model writes a worked solution ending in a final answer. The reference OpenCompass config prepends four generic worked examples (unrelated in content to the Hungarian exam) purely to demonstrate the expected answer format, then samples one completion per problem. There is no multiple-choice option and no programmatic answer key; every response is graded by a person against the exam's official point-by-point rubric.
No model card in ModelSpec reports this benchmark yet.