Hungarian National HS Finals Exam (Mathematics)

Keiran Paster's late-2023 test of language models against that year's Hungarian national high-school mathematics finals, evaluated by hand because the exam was published after every tested model's training cutoff.

Also known as: Hungarian Math Exam, HungarianExamMath, Testing Language Models on a Held-Out High School National Finals Exam

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorymath
Subcategoryheld-out contamination check: the 2023 Hungarian national high-school mathematics finals, hand-graded
Page statusactive
Metricpercentage of exam points earned, hand-graded against the official rubric
Directionhigher_is_better
Unit%
Dataset size33
PublisherIndependent project (Keiran Paster); source exam published by Hungary's national exam authority (Oktatási Hivatal, distributed via educatio.hu)

What it measures

This benchmark is the 2023 Hungarian national high-school mathematics finals (the "matematika érettségi," standard/"közép" level), given to language models as a free-response test rather than reused from a standard training corpus. Each of 33 problems asks for a worked solution -- domain restrictions, combinatorics, percentage change, vector geometry, number bases, inequalities and similar secondary-school topics -- translated into English from the original Hungarian. The point was never the mathematics itself, which is no harder than GSM8K or MATH; it is that the exam was sat and published in May 2023, after the training-data cutoff of every model in the original comparison, so a model's score could not be inflated by having memorised the answers.

Task format

Single-turn, free-response generation: given one problem, the model writes a worked solution ending in a final answer. The reference OpenCompass config prepends four generic worked examples (unrelated in content to the Hungarian exam) purely to demonstrate the expected answer format, then samples one completion per problem. There is no multiple-choice option and no programmatic answer key; every response is graded by a person against the exam's official point-by-point rubric.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub