187 Korean CSAT exam questions across six subjects, each with a recorded human student accuracy; GPT-4 already beat the average human score in the original 2023 evaluation.
unassessed
| Category | domain |
|---|---|
| Subcategory | Korean College Scholastic Ability Test (CSAT) multi-subject exam questions |
| Page status | active |
| Metric | accuracy and length-normalized accuracy (acc, acc_norm) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 187 |
| Dataset licence | Not an open licence: per the Hugging Face dataset card, "the copyright of this material belongs to the Korea Institute for Curriculum and Evaluation (한국교육과정평가원) and may be used for research purposes only." |
| Publisher | HAE-RAE project (open community initiative built around the Polyglot-Ko model family) |
csatqa is EleutherAI lm-evaluation-harness's task group for CSAT-QA, a set of multiple-choice questions manually collected by the HAE-RAE project from South Korea's College Scholastic Ability Test (CSAT, or Suneung), the standardized exam required for university admission. The harness's group covers six subject categories: Writing (WR), Grammar (GR), Reading Comprehension: Science (RCS), Reading Comprehension: Social Science (RCSS), Reading Comprehension: Humanities (RCH), and Literature (LI); each question presents a Korean-language passage or prompt and five numbered answer options. HAE-RAE separately released a larger, 936-question "full" collection of CSAT items spanning exams from 2007 to 2022, but that full collection is not wired into lm-evaluation-harness -- only the smaller, six-category subset described here, chosen because it has attached human student-accuracy data, is.
Five-option multiple-choice question, answered zero-shot. The harness's own prompt is in Korean: "다음을 읽고 정답으로 알맞은 것을 고르시요" (read the following and choose the correct answer), followed by the context, question and five options, and the model completes "주어진 문제의 정답은" (the answer to the given question is) with an option number.
No model card in ModelSpec reports this benchmark yet.