FinanceIQ

7,173 Chinese multiple-choice questions across 10 financial-licensing exam subjects, GPT-4-paraphrased and option-shuffled by its publisher specifically to resist pretraining leakage.

Also known as: FinanceIQ(中文金融领域知识评估数据集)

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryChinese financial-qualification exam multiple-choice knowledge (10 subjects, 36 sub-fields)
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
Dataset size7173
Dataset licenceCC BY-NC-SA 4.0, stated identically by the GitHub repository's own README and the Hugging Face dataset card's licence tag
PublisherDu Xiaoman (Duxiaoman-DI), a Chinese fintech company, as part of its open-source XuanYuan (轩辕) financial large-language-model project

What it measures

FinanceIQ tests a model's command of the specialist knowledge covered by China's major financial professional-qualification exams: certified public accountant (CPA), tax accountant, economist, banking/securities/fund/futures/insurance qualification exams, certified financial planner, and the "financial mathematics" subject from the actuarial exam (added specifically to test harder quantitative material). Each of the 10 subjects is further broken into 36 finer sub-fields. It is a single-turn, Chinese-language, four-option multiple-choice knowledge task, positioned by its publisher as filling a gap it says general Chinese benchmarks such as C-Eval and CMMLU cover only thinly for finance-specific professional knowledge.

Task format

Four-option multiple-choice question (A-D), one correct answer, drawn from one of 10 financial subjects; base models are evaluated five-shot from a fixed 5-question development set per subject, chat models zero-shot, and the answer letter is extracted from the model's generated free text rather than read off token log-probabilities.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub