Patronus AI's open-book financial QA test of 10,231 questions over 40 companies' public filings; only a 150-question human-graded sample is publicly released with answers.
unassessed
| Category | domain |
|---|---|
| Subcategory | open-book financial question answering (SEC filings) |
| Page status | active |
| Metric | human-graded correct / incorrect / failed-to-answer rate |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 150 |
| Dataset licence | not established |
| Publisher | Patronus AI, with Contextual AI and Stanford University |
FinanceBench tests whether a model can answer a financial-analyst-style question about a publicly traded company when given the relevant excerpt from that company's own SEC filing or earnings report as context. Questions range from simple extraction ("what was Boeing's FY2022 cost of goods sold?") to questions requiring a short calculation over reported figures. It is single-turn, open-book (evidence is supplied, not retrieved from scratch, in the paper's own "oracle" setting; other tested configurations force the model to retrieve the right passage itself from a vector store or a long context window), text-only, and English-language.
Given a question and either a directly supplied evidence excerpt, a retrieved passage, or a full document (depending on configuration), the model must produce a short free-text answer, typically a figure, a yes/no judgement, or a brief explanation, which is then graded against a human-written gold answer.
No model card in ModelSpec reports this benchmark yet.