FinanceBench

Patronus AI's open-book financial QA test of 10,231 questions over 40 companies' public filings; only a 150-question human-graded sample is publicly released with answers.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategoryopen-book financial question answering (SEC filings)
Page statusactive
Metrichuman-graded correct / incorrect / failed-to-answer rate
Directionhigher_is_better
Unit%
Dataset size150
Dataset licencenot established
PublisherPatronus AI, with Contextual AI and Stanford University

What it measures

FinanceBench tests whether a model can answer a financial-analyst-style question about a publicly traded company when given the relevant excerpt from that company's own SEC filing or earnings report as context. Questions range from simple extraction ("what was Boeing's FY2022 cost of goods sold?") to questions requiring a short calculation over reported figures. It is single-turn, open-book (evidence is supplied, not retrieved from scratch, in the paper's own "oracle" setting; other tested configurations force the model to retrieve the right passage itself from a vector store or a long context window), text-only, and English-language.

Task format

Given a question and either a directly supplied evidence excerpt, a retrieved passage, or a full document (depending on configuration), the model must produce a short free-text answer, typically a figure, a yes/no judgement, or a brief explanation, which is then graded against a human-written gold answer.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub