V-FiLLM

V-FiLLM builds financial table-reasoning questions from executable computation trees so answers are correct by construction and difficulty is controllable.

Also known as: Verified Financial LLM Reasoning Benchmark

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorycompositional financial table reasoning
Page statusactive
Metricaccuracy
Directionhigher_is_better
Unit%
PublisherETH Zürich and Aisot Technologies Ltd

What it measures

V-FiLLM tests whether a language model can retrieve values from a financial table and compose arithmetic over them. Questions are rendered from typed expression trees whose leaves are spreadsheet cells (revenue, costs, assets, and named concepts such as gross profit). Difficulty is set independently along computation depth, expression breadth, financial-concept complexity, and context size. Optional multi-turn dialogues expose intermediate nodes, and adversarial table or wording perturbations test robustness without changing the gold answer.

Task format

English question over a synthetic 10-Q-style or regularized spreadsheet; the model returns a numeric answer scored against the tree-evaluated gold value, optionally across multi-turn sub-questions.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub