ConvFinQACalc

HELM Enterprise wrap of ConvFinQA: given a table, gold text facts, and prior turns, emit the last-turn number and score float equality.

Also known as: conv_fin_qa_calc, HELM ConvFinQACalc

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryHELM Enterprise numeric last-turn answers on ConvFinQA tables
Page statusunknown
Metricfloat_equiv
Directionhigher_is_better
Dataset size12594
Dataset licenceMIT
PublisherStanford CRFM (HELM scenario); ConvFinQA data from UC Santa Barbara and J.P. Morgan

What it measures

conv_fin_qa_calc is Stanford HELM's calculation scenario built on ConvFinQA (Chen et al., EMNLP 2022). Each instance is one conversation turn: a markdown table from an earnings report, optional gold supporting-fact sentences, earlier questions with their gold numeric answers, and the last unanswered question. The model must output that last number. It does not ask for ConvFinQA's reasoning program. English text plus a table. HELM Enterprise run spec, not the HELM Finance fin_qa program-generation task.

Task format

Generation. Instructions: "Based on the table, answer the final question. Respond with the answer only, with no additional explanation." Output noun Answer. Default adapter max_tokens 5 and max_train_instances 5 (five-shot from TRAIN_SPLIT). Main split valid. Main metric float_equiv.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub