Google FACTS suite averages multimodal, parametric, search and grounding-v2 tracks into one FACTS Score on public and private splits.
unassessed
| Category | composite |
|---|---|
| Subcategory | four-track factuality suite: multimodal, parametric, search and grounding v2 |
| Page status | active |
| Metric | FACTS Score (mean of four track accuracies) |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 7229 |
| Publisher | Google DeepMind, Google Research, Google Cloud, and Kaggle |
The FACTS Leaderboard is a four-track factuality suite, not a single prompt set. FACTS Multimodal asks image questions that need visual grounding plus world knowledge. FACTS Parametric asks closed-book factoids that users care about and that Wikipedia supports. FACTS Search requires a shared Brave Search tool on tail and multi-hop questions. FACTS Grounding v2 reuses the v1 long-document prompts with newer judges. The headline FACTS Score is the unweighted mean of the four track accuracies, each already averaged over that track's public and private items.
Mixed: image-plus-text QA with a human rubric; short closed-book factoids; tool-using search with Brave Search API; long-form grounded generation from a supplied document. Kaggle runs official scoring.
No model card in ModelSpec reports this benchmark yet.