The FACTS Leaderboard

Google FACTS suite averages multimodal, parametric, search and grounding-v2 tracks into one FACTS Score on public and private splits.

Also known as: FACTS Leaderboard Suite, FACTS Score

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorycomposite
Subcategoryfour-track factuality suite: multimodal, parametric, search and grounding v2
Page statusactive
MetricFACTS Score (mean of four track accuracies)
Directionhigher_is_better
Unit%
Dataset size7229
PublisherGoogle DeepMind, Google Research, Google Cloud, and Kaggle

What it measures

The FACTS Leaderboard is a four-track factuality suite, not a single prompt set. FACTS Multimodal asks image questions that need visual grounding plus world knowledge. FACTS Parametric asks closed-book factoids that users care about and that Wikipedia supports. FACTS Search requires a shared Brave Search tool on tail and multi-hop questions. FACTS Grounding v2 reuses the v1 long-document prompts with newer judges. The headline FACTS Score is the unweighted mean of the four track accuracies, each already averaged over that track's public and private items.

Task format

Mixed: image-plus-text QA with a human rubric; short closed-book factoids; tool-using search with Brave Search API; long-form grounded generation from a supplied document. Kaggle runs official scoring.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub