AA-Omniscience

AA-Omniscience evaluates factual knowledge and hallucination behavior across public-domain questions in several professional and academic domains.

Also known as: AA-Omniscience Public, AA-Omniscience Accuracy

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryknowledge
MetricAA-Omniscience Index
Directionhigher_is_better
Unitpoints
Dataset size6000
PublisherArtificial Analysis

What it measures

AA-Omniscience measures whether a language model answers knowledge questions correctly and whether it invents answers when it should abstain. Artificial Analysis reports results across business, humanities and social sciences, science and mathematics, health, law, and software engineering.

Task format

English open-answer questions with a correct-answer and hallucination analysis.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub