AA-Omniscience evaluates factual knowledge and hallucination behavior across public-domain questions in several professional and academic domains.
unassessed
| Category | knowledge |
|---|---|
| Metric | AA-Omniscience Index |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 6000 |
| Publisher | Artificial Analysis |
AA-Omniscience measures whether a language model answers knowledge questions correctly and whether it invents answers when it should abstain. Artificial Analysis reports results across business, humanities and social sciences, science and mathematics, health, law, and software engineering.
English open-answer questions with a correct-answer and hallucination analysis.
No model card in ModelSpec reports this benchmark yet.