{
 "body": "## What it measures\n\nAA-Omniscience tests factual knowledge and the tendency to hallucinate. It uses English open-answer questions covering business, humanities and social sciences, science, engineering and mathematics, health, law, and software engineering. Example questions on the publisher\u2019s page ask for precise historical, medical, economic, and programming facts.\n\nThe evaluation distinguishes knowing from guessing. A model can refuse or say it does not know; the index does not penalize refusal, while an incorrect confident answer is treated as a hallucination.\n\n## How it is scored\n\nArtificial Analysis reports three related outputs. Accuracy is the proportion of correct answers across all questions. Hallucination rate is incorrect answers divided by all non-correct responses, including partial and not-attempted responses. The AA-Omniscience Index rewards correct answers, penalizes hallucinations, and ranges from -100 to 100; zero means correct and incorrect answers balance under the index definition.\n\nThe publisher\u2019s methodology lists 6,000 questions, one repeat, open-answer responses, and separate accuracy and one-minus-hallucination components in its Intelligence Index. Prompting, answer parsing, and refusal handling must match the current publisher protocol.\n\n## Dataset and licence\n\nThe official methodology states that AA-Omniscience has 6,000 questions. Artificial Analysis says it maintains internal copies of selected evaluation datasets, and the evaluation page presents domain distributions and example tasks. The consulted sources do not publish a dataset licence or answer release policy; those fields remain unknown. The page should therefore be treated as a publisher-run evaluation, not a freely downloadable test set.\n\n## Who publishes it\n\nArtificial Analysis develops and runs AA-Omniscience and reports it as a general evaluation in the Artificial Analysis Intelligence Index. The evaluation page is the publisher\u2019s current score and methodology surface. It reports model leaderboards for the index, accuracy, hallucination rate, and domain views.\n\n## Lineage\n\nAA-Omniscience is one evaluation within the Artificial Analysis Intelligence Index. The publisher\u2019s methodology places it in the General category and combines its accuracy and hallucination components in the composite index. No predecessor or successor benchmark is identified on the consulted official pages.\n\n## Saturation and contamination\n\nArtificial Analysis continues to display AA-Omniscience results for hundreds of models, and its score definition still separates accuracy from hallucination behavior. No ceiling or saturation claim is made. The publisher maintains internal copies of the evaluation data, which limits direct public inspection, but the consulted sources do not establish whether specific model training runs contained the questions. Contamination risk is therefore unknown.\n\n## How to run it\n\nAA-Omniscience is primarily a publisher-run evaluation. To compare with its numbers, use the current Artificial Analysis prompt and answer evaluation, zero-shot instruction prompting, the stated language and output settings, and the same refusal and partial-answer rules. The official methodology records temperature conventions and general pass@1 practice, but the benchmark\u2019s answer set is not published for local reruns.\n\n## Reading the numbers\n\nA high accuracy score means the model answered many questions correctly. A low hallucination rate means it avoided incorrect answers among non-correct responses. The combined index rewards useful knowledge while allowing abstention, so it should not be read as a pure recall score. Compare all three measures and inspect domain breakdowns when choosing a model for factual work.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "artificialanalysis_aa_omniscience_public",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [
   "AA-Omniscience Public",
   "AA-Omniscience Accuracy"
  ],
  "category": "knowledge",
  "contamination": {
   "note": "The publisher maintains internal copies of evaluation datasets, but the consulted methodology does not establish model-specific training exposure.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "English"
   ],
   "modalities": [
    "text"
   ],
   "public_test_set": false,
   "size": 6000,
   "size_note": "6,000 questions in the Artificial Analysis Intelligence Index methodology.",
   "url": "https://artificialanalysis.ai/evaluations/omniscience"
  },
  "freshness": {
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna, luna-stream-a-003 (Codex coordinated)",
   "reviewed": "",
   "reviewed_by": ""
  },
  "harness": {
   "other": "Artificial Analysis evaluation methodology; zero-shot instruction prompting with its published answer evaluation."
  },
  "id": "artificialanalysis_aa_omniscience_public",
  "last_updated": "2026-09",
  "leaderboard_url": "https://artificialanalysis.ai/evaluations/omniscience",
  "measures": "AA-Omniscience measures whether a language model answers knowledge questions correctly and whether it invents answers when it should abstain. Artificial Analysis reports results across business, humanities and social sciences, science and mathematics, health, law, and software engineering.",
  "metric": {
   "baseline_note": "The index ranges from -100 to 100; Artificial Analysis also reports accuracy and hallucination rate separately.",
   "direction": "higher_is_better",
   "max_score": 100,
   "name": "AA-Omniscience Index",
   "unit": "points"
  },
  "name": "AA-Omniscience",
  "page_kind": "benchmark",
  "paper": {},
  "publisher": {
   "org": "Artificial Analysis",
   "url": "https://artificialanalysis.ai/"
  },
  "released": "",
  "saturation": {
   "note": "The publisher continues to report results for new models; no ceiling is established.",
   "status": "open"
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "AA-Omniscience evaluation page",
    "url": "https://artificialanalysis.ai/evaluations/omniscience"
   },
   {
    "accessed": "2026-09-08",
    "title": "Artificial Analysis Intelligence Benchmarking Methodology",
    "url": "https://artificialanalysis.ai/methodology/intelligence-benchmarking"
   }
  ],
  "summary": "AA-Omniscience evaluates factual knowledge and hallucination behavior across public-domain questions in several professional and academic domains.",
  "tags": [
   "knowledge",
   "hallucination",
   "factuality",
   "private-dataset"
  ],
  "task_format": "English open-answer questions with a correct-answer and hallucination analysis."
 }
}