{
 "body": "\n## What it measures\n\nNorEval is a collection of Norwegian-language tasks in lm-evaluation-harness. The repository tree includes grammar correction, Norwegian commonsense, Norwegian Belebele, idiom, OpenBookQA, reading comprehension, summarization, and generation directories. The README describes 24 datasets across nine categories and supports both Bokm\u00e5l and Nynorsk.\n\nBecause the members test different abilities, NorEval does not define one single prompt or one universal metric. A run must identify its member tasks and language variant.\n\n## How it is scored\n\nMetrics are task-specific. The sources read identify the collection and its task directories, but do not establish an aggregate scoring formula, human baseline, or common maximum. Report each task\u2019s metric separately.\n\n## Dataset and licence\n\nThe suite README does not provide an aggregate item count or common licence. Member datasets may have different sizes and terms. Norwegian is the documented language context, but exact dialect and split details should be taken from each member task.\n\n## Who publishes it\n\nThe task integration is distributed in EleutherAI\u2019s lm-evaluation-harness. The README identifies the paper \u201cNorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark\u201d (arXiv:2504.07749) and the NorEval project at `ltgoslo/noreval`. No current aggregate leaderboard was established.\n\n## Lineage\n\nNorEval is a family page. The repository exposes member task directories including `ask_gec`, `ncb`, `norbelebele`, `norcommonsenseqa`, `norec`, `noridiom`, `noropenbookqa`, `norquad`, `norsumm`, `nortruthfulqa`, `nrk_quiz_qa`, `norrewrite-instruct`, `norsummarize-instruct`, and `tatoeba`. These are member evaluations, not interchangeable aliases.\n\n## Saturation and contamination\n\nAggregate saturation and contamination are unknown. Different members have different publication histories and exposure profiles, so a suite-level label would conceal important differences.\n\n## How to run it\n\nUse the lm-evaluation-harness task name `noreval` only after checking the current task registry and selecting the desired members. Record the exact subtask, dataset revision, prompt, and metric. Scores from Norwegian subdatasets should not be averaged without a documented weighting rule.\n\n## Reading the numbers\n\nNorEval results can indicate Norwegian performance across several task types. They do not imply uniform competence across the suite or broad Norwegian cultural and linguistic coverage. Inspect member-level scores, dialect, prompt language, and dataset provenance before drawing conclusions.\n\nAn aggregate number can hide a model\u2019s strengths and weaknesses because the member tasks differ in objective, format, and difficulty. A reproducible report should list every selected task and preserve the harness version used to load it.\n",
 "build": {
  "built_at": "2026-09-09T16:56:50+00:00",
  "commit": "0a599558854c0e238c03a0f0d725239cb28f9d11",
  "eligibility_as_of": "2026-09-09"
 },
 "disposition": {
  "canonical_id": "noreval",
  "reasons": [],
  "status": "unassessed",
  "verified_results": []
 },
 "models_covered": [],
 "page": {
  "aliases": [],
  "category": "composite",
  "contamination": {
   "note": "The suite mixes datasets with different publication histories; aggregate risk is not established.",
   "risk": "unknown"
  },
  "dataset": {
   "languages": [
    "Norwegian"
   ],
   "license": "",
   "modalities": [
    "text"
   ],
   "public_test_set": null,
   "size": 24,
   "size_note": "The NorEval README describes 24 datasets: 19 existing peer-reviewed datasets and five created for the benchmark; member item counts vary.",
   "splits": "",
   "url": "https://github.com/ltgoslo/noreval"
  },
  "freshness": {
   "luna-batch-017": null,
   "luna-batch-017 (Codex coordinated)": null,
   "researched": "2026-09-08",
   "researched_by": "GPT-5.6 Luna",
   "reviewed": "2026-09-08",
   "reviewed_by": "GPT-5.6 Luna independent review"
  },
  "harness": {
   "bigbench": "",
   "helm": "",
   "inspect_evals": "",
   "lm_eval": "noreval",
   "opencompass": "",
   "other": ""
  },
  "id": "noreval",
  "last_updated": "",
  "leaderboard_url": "",
  "lineage": {
   "family": "",
   "predecessor": "",
   "successors": [],
   "variants": [
    "ask_gec",
    "ncb",
    "norbelebele",
    "norcommonsenseqa",
    "norec",
    "noridiom",
    "noropenbookqa",
    "norquad"
   ]
  },
  "measures": "NorEval groups Norwegian-language tasks across nine categories, including grammar correction, commonsense, Belebele, idioms, OpenBookQA, reading comprehension, summarization, and generation. It is a suite rather than one item-level evaluation.",
  "metric": {
   "baseline_note": "Metrics are task-specific.",
   "direction": "higher_is_better",
   "human_baseline": null,
   "max_score": null,
   "name": "",
   "random_baseline": null,
   "unit": ""
  },
  "name": "NorEval",
  "page_kind": "family",
  "paper": {
   "arxiv": "2504.07749",
   "title": "NorEval: A Norwegian Language Understanding and Generation Evaluation Benchmark",
   "url": "https://arxiv.org/abs/2504.07749",
   "year": 2025
  },
  "publisher": {
   "authors": [],
   "org": "NorEval / EleutherAI lm-evaluation-harness integration",
   "url": "https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/noreval"
  },
  "released": "",
  "repo_url": "https://github.com/EleutherAI/lm-evaluation-harness",
  "saturation": {
   "as_of": "",
   "note": "No aggregate leaderboard or ceiling analysis was established.",
   "status": "unknown",
   "top_score": null
  },
  "sources": [
   {
    "accessed": "2026-09-08",
    "title": "lm-evaluation-harness NorEval task collection",
    "url": "https://github.com/EleutherAI/lm-evaluation-harness/tree/main/lm_eval/tasks/noreval"
   },
   {
    "accessed": "2026-09-08",
    "title": "NorEval README",
    "url": "https://raw.githubusercontent.com/EleutherAI/lm-evaluation-harness/main/lm_eval/tasks/noreval/README.md"
   }
  ],
  "status": "active",
  "subcategory": "Norwegian language evaluation",
  "summary": "NorEval is a 24-dataset Norwegian language evaluation benchmark integrated into lm-evaluation-harness.",
  "tags": [
   "norwegian",
   "multilingual",
   "suite"
  ],
  "task_format": "Varies by member task; text classification, question answering, and generation tasks are represented."
 }
}