NeedleBench V2

OpenCompass revision of NeedleBench: equal-weight retrieval scores, fictional multi-needle facts, and power-of-two Ancestral Trace Challenge depths.

Also known as: needlebench_v2, NeedleBench v2

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorylong-context
Subcategoryrevised bilingual needle retrieval, fictional multi-needle reasoning, power-of-two ATC
Page statusactive
Metricequal-weight mean of S-RT, M-RT and M-RS (OpenCompass overall); ATC exact match with weighted average and ENL-50 in the paper
Directionhigher_is_better
Unit%
Dataset licenceMIT
PublisherShanghai AI Laboratory and Tsinghua University

What it measures

NeedleBench V2 keeps the bilingual long-context layout of NeedleBench and changes the parts that leaked knowledge or unbalanced the score. Sparse tasks still plant needles in Paul Graham essays or Chinese domain text at a chosen length and depth: single-needle retrieval, multi-needle retrieval, and multi-needle reasoning. Multi-needle reasoning no longer uses R4C/MultiHop-style needles; it uses fictional kinship facts of the same kind as ATC, so a model cannot answer from pretraining. ATC remains information-dense (no filler) but spaces needle counts on powers of two (2, 4, …, 512) instead of 1–5. English and Chinese. Text only. Length packs at 4k, 8k, 32k, 128k, 200k, 256k and 1000k.

Task format

Long generated prompt in; short free-text answer out. Sparse tasks: keyword-aware exact match, 10 repeats, GPT-4 tokenizer length. ATC: NeedleBenchATCDataset with NeedleBenchATCEvaluator and needlebench_atc_postprocess_v2; OpenCompass English 0-shot config uses needle counts [2, 4, 8, 16, 32, 64, 128, 256, 512], path opencompass/needlebench, names.json, 10 repeats.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub