OpenCompass revision of NeedleBench: equal-weight retrieval scores, fictional multi-needle facts, and power-of-two Ancestral Trace Challenge depths.
unassessed
| Category | long-context |
|---|---|
| Subcategory | revised bilingual needle retrieval, fictional multi-needle reasoning, power-of-two ATC |
| Page status | active |
| Metric | equal-weight mean of S-RT, M-RT and M-RS (OpenCompass overall); ATC exact match with weighted average and ENL-50 in the paper |
| Direction | higher_is_better |
| Unit | % |
| Dataset licence | MIT |
| Publisher | Shanghai AI Laboratory and Tsinghua University |
NeedleBench V2 keeps the bilingual long-context layout of NeedleBench and changes the parts that leaked knowledge or unbalanced the score. Sparse tasks still plant needles in Paul Graham essays or Chinese domain text at a chosen length and depth: single-needle retrieval, multi-needle retrieval, and multi-needle reasoning. Multi-needle reasoning no longer uses R4C/MultiHop-style needles; it uses fictional kinship facts of the same kind as ATC, so a model cannot answer from pretraining. ATC remains information-dense (no filler) but spaces needle counts on powers of two (2, 4, …, 512) instead of 1–5. English and Chinese. Text only. Length packs at 4k, 8k, 32k, 128k, 200k, 256k and 1000k.
Long generated prompt in; short free-text answer out. Sparse tasks: keyword-aware exact match, 10 repeats, GPT-4 tokenizer length. ATC: NeedleBenchATCDataset with NeedleBenchATCEvaluator and needlebench_atc_postprocess_v2; OpenCompass English 0-shot config uses needle counts [2, 4, 8, 16, 32, 64, 128, 256, 512], path opencompass/needlebench, names.json, 10 repeats.
No model card in ModelSpec reports this benchmark yet.