Meta's 1,645-prompt instruction-following benchmark, expert-written and LLM-judged against per-prompt rubrics, covering complex single-turn instructions, multi-turn carried context and system-prompt steerability.
unassessed
| Category | instruction-following |
|---|---|
| Subcategory | expert-written, rubric-graded instruction following: complex single-turn, multi-turn carried context, and system-prompt steerability |
| Page status | active |
| Metric | rubric satisfaction rate (share of responses that satisfy every item in their prompt's rubric), reported per subset and averaged across the three |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1645 |
| Dataset licence | CC BY-NC 4.0 |
| Publisher | Meta Superintelligence Labs (Meta), with contributions from Princeton University and Carnegie Mellon University |
AdvancedIF tests instruction following across three capabilities that IFEval-style mechanically-checked benchmarks were not built to cover: complex single-turn instructions (each prompt combines six or more simultaneous constraints -- tone, format, style, structure, length, negative constraints, spelling and instructions that depend on one another), multi-turn carried context (whether a model keeps honouring an instruction given earlier in a conversation once the conversation has moved past it), and system-prompt steerability (whether a model follows instructions placed in the system prompt rather than the user turn, including persona and scope restrictions). Every prompt is expert-written rather than built from a template, and is paired with an expert-curated rubric -- a checklist of specific yes/no questions a correct response must satisfy, many of which (tone, persona consistency, whether an explanation "feels" grounded) cannot be checked by a program the way IFEval's or IFBench's constraints can, which is why AdvancedIF grades with an LLM judge instead of verification code.
Free-text response to a single-turn or multi-turn conversation (including, for the system-steerability subset, a system prompt the response must respect). An LLM judge reads the full conversation, the response and the prompt's rubric, answers each rubric question, and outputs whether the response satisfied every rubric item.
No model card in ModelSpec reports this benchmark yet.