IFEval

IFEval scores whether a model's response obeys objectively checkable instructions, such as word counts or keyword frequency, using code rather than human or LLM judgment.

Also known as: Instruction-Following Eval, Instruction-Following Evaluation for Large Language Models

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryinstruction-following
Subcategoryobjectively verifiable natural-language instruction compliance
Page statusactive
Metricinstruction-following accuracy (strict and loose, at the prompt level and the instruction level)
Directionhigher_is_better
Unit%
Dataset size500
Dataset licenceCC BY 4.0
PublisherGoogle Research

What it measures

IFEval tests whether a model follows explicit, machine-checkable instructions layered onto a prompt, such as "write more than 400 words," "mention the keyword 'AI' at least 3 times," or "wrap your answer in double quotation marks." Each of the roughly 500 prompts combines one or more of 25 instruction types covering things like length, keyword use, formatting, punctuation and casing. Because every instruction can be checked by a program rather than a human or another model, scoring needs no subjective judgment call, unlike most instruction-following or helpfulness evaluations.

Task format

A model is given a prompt containing one or more verifiable instructions and generates a free-form response in a single turn, zero-shot, with no few-shot examples. A separate verification function checks the response against each instruction mechanically (for example, counting words or scanning for a keyword) and returns pass/fail per instruction.

Models reporting this benchmark

These figures come from the model cards, which carry one collection date per card and no per-score attribution. They are shown as reported, not as verified evidence.
ModelProviderScoreCard as of
GPT-5.1OpenAI92.52026-04
GPT-5.1 ChatOpenAI92.52026-04
GPT-5.1 CodexOpenAI92.52026-04
GPT-5.1 Codex MaxOpenAI92.52026-04
GPT-5.1 Codex miniOpenAI92.52026-04
Claude Mythos PreviewAnthropic92.12026-04
Claude Opus 4Anthropic92.12026-04
Claude Opus 4.6Anthropic92.12026-04
o3-proOpenAI922026-04
GPT-5OpenAI91.82026-04
GPT-5 Chat (latest)OpenAI91.82026-04
GPT-5 ProOpenAI91.82026-04
GPT-5-CodexOpenAI91.82026-04
GPT-5.2OpenAI91.82026-04
GPT-5.2 ChatOpenAI91.82026-04
GPT-5.2 CodexOpenAI91.82026-04
GPT-5.2 ProOpenAI91.82026-04
GPT-5.3 Chat (latest)OpenAI91.82026-04
GPT-5.3 CodexOpenAI91.82026-04
GPT-5.3 Codex SparkOpenAI91.82026-04
GPT-5.4OpenAI91.82026-04
GPT-5.4 miniOpenAI91.82026-04
GPT-5.4 nanoOpenAI91.82026-04
GPT-5.4 ProOpenAI91.82026-04
Claude Opus 4.5Anthropic91.52026-04
Claude Opus 4.5 (latest)Anthropic91.52026-04
Claude Opus 4.1Anthropic91.02026-04
Claude Opus 4.1 (latest)Anthropic91.02026-04
Gemini 2.5 ProGoogle DeepMind91.02026-04
Gemini 2.5 Pro Preview 05-06Google DeepMind91.02026-04
Gemini 2.5 Pro Preview 06-05Google DeepMind91.02026-04
Gemini 2.5 Pro Preview TTSGoogle DeepMind91.02026-04
o3OpenAI91.02026-04
o3-deep-researchOpenAI912026-04
Claude Sonnet 4Anthropic90.42026-04
Claude Sonnet 4.6Anthropic90.42026-04
Grok 4xAI90.02026-04
Grok 4 FastxAI90.02026-04
Grok 4 Fast (Non-Reasoning)xAI90.02026-04
Grok 4.1 FastxAI90.02026-04
Grok 4.1 Fast (Non-Reasoning)xAI90.02026-04
Grok 4.20 (Non-Reasoning)xAI90.02026-04
Grok 4.20 (Reasoning)xAI90.02026-04
Grok 4.20 Multi-AgentxAI90.02026-04
Claude Opus 4 (latest)Anthropic89.72026-04
GPT-4.1OpenAI89.52026-04
Claude Sonnet 4.5Anthropic89.32026-04
Claude Sonnet 4.5 (latest)Anthropic89.32026-04
o4-miniOpenAI89.02026-04
o4-mini-deep-researchOpenAI89.02026-04
o1-proOpenAI88.52026-04
Qwen3 235B-A22BAlibaba / Qwen Team88.52026-04
DeepSeek R1 0528DeepSeek88.02026-04
DeepSeek R1 0528 NVFP4 v2NVIDIA88.02026-04
GPT-5 MiniOpenAI88.02026-04
Qwen3-Coder 480B-A35B InstructAlibaba / Qwen Team88.02026-04
Gemini 2.5 FlashGoogle DeepMind87.52026-04
Gemini 2.5 Flash ImageGoogle DeepMind87.52026-04
Gemini 2.5 Flash Image (Preview)Google DeepMind87.52026-04
Gemini 2.5 Flash LiteGoogle DeepMind87.52026-04
Gemini 2.5 Flash Lite Preview 06-17Google DeepMind87.52026-04
Gemini 2.5 Flash Lite Preview 09-25Google DeepMind87.52026-04
Gemini 2.5 Flash Preview 04-17Google DeepMind87.52026-04
Gemini 2.5 Flash Preview 05-20Google DeepMind87.52026-04
Gemini 2.5 Flash Preview 09-25Google DeepMind87.52026-04
Gemini 2.5 Flash Preview TTSGoogle DeepMind87.52026-04
Claude Sonnet 4 (latest)Anthropic87.22026-04
Claude Sonnet 3.7Anthropic87.02026-04
DeepSeek V3.2DeepSeek87.02026-04
DeepSeek V3.2 ExpDeepSeek87.02026-04
DeepSeek R1DeepSeek86.52026-04
DeepSeek ReasonerDeepSeek86.52026-04
o3-miniOpenAI86.52026-04
o1OpenAI86.02026-04
o1-previewOpenAI86.02026-04
DeepSeek V3.1DeepSeek85.52026-04
Claude Sonnet 3.5Anthropic85.42026-04
Claude Sonnet 3.5 v2Anthropic85.42026-04
Claude Haiku 4.5Anthropic85.02026-04
Claude Haiku 4.5 (latest)Anthropic85.02026-04
GPT-4.1 miniOpenAI85.02026-04
Grok 3xAI85.02026-04
Grok 3 FastxAI85.02026-04
Grok 3 Fast LatestxAI85.02026-04
Grok 3 LatestxAI85.02026-04
Magistral Medium (latest)Mistral AI85.02026-04
GPT-4oOpenAI84.32026-04
GPT-4o (2024-05-13)OpenAI84.32026-04
GPT-4o (2024-08-06)OpenAI84.32026-04
GPT-4o (2024-11-20)OpenAI84.32026-04
DeepSeek ChatDeepSeek84.02026-04
DeepSeek V2DeepSeek842026-04
DeepSeek V2 LiteDeepSeek842026-04
DeepSeek V2 Lite ChatDeepSeek842026-04
DeepSeek V3DeepSeek84.02026-04
DeepSeek V3 0324DeepSeek84.02026-04
Gemini 2.0 FlashGoogle DeepMind84.02026-04
Qwen3 32BAlibaba / Qwen Team84.02026-04
Qwen3 32B AWQAlibaba / Qwen Team84.02026-04
Qwen3 32B NVFP4NVIDIA84.02026-04
Gemma 4 31BGoogle DeepMind83.02026-04
gemma 4 31B itGoogle DeepMind83.02026-04
gemma 4 31B it GGUFUnsloth83.02026-04
Gemma 4 31B IT NVFP4NVIDIA83.02026-04
Llama 4 Maverick 17B 128E InstructMeta83.02026-04
Llama-4-Maverick-17B-128E-Instruct-FP8Meta83.02026-04
GPT-5 NanoOpenAI82.52026-04
Gemini 1.5 ProGoogle DeepMind82.02026-04
GPT-4 TurboOpenAI82.02026-04
Mistral Large (latest)Mistral AI82.02026-04
Mistral Large 2.1Mistral AI82.02026-04
Mistral Large 3Mistral AI82.02026-04
o1-miniOpenAI82.02026-04
Qwen2.5 72B InstructAlibaba / Qwen Team82.02026-04
Gemma 4 26BGoogle DeepMind81.52026-04
gemma 4 26B A4B itGoogle DeepMind81.52026-04
gemma 4 26B A4B it GGUFUnsloth81.52026-04
Claude Opus 3Anthropic81.02026-04
Meta Llama 3 70B InstructMeta81.02026-04
Meta Llama 3 70B InstructNous Research81.02026-04
Claude Haiku 3.5Anthropic80.52026-04
Claude Haiku 3.5 (latest)Anthropic80.52026-04
GPT-4o miniOpenAI80.52026-04
Grok 3 MinixAI80.52026-04
Grok 3 Mini FastxAI80.52026-04
Grok 3 Mini Fast LatestxAI80.52026-04
Grok 3 Mini LatestxAI80.52026-04
Llama 3.1 405BMeta80.02026-04
Llama 3.1 405B FP8Meta80.02026-04
Llama 3.1 405B InstructMeta80.02026-04
Llama 3.1 405B Instruct FP8Meta80.02026-04
Qwen3 30B A3B Instruct 2507Alibaba / Qwen Team80.02026-04
Qwen3 30B A3B NVFP4NVIDIA80.02026-04
Qwen3 30B-A3BAlibaba / Qwen Team80.02026-04
QwQ PlusAlibaba / Qwen Team802026-04
Gemma 3 27BGoogle DeepMind79.02026-04
Llama 4 Scout 17B 16EMeta79.02026-04
Llama 4 Scout 17B 16E InstructMeta79.02026-04
Llama-4-Scout-17B-16E-Instruct-FP8Meta79.02026-04
Gemini 2.0 Flash LiteGoogle DeepMind78.52026-04
Llama 3.3 70B Instruct NVFP4NVIDIA78.52026-04
Llama-3.3-70B-InstructMeta78.52026-04
Command ACohere782026-04
Command A ReasoningCohere782026-04
DeepSeek R1 Distill Llama 70BDeepSeek78.02026-04
GPT-4.1 nanoOpenAI78.02026-04
Grok 2xAI782026-04
Grok 2 (1212)xAI782026-04
Grok 2 LatestxAI782026-04
Grok 2 VisionxAI782026-04
Grok 2 Vision (1212)xAI782026-04
Grok 2 Vision LatestxAI782026-04
Magistral SmallMistral AI78.02026-04
Magistral Small 2506Mistral AI78.02026-04
phi 4Microsoft78.02026-04
Phi 4 multimodal instructMicrosoft782026-04
Qwen2.5 32B InstructAlibaba / Qwen Team782026-04
Qwen2.5 32B Instruct AWQAlibaba / Qwen Team782026-04
Mixtral 8x22BMistral AI77.82026-04
Mixtral 8x22B Instruct v0.1Mistral AI77.82026-04
Yi 1.5 34B01.AI77.32026-04
Yi 1.5 34B Chat01.AI77.32026-04
Gemini 1.5 FlashGoogle DeepMind77.02026-04
Gemini 1.5 Flash-8BGoogle DeepMind77.02026-04
Qwen3 14BAlibaba / Qwen Team772026-04
Qwen3 14B AWQAlibaba / Qwen Team772026-04
Qwen3 14B NVFP4NVIDIA772026-04
GPT-4OpenAI76.52026-04
Falcon3 7B InstructTII76.12025-03
DeepSeek R1 Distill Qwen 32BDeepSeek76.02026-04
Llama 3.1 70BMeta76.02026-04
Llama 3.1 70B InstructMeta76.02026-04
Qwen2.5 Coder 32B InstructAlibaba / Qwen Team76.02026-04
Qwen2.5 Coder 32B Instruct AWQAlibaba / Qwen Team76.02026-04
Claude Sonnet 3Anthropic75.52026-04
Mistral NemoMistral AI75.12026-04
Mistral Nemo Base 2407Mistral AI75.12026-04
Mistral Nemo Instruct 2407Mistral AI75.12026-04
Command R+Cohere75.02026-04
gemma 2 27B itGoogle DeepMind75.02026-04
Codestral (latest)Mistral AI74.02026-04
Qwen2.5 14B InstructAlibaba / Qwen Team742026-04
Qwen2.5 14B Instruct AWQAlibaba / Qwen Team742026-04
Llama 3.2 3B InstructMeta73.92026-04
Gemma 3 12BGoogle DeepMind73.02026-04
Qwen3 8BAlibaba / Qwen Team732026-04
Qwen3 8B AWQAlibaba / Qwen Team732026-04
Qwen3 8B BaseAlibaba / Qwen Team732026-04
OLMo 2 1124 7BAllen AI72.42025-03
Phi 3.5 mini instructMicrosoft72.42026-04
granite 3.1 8B instructIBM72.12026-04
Claude Haiku 3Anthropic72.02026-04
DeepSeek R1 Distill Qwen 14BDeepSeek72.02026-04
Phi 4 mini instructMicrosoft72.02026-04
Falcon3 Mamba 7B InstructTII71.72025-03
Command RCohere702026-04
Falcon3 3B InstructTII69.82025-03
Mixtral 8x7BMistral AI69.52026-04
Mixtral 8x7B Instruct v0.1Mistral AI69.52026-04
Mixtral 8x7B v0.1Mistral AI69.52026-04

Showing the top 200 of 354.

Data

This page as JSON · Edit on GitHub