1106 pages. 7 hold active eligibility as of 2026-09-09.
Verified against the catalogue contract as of 2026-09-09. These are the only benchmarks presented as current.
| Benchmark | ID | Category | Models reporting |
|---|---|---|---|
| AA-Briefcase | aa_briefcase | agentic | 0 |
| AA-LCR (Artificial Analysis Long Context Reasoning) | aa_lcr | long-context | 0 |
| AutomationBench-AA | automationbench_aa | agentic | 0 |
| CritPt (Complex Research using Integrated Thinking — Physics Test) | critpt | reasoning | 0 |
| GDP.pdf-AA | gdp_pdf_aa | long-context | 0 |
| GDPval-AA | gdpval_aa | agentic | 0 |
| SciCode | scicode | coding | 0 |
Assessed, but missing evidence the contract requires, or not yet approved by a reviewer. Not a claim of staleness.
| Benchmark | ID | Category | Models reporting |
|---|---|---|---|
| AutomationBench | automationbench | agentic | 0 |
| LiveBench | livebench | composite | 0 |
| OpenBookQA | openbookqa | reasoning | 0 |
| OpenML Benchmark | openml_benchmark | knowledge | 0 |
| OpenML Benchmarks | openml_benchmarks | knowledge | 0 |
| SWE-Together | swe_together | knowledge | 0 |
| Terminal-Bench v4.0 | terminal_bench_v4_0 | agentic | 0 |
Alternate identifiers that resolve to a canonical benchmark.
| Benchmark | ID | Category | Models reporting |
|---|---|---|---|
| OpenBookQA | obqa | knowledge | 0 |
Pages nobody has yet assessed against the contract. Reported separately from evaluated dispositions, as the spec requires.
| Benchmark | ID | Category | Models reporting |
|---|---|---|---|
| AA-Omniscience | artificialanalysis_aa_omniscience_public | knowledge | 0 |
| AbstentionBench | abstention_bench | safety | 0 |
| Abstract Narrative Understanding (BIG-bench) | abstract_narrative_understanding | reasoning | 0 |
| Abstraction and Reasoning Corpus (BIG-bench wrapping) | abstraction_and_reasoning_corpus | reasoning | 0 |
| ACI-Bench | aci_bench | domain | 0 |
| ACLUE (Ancient Chinese Language Understanding Evaluation) | aclue | knowledge | 0 |
| ACPBench | acp_bench | reasoning | 0 |
| ACPBench-Hard | acp_bench_hard | reasoning | 0 |
| ADHD-Behavior | shc_ptbm | domain | 0 |
| ADHD-MedEffects | shc_sei | domain | 0 |
| AdvancedIF | advancedif | instruction-following | 0 |
| AdvGLUE (Adversarial GLUE) | adv_glue | composite | 0 |
| Agent Memory Benchmark (AMB) | agent_memory_benchmark_amb | knowledge | 0 |
| Agent Memory Leaderboard | agent_memory_leaderboard | knowledge | 0 |
| agent-memory-bench | agent_memory_bench | knowledge | 0 |
| AgentBench | agent_bench | agentic | 0 |
| AgentDojo | agentdojo | agentic | 0 |
| AgentHarm | agentharm | safety | 0 |
| Agentic misalignment | agentic_misalignment | safety | 0 |
| AgentThreatBench | agent_threat_bench | safety | 0 |
| AGIEval | agieval | composite | 0 |
| AI Agent Benchmark Leaderboards | ai_agent_benchmark_leaderboards | knowledge | 0 |
| AI Agent Benchmark Results | ai_agent_benchmark_results | knowledge | 0 |
| ai-agent-benchmark | ai_agent_benchmark | knowledge | 0 |
| AI2D (AI2 Diagrams) | ai2d | multimodal | 39 |
| Aider Polyglot Benchmark | aider_polyglot | coding | 116 |
| AIME (American Invitational Mathematics Examination) | aime | math | 0 |
| AIME 2024 | aime_2024 | math | 0 |
| AIME 2025 | aime_2025 | math | 29 |
| AIME 2026 | aime_2026 | math | 0 |
| AIR-Bench 2024 | air_bench | safety | 0 |
| AIR-Bench 2024 | air_bench_2024 | knowledge | 0 |
| AIR-BENCH Live | air_bench_live | knowledge | 0 |
| AIR-Bench-Dataset | air_bench_dataset | knowledge | 0 |
| AlGhafa | alghafa | composite | 0 |
| Alignment of Simplicity Priors for Turing-Complete Concept Learning | simp_turing_concept | reasoning | 0 |
| AlpacaEval | alpaca_eval | human-preference | 56 |
| ALRAGE | alrage | knowledge | 0 |
| Anachronisms | anachronisms | knowledge | 0 |
| Analogical Similarity | analogical_similarity | reasoning | 0 |
| Analytic Entailment | analytic_entailment | reasoning | 0 |
| ANIMA (Animal Norms In Moral Assessment) | anima | safety | 0 |
| ANLI (Adversarial NLI) | anli | reasoning | 0 |
| Anthropic HH-RLHF (HELM Instruct) | anthropic_hh_rlhf | instruction-following | 0 |
| Anthropic Red Team (HELM) | anthropic_red_team | safety | 0 |
| APE (Attempt to Persuade Eval) | ape | safety | 0 |
| APPS (Automated Programming Progress Standard) | apps | coding | 0 |
| ArabCulture | arabculture | reasoning | 0 |
| Arabic Content Generation (HELM Arabic Enterprise) | arabic_content_generation | generation | 0 |
| Arabic EXAMS | arabic_exams | knowledge | 0 |
| Arabic Finance (HELM Arabic Enterprise) | arabic_finance | domain | 0 |
| Arabic Legal (HELM Arabic Enterprise) | arabic_legal | domain | 0 |
| ArabicMMLU | arabic_mmlu | knowledge | 0 |
| AraDiCE | aradice | composite | 0 |
| AraTrust | aratrust | safety | 0 |
| ARC (AI2 Reasoning Challenge) | arc | reasoning | 0 |
| ARC Prize Public Evaluation | arc_prize_public_evaluation | reasoning | 0 |
| ARC-AGI-2 | arc_agi_2 | reasoning | 1 |
| ARC-Challenge | arc_challenge | reasoning | 50 |
| ARC-Easy | arc_easy | reasoning | 0 |
| Arena Elo (Chatbot Arena / LMArena) | arena_elo | human-preference | 0 |
| Arena Elo — Coding | arena_elo_coding | human-preference | 210 |
| Arena Elo — Hard Prompts | arena_elo_hard_prompts | human-preference | 103 |
| Arena Elo — Math | arena_elo_math | human-preference | 209 |
| Arena Elo — Overall (Text) | arena_elo_overall | human-preference | 211 |
| Arena Elo — Style Control | arena_elo_style_control | human-preference | 103 |
| Arena Elo — Vision | arena_elo_vision | human-preference | 55 |
| Arithmetic (GPT-3 synthetic arithmetic tasks) | arithmetic | math | 0 |
| Artificial Analysis | artificial_analysis | composite | 0 |
| Artificial Analysis Intelligence Index | artificial_analysis_quality_index | composite | 130 |
| Artificial Analysis Output Speed | artificial_analysis_speed_index | composite | 129 |
| ArtificialAnalysis/AA-Briefcase-Lite | artificialanalysis_aa_briefcase_lite | knowledge | 0 |
| ArtificialAnalysis/AA-LCR | artificialanalysis_aa_lcr | knowledge | 0 |
| ArxivRollBench | arxivrollbench | reasoning | 0 |
| ASCII MNIST | mnist_ascii | multimodal | 0 |
| ASCII Word Recognition (BIG-bench) | ascii_word_recognition | reasoning | 0 |
| ASDiv | asdiv | math | 0 |
| AssistantBench | assistant_bench | agentic | 0 |
| ASTRA-bench | astra_bench | agentic | 0 |
| ASTRA-QA | astra_qa | knowledge | 0 |
| ATLAS (AGI-Oriented Testbed for Logical Application in Science) | atlas | reasoning | 0 |
| Authorship Verification (BIG-bench) | authorship_verification | reasoning | 0 |
| Auto Categorization | auto_categorization | knowledge | 0 |
| Auto Debugging | auto_debugging | coding | 0 |
| AutoBencher Capabilities | autobencher_capabilities | composite | 0 |
| AutoBencher Safety | autobencher_safety | safety | 0 |
| bAbI (Question-Answering Tasks) | babi_qa | reasoning | 0 |
| BABILong | babilong | long-context | 0 |
| Bangla BoolQA | bangla_boolqa | reasoning | 0 |
| Bangla CommonsenseQA | bangla_commonsenseqa | reasoning | 0 |
| Bangla MMLU | bangla_mmlu | knowledge | 0 |
| Bangla OpenBookQA | bangla_openbookqa | reasoning | 0 |
| Bangla PIQA | bangla_piqa | reasoning | 0 |
| BANKING77 | banking77 | domain | 0 |
| BasqueBench | basque_bench | composite | 0 |
| BasqueGLUE | basqueglue | composite | 0 |
| BBQ (Bias Benchmark for QA) | bbq | safety | 66 |
| BBQ-Lite (Bias Benchmark for QA, BIG-bench) | bbq_lite | safety | 0 |
| BBQ-Lite JSON (BIG-bench) | bbq_lite_json | safety | 0 |
| BEIR | beir | embedding | 69 |
| Belebele | belebele | knowledge | 0 |
| Bench-CoE | bench_coe | composite | 0 |
| Bench-MFG | bench_mfg | reasoning | 0 |
| Benchmark Contamination | benchmark_contamination | knowledge | 0 |
| benchmark-bcplus | benchmark_bcplus | knowledge | 0 |
| benchmark-research | benchmark_research | knowledge | 0 |
| benchmark_results | benchmark_results | knowledge | 0 |
| Benchmarking the Benchmarks | benchmarking_the_benchmarks | reasoning | 0 |
| Benchmarking the Domain Gap | benchmarking_the_domain_gap | domain | 0 |
| Benchmarking the Residual | benchmarking_the_residual | long-context | 0 |
| BertaQA | bertaqa | knowledge | 0 |
| Beyond BLEU | beyond_bleu | knowledge | 0 |
| Beyond FLOPs | beyond_flops | knowledge | 0 |
| Beyond Leaderboards | beyond_leaderboards | knowledge | 0 |
| Beyond MSE | beyond_mse | knowledge | 0 |
| BeyondAIME | beyondaime | math | 0 |
| BFCL (Berkeley Function-Calling Leaderboard) | bfcl | agentic | 0 |
| BHS (Basque, Hindi, Swahili syntactic evaluation) | bhs | reasoning | 0 |
| BIG-bench (Beyond the Imitation Game Benchmark) | big_bench | composite | 0 |
| BIG-Bench Extra Hard (BBEH) | bbeh | reasoning | 0 |
| BIG-Bench Hard | bbh | reasoning | 221 |
| BigCodeBench | bigcodebench | coding | 0 |
| BIRD | bird | knowledge | 0 |
| BIRD-CRITIC | bird_critic | knowledge | 0 |
| BIRD-History | bird_history | knowledge | 0 |
| BIRD-INTERACT | bird_interact | knowledge | 0 |
| BIRD-SQL (HELM bird_sql / BIRD Dev) | bird_sql | coding | 0 |
| BLiMP (Benchmark of Linguistic Minimal Pairs) | blimp | knowledge | 0 |
| BLiMP-NL (Benchmark of Linguistic Minimal Pairs for Dutch) | blimp_nl | knowledge | 0 |
| BLUEX (Brazilian Leading Universities Entrance eXams) | bluex | knowledge | 0 |
| BMT-Status | shc_bmt | domain | 0 |
| BOLD (Bias in Open-Ended Language Generation Dataset) | bold | safety | 0 |
| Boolean Expressions (BIG-bench) | boolean_expressions | reasoning | 0 |
| BoolQ | boolq | reasoning | 0 |
| Bridging Anaphora Resolution as Question Answering (BARQA-ISNotes) | bridging_anaphora_resolution_barqa | reasoning | 0 |
| BrowseComp | browsecomp | agentic | 3 |
| BuySideFinBench | buysidefinbench | domain | 0 |
| C-Eval | ceval | knowledge | 0 |
| CaBBQ (Catalan Bias Benchmark for Question Answering) | cabbq | safety | 0 |
| CaLM (Causal Evaluation of Language Models) | calm | reasoning | 0 |
| CARDBiomedBench | cardbiomedbench | domain | 0 |
| CareQA | careqa | domain | 0 |
| CaseHOLD (Case Holdings On Legal Decisions) | casehold | domain | 0 |
| CatalanBench | catalan_bench | composite | 0 |
| Causal Judgment (BIG-bench) | causal_judgment | reasoning | 0 |
| Cause and Effect | cause_and_effect | reasoning | 0 |
| CDI-QA | shc_cdi | domain | 0 |
| CHARM (Benchmarking Chinese Commonsense Reasoning of LLMs) | charm | reasoning | 0 |
| ChartQA | chartqa | multimodal | 64 |
| CharXiv | charxiv | multimodal | 0 |
| CharXiv Reasoning | charxiv_reasoning | multimodal | 3 |
| CharXiv Reasoning (with tool use) | charxiv_reasoning_tools | multimodal | 2 |
| Checkmate in One | checkmate_in_one | reasoning | 0 |
| Chem Exam | chem_exam | domain | 0 |
| ChemBench | chembench | domain | 0 |
| Chess State Tracking | chess_state_tracking | reasoning | 0 |
| Chinese Remainder Theorem | chinese_remainder_theorem | math | 0 |
| Chinese SimpleQA | chinese_simpleqa | knowledge | 0 |
| CIBench | cibench | agentic | 0 |
| CIFAR-10 Classification (BIG-bench encodings) | cifar10_classification | multimodal | 0 |
| CIMCQA | ci_mcqa | domain | 0 |
| CivilComments (HELM) | civil_comments | safety | 0 |
| CL-bench | cl_bench_a_benchmark_for_context_learning | long-context | 0 |
| CL-bench | clbench | reasoning | 0 |
| CL-bench Life | cl_bench_life | long-context | 0 |
| ClassEval | class_eval | coding | 0 |
| CLEAR (MedHELM) | clear | domain | 0 |
| CLEVA | cleva | composite | 0 |
| CLIcK | click | knowledge | 0 |
| ClimaQA | climaqa | domain | 0 |
| ClinicBench | clinicbench | domain | 0 |
| ClinicReferral | shc_sequoia | domain | 0 |
| ClozeTest-maxmin | clozetest_maxmin | coding | 0 |
| CLUE (Chinese Language Understanding Evaluation) | clue | composite | 0 |
| CLUE: AFQMC (Ant Financial Question Matching Corpus) | clue_afqmc | composite | 0 |
| CLUE: C3 (free-form multiple-choice Chinese reading comprehension) | clue_c3 | composite | 0 |
| CLUE: CMNLI (Chinese Multi-Genre NLI) | clue_cmnli | composite | 0 |
| CLUE: CMRC 2018 (Simplified Chinese span-extraction reading comprehension) | clue_cmrc | composite | 0 |
| CLUE: DRCD (Traditional Chinese span-extraction reading comprehension) | clue_drcd | composite | 0 |
| CLUE: OCNLI (Original Chinese Natural Language Inference) | clue_ocnli | composite | 0 |
| CMB (Comprehensive Medical Benchmark in Chinese) | cmb | domain | 0 |
| CMMLU (Chinese Massive Multitask Language Understanding) | cmmlu | knowledge | 0 |
| CMO fill-in-the-blank | cmo_fib | math | 0 |
| CMPhysBench | cmphysbench | reasoning | 0 |
| CNN/DailyMail (lm-eval, See et al. v3.0.0) | cnn_dailymail_abisee | generation | 0 |
| CoCo-Bench | coco_bench | coding | 0 |
| CoCoNot | coconot | safety | 0 |
| Code Line Description | code_line_description | coding | 0 |
| CodeCompass | codecompass | coding | 0 |
| CodeInsights Code Efficiency | codeinsights_code_efficiency | coding | 0 |
| CodeInsights Correct Code | codeinsights_correct_code | coding | 0 |
| CodeInsights Edge Case | codeinsights_edge_case | coding | 0 |
| CodeInsights Student Coding | codeinsights_student_coding | coding | 0 |
| CodeInsights Student Mistake | codeinsights_student_mistake | coding | 0 |
| CodeIPI (Indirect Prompt Injection for Coding Agents) | ipi_coding_agent | safety | 0 |
| Codenames | codenames | reasoning | 0 |
| CodeXGLUE | code_x_glue | coding | 0 |
| Color | color | knowledge | 0 |
| Com2Sense | com2sense | reasoning | 0 |
| Common Morpheme | common_morpheme | knowledge | 0 |
| CommonsenseQA | commonsense_qa | reasoning | 0 |
| CommonsenseQA-CN | commonsenseqa_cn | reasoning | 0 |
| CompassBench v1.1 | compassbench_20_v1_1 | composite | 0 |
| CompassBench v1.1 (public) | compassbench_20_v1_1_public | composite | 0 |
| CompassBench v1.3 | compassbench_v1_3 | composite | 0 |
| ComputeEval | compute_eval | coding | 0 |
| Conceptual Combinations | conceptual_combinations | reasoning | 0 |
| Conlang Translation | conlang_translation | translation | 0 |
| Context Definition Alignment | context_definition_alignment | reasoning | 0 |
| Contextual Parametric Knowledge Conflicts | contextual_parametric_knowledge_conflicts | knowledge | 0 |
| ConvFinQACalc | conv_fin_qa_calc | domain | 0 |
| Convince Me | convinceme | safety | 0 |
| COPAL-ID (Choice of Plausible Alternatives — Local Nuances, Indonesia) | copal_id | reasoning | 0 |
| Copyright (HELM memorisation / extraction) | copyright | safety | 0 |
| CoQA (Conversational Question Answering Challenge) | coqa | reasoning | 0 |
| CORE-Bench | core_bench | agentic | 0 |
| CORE-Bench | core_bench_computational_reproducibility_agent_benchmark | agentic | 0 |
| COVIDDialog (HELM English medical dialogue) | covid_dialog | domain | 0 |
| Crash Blossoms (BIG-bench) | crash_blossom | reasoning | 0 |
| CRASS (BIG-bench crass_ai) | crass_ai | reasoning | 0 |
| CrowS-Pairs | crows_pairs | safety | 0 |
| CrowS-Pairs-CN | crowspairs_cn | safety | 0 |
| CRUXEval | cruxeval | coding | 0 |
| Cryobiology Spanish | cryobiology_spanish | knowledge | 0 |
| Cryptonite | cryptonite | reasoning | 0 |
| CS Algorithms | cs_algorithms | reasoning | 0 |
| CSAT-QA | csatqa | domain | 0 |
| CTI-REALM | cti_realm | agentic | 0 |
| CTI-to-MITRE | cti_to_mitre | domain | 0 |
| CValues | cvalues | safety | 0 |
| CVE-Bench | cve_bench | agentic | 0 |
| Cybench | cybench | agentic | 0 |
| CyberGym | cybergym | agentic | 0 |
| CyberMetric | cybermetric | domain | 0 |
| CyberSecEval 2 | cyberseceval_2 | safety | 0 |
| CyberSecEval 3 | cyberseceval_3 | safety | 0 |
| CyberSecEval 4 | cyberseceval_4 | safety | 0 |
| Cycled Letters (BIG-bench) | cycled_letters | reasoning | 0 |
| CzechBankQA | czech_bank_qa | coding | 0 |
| DarijaBench | darija_bench | composite | 0 |
| DarijaHellaSwag | darijahellaswag | reasoning | 0 |
| DarijaMMLU | darijammlu | knowledge | 0 |
| Dark Humor Detection (BIG-bench) | dark_humor_detection | reasoning | 0 |
| Data imputation (HELM) | entity_data_imputation | reasoning | 0 |
| Data Wrangling (BIG-bench) | mult_data_wrangling | reasoning | 0 |
| Date Understanding | date_understanding | reasoning | 0 |
| DecodingTrust Adversarial Demonstrations | decodingtrust_adv_demonstration | safety | 0 |
| DecodingTrust Adversarial Robustness (AdvGLUE++) | decodingtrust_adv_robustness | safety | 0 |
| DecodingTrust Fairness | decodingtrust_fairness | safety | 0 |
| DecodingTrust Machine Ethics | decodingtrust_machine_ethics | safety | 0 |
| DecodingTrust OoD Robustness | decodingtrust_ood_robustness | safety | 0 |
| DecodingTrust Privacy | decodingtrust_privacy | safety | 0 |
| DecodingTrust Stereotype Bias | decodingtrust_stereotype_bias | safety | 0 |
| DecodingTrust Toxicity Prompts | decodingtrust_toxicity_prompts | safety | 0 |
| DeepMind MRCR v2 | deepmind_mrcr_v2 | long-context | 0 |
| DeepSearchQA | deepsearchqa | agentic | 1 |
| Dingo (OpenCompass wrap) | dingo | generation | 0 |
| Disambiguation QA | disambiguation_qa | reasoning | 0 |
| DischargeMe (MedHELM) | dischargeme | domain | 0 |
| Discourse Marker Prediction | discourse_marker_prediction | knowledge | 0 |
| Discrim-Eval | discrim_eval | safety | 0 |
| DISFL-QA | disfl_qa | reasoning | 0 |
| Disinformation (HELM) | disinformation | safety | 0 |
| Diverse Social Bias | diverse_social_bias | safety | 0 |
| DocVQA | docvqa | multimodal | 67 |
| Dr. Bench | dr_bench | agentic | 0 |
| Dr. DocBench | dr_docbench | multimodal | 0 |
| DROP | drop | reasoning | 0 |
| DS-1000 | ds1000 | coding | 0 |
| Dyck Language (HELM) | dyck_language | reasoning | 0 |
| Dyck Languages (BIG-bench) | dyck_languages | reasoning | 0 |
| Dynamic Counting | dynamic_counting | reasoning | 0 |
| Earth-Silver | earth_silver | domain | 0 |
| ECHR Judgment Classification (HELM) | echr_judgment_classification | domain | 0 |
| EESE (Ever-Evolving Science Exam) | eese | knowledge | 0 |
| EgyHellaSwag | egyhellaswag | reasoning | 0 |
| EgyMMLU | egymmlu | knowledge | 0 |
| EHRSHOT (HELM ehrshot) | ehrshot | domain | 0 |
| EHRSQL (HELM ehr_sql / eICU) | ehr_sql | coding | 0 |
| Elementary Math QA | elementary_math_qa | math | 0 |
| Emoji Movie | emoji_movie | knowledge | 0 |
| Emojis Emotion Prediction | emojis_emotion_prediction | reasoning | 0 |
| Empirical Judgments | empirical_judgments | reasoning | 0 |
| ENEM Challenge | enem_challenge | knowledge | 0 |
| English Proverbs | english_proverbs | reasoning | 0 |
| English to Russian Proverbs | english_russian_proverbs | reasoning | 0 |
| ENT-Referral | shc_ent | domain | 0 |
| Entailed Polarity | entailed_polarity | reasoning | 0 |
| Entailed Polarity in Hindi | entailed_polarity_hindi | reasoning | 0 |
| Entity matching (HELM) | entity_matching | reasoning | 0 |
| Epistemic Reasoning | epistemic_reasoning | reasoning | 0 |
| EQ-Bench | eq_bench | reasoning | 0 |
| EQ-Bench (Catalan) | eq_bench_ca | reasoning | 0 |
| EQ-Bench (Spanish) | eq_bench_es | reasoning | 0 |
| ES-MemEval | es_memeval | long-context | 0 |
| EsBBQ (Spanish Bias Benchmark for Question Answering) | esbbq | safety | 0 |
| Estimating Risk of Suicide | suicide_risk | safety | 0 |
| ETHICS (lm-eval hendrycks_ethics) | hendrycks_ethics | safety | 0 |
| EusExams | eus_exams | knowledge | 0 |
| EusProficiency | eus_proficiency | knowledge | 0 |
| EusReading | eus_reading | knowledge | 0 |
| EusTrivia | eus_trivia | knowledge | 0 |
| Evalita-LLM | evalita_llm | composite | 0 |
| Evaluating Information Essentiality | evaluating_information_essentiality | reasoning | 0 |
| Evo-Bench | evo_bench | agentic | 0 |
| EvoEval | evo_eval | coding | 0 |
| EWoK (Elements of World Knowledge) | ewok | knowledge | 0 |
| EXAMS (Multilingual) | exams_multilingual | knowledge | 0 |
| Fact-Checking (BIG-bench) | fact_checker | knowledge | 0 |
| Factuality of Summary | factuality_of_summary | generation | 0 |
| Fake Alignment (FINE) | fake_alignment | safety | 0 |
| Fantasy Reasoning | fantasy_reasoning | reasoning | 0 |
| FDA (BASED information extraction) | fda | domain | 0 |
| Few-shot NLG (BIG-bench) | few_shot_nlg | generation | 0 |
| FewCLUE (Chinese Few-shot Learning Evaluation Benchmark) | fewclue | composite | 0 |
| FewCLUE: BUSTM (Dialogue Short Text Matching) | fewclue_bustm | reasoning | 0 |
| FewCLUE: CHID (Chinese Idiom Cloze Test) | fewclue_chid | reasoning | 0 |
| FewCLUE: CLUEWSC (Winograd Schema Coreference) | fewclue_cluewsc | reasoning | 0 |
| FewCLUE: CSL (Keyword Recognition) | fewclue_csl | reasoning | 0 |
| FewCLUE: EPRSTMT (E-commerce Sentiment Analysis) | fewclue_eprstmt | reasoning | 0 |
| FewCLUE: OCNLI-FC (Natural Language Inference) | fewclue_ocnli_fc | reasoning | 0 |
| FewCLUE: TNEWS (Short News Classification) | fewclue_tnews | reasoning | 0 |
| Figure of Speech Detection (BIG-bench) | figure_of_speech_detection | reasoning | 0 |
| FinanceBench | financebench | domain | 0 |
| FinanceIQ | financeiq | domain | 0 |
| Financial PhraseBank | financial_phrasebank | domain | 0 |
| FinBench | finbench | domain | 13 |
| FinQA | fin_qa | domain | 0 |
| FLD (Formal Logic Deduction) | fld | reasoning | 0 |
| FLORES English-to-Chinese | flores_en_zh | translation | 15 |
| FLORES English-to-German | flores_en_de | translation | 15 |
| FLORES English-to-Japanese | flores_en_ja | translation | 15 |
| FLORES English-to-Spanish | flores_en_es | translation | 15 |
| FLORES-200 | flores | translation | 0 |
| ForecastBench-Sim | forecastbench_sim | reasoning | 0 |
| Forecasting Subquestions (BIG-bench) | forecasting_subquestions | reasoning | 0 |
| Formal Fallacies Syllogisms Negation (BIG-bench) | formal_fallacies_syllogisms_negation | reasoning | 0 |
| FORTRESS | fortress | safety | 0 |
| FrenchBench | french_bench | composite | 0 |
| Frontier-CS | frontier_cs | coding | 0 |
| FrontierScience | frontierscience | domain | 0 |
| FrontierScience — Research track | frontierscience_research | domain | 1 |
| Full-Duplex-Bench | full_duplex_bench | multimodal | 0 |
| Full-Duplex-Bench-v3 | full_duplex_bench_v3 | multimodal | 0 |
| GAIA | gaia | agentic | 0 |
| GalicianBench | galician_bench | composite | 0 |
| Game of 24 | game24 | math | 0 |
| GAOKAO-Bench | gaokaobench | knowledge | 0 |
| GaoKaoMATH | gaokao_math | math | 0 |
| GDM In-house CTF | gdm_in_house_ctf | agentic | 0 |
| GDM InterCode CTF | gdm_intercode_ctf | agentic | 0 |
| GDM Self-proliferation | gdm_self_proliferation | agentic | 0 |
| GDM Self-reasoning | gdm_self_reasoning | agentic | 0 |
| GDM Stealth | gdm_stealth | agentic | 0 |
| GDPval | gdpval | agentic | 0 |
| GEM (BIG-bench) | gem | generation | 0 |
| GenAI-Bench | genai_bench | multimodal | 0 |
| Gender Inclusive Sentences German | gender_inclusive_sentences_german | generation | 0 |
| Gender Sensitivity Test - Chinese | gender_sensitivity_chinese | safety | 0 |
| Gender Sensitivity Test - English | gender_sensitivity_english | safety | 0 |
| General Knowledge | general_knowledge | knowledge | 0 |
| General365 | general365 | reasoning | 0 |
| Geometric Shapes | geometric_shapes | reasoning | 0 |
| Global-MMLU | global_mmlu | knowledge | 0 |
| GLUE (General Language Understanding Evaluation benchmark) | glue | composite | 0 |
| GLUE: CoLA (Corpus of Linguistic Acceptability) | glue_cola | composite | 0 |
| GLUE: MRPC (Microsoft Research Paraphrase Corpus) | glue_mrpc | composite | 0 |
| GLUE: QQP (Quora Question Pairs) | glue_qqp | composite | 0 |
| Goal-step WikiHow (BIG-bench) | goal_step_wikihow | reasoning | 0 |
| Gold Commodity News (HELM) | gold_commodity_news | domain | 0 |
| GovRepcrs (OpenCompass / GovReport CRS) | govrepcrs | long-context | 0 |
| GPQA | gpqa | reasoning | 0 |
| GPQA Diamond | gpqa_diamond | reasoning | 356 |
| Grammar / Best ChatGPT Prompts (HELM Instruct) | grammar | instruction-following | 0 |
| GraphWalks | graphwalks | long-context | 0 |
| GraphWalks BFS (256K-1M context) | graphwalks_bfs_256k_1m | long-context | 3 |
| GraphWalks Parents (256K-1M context) | graphwalks_parents_256k_1m | long-context | 1 |
| GRE Reading Comprehension (BIG-bench) | gre_reading_comprehension | knowledge | 0 |
| GreekMMLU | greekmmlu | knowledge | 0 |
| GroundCocoa | groundcocoa | reasoning | 0 |
| GSM-Hard | gsm_hard | math | 0 |
| GSM8K | gsm8k | math | 130 |
| GSM8K contamination (OpenCompass PPL probe) | gsm8k_contamination | math | 0 |
| GUI-CC | gui_cc | agentic | 0 |
| HAE-RAE Bench (lm-eval haerae) | haerae | knowledge | 0 |
| HarmBench | harm_bench | safety | 0 |
| HarmBench GCG-Transfer (GCG-T) | harm_bench_gcg_transfer | safety | 0 |
| HEAD-QA | headqa | domain | 0 |
| HealthBench | healthbench | domain | 0 |
| HealthBench Hard | healthbench_hard | domain | 1 |
| HealthQA-BR (HELM) | healthqa_br | domain | 0 |
| HEART | heart | human-preference | 0 |
| HEART-Bench | heart_bench | human-preference | 0 |
| HellaSwag | hellaswag | reasoning | 50 |
| HELM MELT information retrieval (Vietnamese mMARCO and mRobust) | melt_ir | knowledge | 0 |
| HELM MELT knowledge (ZaloE2E and ViMMRC) | melt_knowledge | knowledge | 0 |
| HELM MELT synthetic reasoning (abstract symbols) | melt_synthetic_reasoning | reasoning | 0 |
| HELM MELT synthetic reasoning (natural language) | melt_srn | reasoning | 0 |
| HELM Safety | helm_safety | safety | 71 |
| HELM Summarization | summarization | generation | 0 |
| HHH Alignment (BIG-bench) | hhh_alignment | safety | 0 |
| High Low Game (BIG-bench) | high_low_game | reasoning | 0 |
| Hindi Question Answering (BIG-bench) | hindi_question_answering | knowledge | 0 |
| Hindu Knowledge (BIG-bench) | hindu_knowledge | knowledge | 0 |
| Hinglish Toxicity (BIG-bench) | hinglish_toxicity | safety | 0 |
| Histoires Morales | histoires_morales | safety | 0 |
| HMMT 2026 | hmmt2026 | math | 0 |
| HospiceReferral | shc_gip | domain | 0 |
| HRM8K | hrm8k | math | 0 |
| Human Organs and Senses (BIG-bench) | human_organs_senses | knowledge | 0 |
| HumanEval | humaneval | coding | 200 |
| HumanEval Pro | humaneval_pro | coding | 0 |
| HumanEval+ | humaneval_plus | coding | 0 |
| HumanEval-CN | humaneval_cn | coding | 0 |
| HumanEval-Infilling | humaneval_infilling | coding | 0 |
| HumanEval-X | humanevalx | coding | 0 |
| Humanity's Last Exam | hle | reasoning | 6 |
| Humanity's Last Exam (with tools) | hle_tools | reasoning | 5 |
| Hungarian National HS Finals Exam (Mathematics) | hungarian_exam | math | 0 |
| Hyperbaton (BIG-bench) | hyperbaton | instruction-following | 0 |
| Icelandic WinoGrande | icelandic_winogrande | reasoning | 0 |
| Identify Math Theorems (BIG-bench) | identify_math_theorems | math | 0 |
| Identify Odd Metaphor (BIG-bench) | identify_odd_metaphor | reasoning | 0 |
| IFBench | ifbench | instruction-following | 0 |
| IFEval | ifeval | instruction-following | 354 |
| IFEval_ca (Catalan IFEval) | ifeval_ca | instruction-following | 0 |
| IFEval_es (Spanish IFEval) | ifeval_es | instruction-following | 0 |
| IFEvalCode | ifevalcode | instruction-following | 0 |
| IMDb (Large Movie Review Dataset) | imdb | reasoning | 0 |
| IMDb PT-BR (HELM) | imdb_ptbr | reasoning | 0 |
| Implicatures (BIG-bench) | implicatures | reasoning | 0 |
| Implicit Relations (BIG-bench) | implicit_relations | reasoning | 0 |
| Indic Cause and Effect (BIG-bench) | indic_cause_and_effect | reasoning | 0 |
| Indic DiarBench | indic_diarbench | multimodal | 0 |
| INDIC-DIALECT | indic_dialect | composite | 0 |
| IndicXNLI | indicxnli | reasoning | 0 |
| Inference-PPL | inference_ppl | generation | 0 |
| InstrumentalEval | instrumentaleval | safety | 0 |
| Intent Recognition | intent_recognition | domain | 0 |
| InteractiveQA MMLU | interactive_qa_mmlu | knowledge | 0 |
| International Corpus of English | ice | generation | 0 |
| InternSandbox | internsandbox | reasoning | 0 |
| Intersect Geometry | intersect_geometry | math | 0 |
| Inverse IFEval | inverseifeval | instruction-following | 0 |
| Inverse Scaling Prize | inverse_scaling | composite | 0 |
| IPA Natural Language Inference | international_phonetic_alphabet_nli | reasoning | 0 |
| IPA Transliteration | international_phonetic_alphabet_transliterate | translation | 0 |
| IPhO 2025 Theory | ipho_2025_theory | reasoning | 1 |
| Irony Identification (BIG-bench) | irony_identification | reasoning | 0 |
| IWSLT 2017 (OpenCompass English-German) | iwslt2017 | translation | 0 |
| Japanese Leaderboard (lm-evaluation-harness) | japanese_leaderboard | composite | 0 |
| jfinqa | jfinqa | domain | 0 |
| Jigsaw Multilingual Toxic Comment Classification (OpenCompass) | jigsawmultilingual | safety | 0 |
| JSONSchemaBench | jsonschema_bench | generation | 0 |
| K-Bench | k_bench | agentic | 0 |
| K-MetBench | k_metbench | domain | 0 |
| Kanji ASCII Art (BIG-bench) | kanji_ascii | reasoning | 0 |
| Kannada Riddles (BIG-bench) | kannada | reasoning | 0 |
| Kaoshi (OpenCompass) | kaoshi | knowledge | 0 |
| KBL (Korean Benchmark for Legal Language Understanding) | kbl | domain | 0 |
| KCLE (OpenCompass) | kcle | knowledge | 0 |
| KernelBench | kernelbench | coding | 0 |
| Key/Value Maps | key_value_maps | reasoning | 0 |
| KMMLU (Korean-MMLU) | kmmlu | knowledge | 0 |
| Known Unknowns | known_unknowns | knowledge | 0 |
| Koala (HELM Instruct) | koala | instruction-following | 0 |
| KoBEST (Korean Balanced Evaluation of Significant Tasks) | kobest | composite | 0 |
| KOR-Bench | korbench | reasoning | 0 |
| KorMedMCQA | kormedmcqa | domain | 0 |
| KPI-EDGAR | kpi_edgar | domain | 0 |
| L-Eval | leval | long-context | 0 |
| LAB-Bench | lab_bench | domain | 0 |
| LAB-Bench: FigQA | lab_bench_figqa | domain | 2 |
| LAB-Bench: FigQA (with tools) | lab_bench_figqa_tools | domain | 2 |
| LAMBADA | lambada | reasoning | 0 |
| LAMBADA Cloze | lambada_cloze | reasoning | 0 |
| LAMBADA multilingual (OpenAI MT) | lambada_multilingual | reasoning | 0 |
| LAMBADA multilingual (Stable LM translations) | lambada_multilingual_stablelm | reasoning | 0 |
| Language Games (BIG-bench) | language_games | translation | 0 |
| Language Identification (BIG-bench Lite) | language_identification | knowledge | 0 |
| LawBench | lawbench | domain | 0 |
| LCBench2023 | lcbench | coding | 0 |
| LCSTS (Large-scale Chinese Short Text Summarization) | lcsts | generation | 0 |
| leaderboard-dataset | leaderboard_dataset | knowledge | 0 |
| leaderboard-details | leaderboard_details | knowledge | 0 |
| leaderboard-requests | leaderboard_requests | knowledge | 0 |
| leaderboard-results | leaderboard_results | knowledge | 0 |
| Legal Contract Summarization (HELM) | legal_contract_summarization | domain | 0 |
| Legal Opinion Sentiment Classification (HELM) | legal_opinion_sentiment_classification | domain | 0 |
| Legal summarization (HELM) | legal_summarization | domain | 0 |
| LegalBench | legalbench | domain | 20 |
| LegalSupport | legal_support | domain | 0 |
| LexGLUE (Legal General Language Understanding Evaluation) | lex_glue | domain | 0 |
| LEXTREME | lextreme | domain | 0 |
| LIBRA (Long Input Benchmark for Russian Analysis) | libra | long-context | 0 |
| LingOly | lingoly | reasoning | 0 |
| Linguistic Mappings | linguistic_mappings | reasoning | 0 |
| Linguistics Puzzles | linguistics_puzzles | reasoning | 0 |
| List Functions | list_functions | reasoning | 0 |
| LIT-RAGBench | lit_ragbench | reasoning | 0 |
| LiveCodeBench | live_code_bench | coding | 144 |
| LiveCodeBench Pro | livecodebench_pro | coding | 0 |
| LiveMathBench | livemathbench | math | 0 |
| LiveQA (TREC-2017 Medical Task) | live_qa | domain | 0 |
| LiveReasonBench | livereasonbench | reasoning | 0 |
| LiveStemBench | livestembench | domain | 0 |
| LLM Compression | llm_compression | generation | 0 |
| LLM Stats | llm_stats | knowledge | 0 |
| LLM-QBench | llm_qbench | knowledge | 0 |
| LLM-SoccerArena | llm_soccerarena | knowledge | 0 |
| LM-SynEval (Targeted Syntactic Evaluation of Language Models) | lm_syneval | knowledge | 0 |
| LMentry | lm_entry | reasoning | 0 |
| Logic Grid Puzzle | logic_grid_puzzle | reasoning | 0 |
| Logical Arguments | logical_args | reasoning | 0 |
| Logical Deduction | logical_deduction | reasoning | 0 |
| Logical Fallacy Detection | logical_fallacy_detection | reasoning | 0 |
| Logical Sequence | logical_sequence | reasoning | 0 |
| LogiQA | logiqa | reasoning | 0 |
| LogiQA 2.0 | logiqa2 | reasoning | 0 |
| Long Context Integration | long_context_integration | long-context | 0 |
| LongBench | longbench | long-context | 0 |
| LongBench v2 | longbenchv2 | long-context | 0 |
| LongProc | longproc | long-context | 0 |
| LSAT (Analytical Reasoning) | lsat_qa | reasoning | 0 |
| LV-Eval | lveval | long-context | 0 |
| M3-BENCH | m3_bench | human-preference | 0 |
| M3-DuplexBench | m3_duplexbench | generation | 0 |
| MaCBench | macbench | domain | 0 |
| MadinahQA | madinah_qa | domain | 0 |
| Make Me Pay | make_me_pay | safety | 0 |
| MakeMeSay | makemesay | safety | 0 |
| MAS-Bench | mas_bench | agentic | 0 |
| MASBench | mas_orchestra | agentic | 0 |
| MASK (Model Alignment between Statements and Knowledge) | mask | safety | 0 |
| Mastermath2024v1 (OpenCompass) | mastermath2024v1 | math | 0 |
| MastermindEval | mastermind | reasoning | 0 |
| Matbench | matbench | domain | 0 |
| MATH (Mathematics Aptitude Test of Heuristics) | math | math | 0 |
| MATH 401 | math401 | math | 0 |
| MATH-500 | math_500 | math | 354 |
| MathBench | mathbench | math | 0 |
| Mathematical Induction (BIG-bench) | mathematical_induction | math | 0 |
| MathQA | mathqa | math | 0 |
| MathVista | mathvista | math | 66 |
| Matrix Shapes (BIG-bench) | matrixshapes | math | 0 |
| MBPP (Mostly Basic Python Problems) | mbpp | coding | 0 |
| MBPP Pro | mbpp_pro | coding | 0 |
| MBPP+ | mbpp_plus | coding | 0 |
| MBPP-CN | mbpp_cn | coding | 0 |
| MC-TACO | mc_taco | reasoning | 0 |
| MCP-Bench | mcp_bench | agentic | 0 |
| MedAlign | medalign | instruction-following | 0 |
| MedBench | medbench | domain | 0 |
| Medbullets | medbullets | domain | 0 |
| MedCalc-Bench | medcalc_bench | domain | 0 |
| MedConceptsQA | med_concepts_qa | domain | 0 |
| MedConfInfo | shc_conf | domain | 0 |
| MedDialog | med_dialog | domain | 0 |
| MEDEC | medec | domain | 0 |
| MedHallu | medhallu | safety | 0 |
| MedHELM Configurable | medhelm_configurable | domain | 0 |
| Medical Questions Russian | medical_questions_russian | domain | 0 |
| MedicationQA | medication_qa | domain | 0 |
| MEDIQA (HELM) | medi_qa | domain | 0 |
| MEDIQA 2019 QA (lm-eval) | mediqa_qa2019 | domain | 0 |
| MedMCQA | medmcqa | domain | 4 |
| MedQA | medqa | domain | 51 |
| MedText (lm-eval) | medtext | domain | 0 |
| MedXpertQA MM | medxpertqa_multimodal | domain | 1 |
| MedXpertQA Text | medxpertqa | domain | 0 |
| MELT translation (HELM Vietnamese OPUS-100 and PhoMT) | melt_translation | translation | 0 |
| MentalHealth (MedHELM) | mental_health | domain | 0 |
| MeQSum | meqsum | generation | 0 |
| metabench | metabench | composite | 0 |
| Metaphor Boolean (BIG-bench) | metaphor_boolean | reasoning | 0 |
| Metaphor Understanding (BIG-bench) | metaphor_understanding | reasoning | 0 |
| MGSM (Multilingual Grade School Math) | mgsm | math | 45 |
| MIMIC-BHC (MedHELM) | mimic_bhc | domain | 0 |
| MIMIC-III Report Summarization (lm-eval) | mimic_repsum | domain | 0 |
| MIMIC-IV Billing Code (MedHELM) | mimiciv_billing_code | domain | 0 |
| MIMIC-RRS (MedHELM) | mimic_rrs | domain | 0 |
| Mind2Web | mind2web | agentic | 0 |
| Mind2Web-SC | mind2web_sc | safety | 0 |
| MINT (Medical Incremental N-Turn Benchmark) | benchmarking_multi_turn_medical_diagnosis | domain | 0 |
| Minute Mysteries QA | minute_mysteries_qa | reasoning | 0 |
| MIRACL | miracl | embedding | 90 |
| Misconceptions | misconceptions | knowledge | 0 |
| Misconceptions (Russian) | misconceptions_russian | knowledge | 0 |
| MLE-bench | mle_bench | agentic | 0 |
| MLQA | mlqa | knowledge | 0 |
| MLRC-Bench | mlrc_bench | agentic | 0 |
| MMBench | mmbench | multimodal | 0 |
| MMIU | mmiu | multimodal | 0 |
| MMLU (Massive Multitask Language Understanding) | mmlu | knowledge | 0 |
| MMLU clinical African languages (HELM) | mmlu_clinical_afr | knowledge | 0 |
| MMLU-CF | mmlu_cf | knowledge | 0 |
| MMLU-Pro | mmlu_pro | knowledge | 354 |
| MMLU-Pro+ | mmlu_pro_plus | reasoning | 0 |
| MMLU-ProX | mmlu_prox | knowledge | 0 |
| MMLU-Redux | mmlu_redux | knowledge | 0 |
| MMLU-SR | mmlusr | reasoning | 0 |
| MMLU: Abstract Algebra | mmlu_abstract_algebra | knowledge | 50 |
| MMLU: Anatomy | mmlu_anatomy | knowledge | 50 |
| MMLU: Astronomy | mmlu_astronomy | knowledge | 90 |
| MMLU: Biology (subcategory) | mmlu_biology | knowledge | 40 |
| MMLU: Business Ethics | mmlu_business_ethics | knowledge | 90 |
| MMLU: Chemistry (subcategory) | mmlu_chemistry | knowledge | 40 |
| MMLU: Clinical Knowledge | mmlu_clinical_knowledge | knowledge | 90 |
| MMLU: College Biology | mmlu_college_biology | knowledge | 50 |
| MMLU: College Chemistry | mmlu_college_chemistry | knowledge | 50 |
| MMLU: College Computer Science | mmlu_college_computer_science | knowledge | 50 |
| MMLU: College Mathematics | mmlu_college_mathematics | knowledge | 50 |
| MMLU: College Medicine | mmlu_college_medicine | knowledge | 50 |
| MMLU: College Physics | mmlu_college_physics | knowledge | 50 |
| MMLU: Computer Science (subcategory) | mmlu_computer_science | knowledge | 40 |
| MMLU: Computer Security | mmlu_computer_security | knowledge | 50 |
| MMLU: Conceptual Physics | mmlu_conceptual_physics | knowledge | 50 |
| MMLU: Econometrics | mmlu_econometrics | knowledge | 50 |
| MMLU: Electrical Engineering | mmlu_electrical_engineering | knowledge | 50 |
| MMLU: Elementary Mathematics | mmlu_elementary_mathematics | knowledge | 50 |
| MMLU: Formal Logic | mmlu_formal_logic | knowledge | 50 |
| MMLU: Global Facts | mmlu_global_facts | knowledge | 50 |
| MMLU: High School Biology | mmlu_high_school_biology | knowledge | 50 |
| MMLU: High School Chemistry | mmlu_high_school_chemistry | knowledge | 50 |
| MMLU: High School Computer Science | mmlu_high_school_computer_science | knowledge | 50 |
| MMLU: High School European History | mmlu_high_school_european_history | knowledge | 50 |
| MMLU: High School Geography | mmlu_high_school_geography | knowledge | 50 |
| MMLU: High School Government and Politics | mmlu_high_school_government_and_politics | knowledge | 50 |
| MMLU: High School Macroeconomics | mmlu_high_school_macroeconomics | knowledge | 50 |
| MMLU: High School Mathematics | mmlu_high_school_mathematics | knowledge | 50 |
| MMLU: High School Microeconomics | mmlu_high_school_microeconomics | knowledge | 50 |
| MMLU: High School Physics | mmlu_high_school_physics | knowledge | 50 |
| MMLU: High School Psychology | mmlu_high_school_psychology | knowledge | 50 |
| MMLU: High School Statistics | mmlu_high_school_statistics | knowledge | 50 |
| MMLU: High School US History | mmlu_high_school_us_history | knowledge | 50 |
| MMLU: High School World History | mmlu_high_school_world_history | knowledge | 50 |
| MMLU: Human Aging | mmlu_human_aging | knowledge | 50 |
| MMLU: Human Sexuality | mmlu_human_sexuality | knowledge | 50 |
| MMLU: International Law | mmlu_international_law | knowledge | 50 |
| MMLU: Jurisprudence | mmlu_jurisprudence | knowledge | 90 |
| MMLU: Logical Fallacies | mmlu_logical_fallacies | knowledge | 50 |
| MMLU: Machine Learning | mmlu_machine_learning | knowledge | 50 |
| MMLU: Management | mmlu_management | knowledge | 50 |
| MMLU: Marketing | mmlu_marketing | knowledge | 50 |
| MMLU: Medical Genetics | mmlu_medical_genetics | knowledge | 50 |
| MMLU: Miscellaneous | mmlu_miscellaneous | knowledge | 50 |
| MMLU: Moral Disputes | mmlu_moral_disputes | knowledge | 50 |
| MMLU: Moral Scenarios | mmlu_moral_scenarios | knowledge | 50 |
| MMLU: Nutrition | mmlu_nutrition | knowledge | 50 |
| MMLU: Philosophy | mmlu_philosophy | knowledge | 50 |
| MMLU: Physics (unresolved key) | mmlu_physics | knowledge | 40 |
| MMLU: Prehistory | mmlu_prehistory | knowledge | 50 |
| MMLU: Professional Accounting | mmlu_professional_accounting | knowledge | 90 |
| MMLU: Professional Law | mmlu_professional_law | knowledge | 90 |
| MMLU: Professional Medicine | mmlu_professional_medicine | knowledge | 50 |
| MMLU: Professional Psychology | mmlu_professional_psychology | knowledge | 50 |
| MMLU: Public Relations | mmlu_public_relations | knowledge | 50 |
| MMLU: Security Studies | mmlu_security_studies | knowledge | 50 |
| MMLU: Sociology | mmlu_sociology | knowledge | 50 |
| MMLU: US Foreign Policy | mmlu_us_foreign_policy | knowledge | 50 |
| MMLU: Virology | mmlu_virology | knowledge | 50 |
| MMLU: World Religions | mmlu_world_religions | knowledge | 50 |
| MMLUArabic (AceGPT translated MMLU) | mmluarabic | knowledge | 0 |
| MMMLU-lite | mmmlu_lite | knowledge | 0 |
| MMMU | mmmu | multimodal | 68 |
| MMMU-Pro | mmmu_pro | multimodal | 0 |
| Model-Written Evaluations | model_written_evals | safety | 0 |
| Modified Arithmetic | modified_arithmetic | math | 0 |
| Mol-Instructions molecule-oriented tasks | molinstructions_chem | domain | 0 |
| MolecularIQ | molculariq | domain | 0 |
| Moral Permissibility | moral_permissibility | reasoning | 0 |
| Moral Stories | moral_stories | safety | 0 |
| MORU (Moral Reasoning under Uncertainty) | moru | safety | 0 |
| Movie Dialogue Same or Different (BIG-bench) | movie_dialog_same_or_different | reasoning | 0 |
| Movie Recommendation (BIG-bench) | movie_recommendation | knowledge | 0 |
| MP-20 (OpenCompass) | mp20 | domain | 0 |
| MRAG | mrag | domain | 0 |
| MRAG-Bench | mrag_bench | multimodal | 0 |
| MS MARCO (HELM passage ranking) | msmarco | knowledge | 0 |
| MT-Bench | mt_bench | human-preference | 46 |
| MTEB (Massive Text Embedding Benchmark) | mteb | embedding | 0 |
| MTEB Classification | mteb_classification | embedding | 94 |
| MTEB Clustering | mteb_clustering | embedding | 94 |
| MTEB Overall (leaderboard average) | mteb_overall | embedding | 96 |
| MTEB Pair Classification | mteb_pair_classification | embedding | 93 |
| MTEB Reranking | mteb_reranking | embedding | 93 |
| MTEB Retrieval | mteb_retrieval | embedding | 96 |
| MTEB STS (Semantic Textual Similarity) | mteb_sts | embedding | 93 |
| MTEB Summarization | mteb_summarization | embedding | 93 |
| MTEB-BR | mteb_br | embedding | 0 |
| MTR-Bench | mtr_bench | reasoning | 0 |
| MTR-Suite | mtr_suite | composite | 0 |
| MTS-Dialog (lm-eval) | mts_dialog | domain | 0 |
| MTSamples Procedures (MedHELM) | mtsamples_procedures | domain | 0 |
| MTSamples Replicate (MedHELM) | mtsamples_replicate | domain | 0 |
| MULTI | multi | multimodal | 0 |
| MULTI-Bench | multi_bench | human-preference | 0 |
| Multi-IF | multiif | instruction-following | 0 |
| Multi-SWE-bench | multi_swe_bench | coding | 0 |
| MultiBLiMP 1.0 | multiblimp | knowledge | 0 |
| MultiEmo (BIG-bench) | multiemo | knowledge | 0 |
| Multilingual MMLU (MMMLU) | mmmlu | knowledge | 2 |
| MultiPL-E | multipl_e | coding | 37 |
| MultiPL-E: C# | multipl_e_csharp | coding | 135 |
| MultiPL-E: C++ | multipl_e_cpp | coding | 37 |
| MultiPL-E: Go | multipl_e_go | coding | 37 |
| MultiPL-E: Java | multipl_e_java | coding | 37 |
| MultiPL-E: JavaScript | multipl_e_javascript | coding | 37 |
| MultiPL-E: Julia | multipl_e_julia | coding | 129 |
| MultiPL-E: Kotlin (unconfirmed) | multipl_e_kotlin | coding | 135 |
| MultiPL-E: Lua | multipl_e_lua | coding | 135 |
| MultiPL-E: Perl | multipl_e_perl | coding | 129 |
| MultiPL-E: PHP | multipl_e_php | coding | 135 |
| MultiPL-E: Python | multipl_e_python | coding | 37 |
| MultiPL-E: R | multipl_e_r | coding | 129 |
| MultiPL-E: Ruby | multipl_e_ruby | coding | 135 |
| MultiPL-E: Rust | multipl_e_rust | coding | 37 |
| MultiPL-E: Scala | multipl_e_scala | coding | 135 |
| MultiPL-E: Swift | multipl_e_swift | coding | 135 |
| MultiPL-E: TypeScript | multipl_e_typescript | coding | 37 |
| Multistep Arithmetic (BIG-bench) | multistep_arithmetic | math | 0 |
| Muslim-Violence Bias (BIG-bench) | muslim_violence_bias | safety | 0 |
| MuSR | musr | reasoning | 221 |
| MuTual | mutual | reasoning | 0 |
| MV-Bench | mv_bench | multimodal | 0 |
| MV-dVRK | mv_dvrk | multimodal | 0 |
| N2C2-CT Matching (HELM) | n2c2_ct_matching | domain | 0 |
| NarrativeQA | narrativeqa | long-context | 0 |
| Natural Questions (HELM) | natural_qa | knowledge | 0 |
| natural_instructions (BIG-bench Natural Instructions) | natural_instructions | instruction-following | 0 |
| Navigate (BIG-bench) | navigate | reasoning | 0 |
| NeedleBench | needlebench | long-context | 0 |
| NeedleBench V2 | needlebench_v2 | long-context | 0 |
| NEJMAI / nephSAP nephrology benchmark | nejm_ai_benchmark | domain | 0 |
| NewsQA | newsqa | knowledge | 0 |
| NIAH (Inspect Evals Needle in a Haystack) | niah | long-context | 0 |
| Nonsense Words Grammar (BIG-bench) | nonsense_words_grammar | reasoning | 0 |
| NorEval | noreval | composite | 0 |
| NoteExtract (MedHELM chw_care_plan) | chw_care_plan | domain | 0 |
| Novel Concepts (BIG-bench) | novel_concepts | reasoning | 0 |
| NoveltyBench | novelty_bench | generation | 0 |
| NPHardEval | nphardeval | reasoning | 0 |
| NQ-CN (OpenCompass) | nq_cn | knowledge | 0 |
| NQ-Open | nq_open | knowledge | 0 |
| NYU-LLM-CTF/CTFTiny | nyu_llm_ctf_ctftiny | knowledge | 0 |
| NYU-LLM-CTF/NYU_CTF_Bench | nyu_llm_ctf_nyu_ctf_bench | knowledge | 0 |
| O-NET (Inspect Evals) | onet | knowledge | 0 |
| OAB Exams | oab_exams | domain | 0 |
| Object Counting (BIG-bench) | object_counting | reasoning | 0 |
| OCRBench | ocrbench | multimodal | 21 |
| Odd One Out (BIG-bench) | odd_one_out | reasoning | 0 |
| OECD Integrity | oecd_integrity | knowledge | 0 |
| OECD Outlook | oecd_outlook | knowledge | 0 |
| OECD PISA | oecd_pisa | knowledge | 0 |
| OECD Statistics | oecd_statistics | knowledge | 0 |
| OJBench | ojbench | coding | 0 |
| Okapi ARC multilingual (lm-eval) | okapi_arc_multilingual | reasoning | 0 |
| Okapi multilingual HellaSwag | okapi_hellaswag_multilingual | reasoning | 0 |
| Okapi multilingual MMLU | okapi_mmlu_multilingual | knowledge | 0 |
| Okapi multilingual TruthfulQA | okapi_truthfulqa_multilingual | safety | 0 |
| OLAPH / MedLFQA | olaph | domain | 0 |
| OlymMATH | olymmath | math | 0 |
| OlympiadBench | olympiadbench | multimodal | 0 |
| Omni-MATH | omni_math | math | 0 |
| Open Arabic LLM Leaderboard — Complete configuration | arabic_leaderboard_complete | composite | 0 |
| Open Arabic LLM Leaderboard — Light configuration | arabic_leaderboard_light | composite | 0 |
| Open Assistant (HELM Instruct) | open_assistant | instruction-following | 0 |
| OpenAI MRCR | openai_mrcr | long-context | 0 |
| OpenCompass biodata (biology-instruction) | biodata | domain | 0 |
| OpenCompass safety (Perspective toxicity) | safety | safety | 0 |
| OpenFinData | openfindata | domain | 0 |
| OpenML Explain | openml_explain | knowledge | 0 |
| OpenSWI | openswi | domain | 0 |
| Operators | operators | math | 0 |
| OpinionQA | opinions_qa | safety | 0 |
| OPT-BENCH | opt_bench | knowledge | 0 |
| OPT-Engine | opt_engine | knowledge | 0 |
| ORT (Out-of-Distribution Robustness Testing) | llm_unlearning_should_be_form_independent | safety | 0 |
| OSWorld | osworld | agentic | 3 |
| P-MMEval | pmmeval | composite | 0 |
| Paloma | paloma | generation | 0 |
| PaperBench | paperbench | agentic | 0 |
| Paragraph Segmentation | paragraph_segmentation | generation | 0 |
| Paragraph-level Simplification of Medical Texts | med_paragraph_simplification | generation | 0 |
| PARSINLU QA | parsinlu_qa | knowledge | 0 |
| ParsiNLU Reading Comprehension | parsinlu_reading_comprehension | knowledge | 0 |
| PAWS | paws | reasoning | 0 |
| PAWS-X | paws_x | reasoning | 0 |
| Penguins in a Table | penguins_in_a_table | reasoning | 0 |
| PennyLane QML Benchmarks | pennylane_qml | domain | 0 |
| Periodic Elements (BIG-bench) | periodic_elements | knowledge | 0 |
| Persian Idioms (BIG-bench) | persian_idioms | knowledge | 0 |
| PersistBench | persistbench | safety | 0 |
| Personality (Inspect Evals) | personality | domain | 0 |
| PerspectiveGap | perspectivegap | agentic | 0 |
| Phrase Relatedness (BIG-bench) | phrase_relatedness | knowledge | 0 |
| PHYBench | phybench | reasoning | 0 |
| Physical Intuition (BIG-bench) | physical_intuition | reasoning | 0 |
| PHYSICS (Benchmarking Foundation Models on University-Level Physics Problem Solving) | physics | reasoning | 0 |
| Physics GRE (Inflection-Benchmarks) | physics_gre | knowledge | 0 |
| physics_questions (BIG-bench) | physics_questions | math | 0 |
| PI-LLM | pi_llm | long-context | 0 |
| Pile-10k | pile_10k | generation | 0 |
| PIQA | piqa | reasoning | 0 |
| PISA-Bench | pisa | multimodal | 0 |
| PJExam | pjexam | knowledge | 0 |
| Play Dialogue Same or Different (BIG-bench) | play_dialog_same_or_different | reasoning | 0 |
| PMC-Patients | pmc_patients | embedding | 0 |
| PolEmo 2.0 | polemo2 | domain | 0 |
| Polish Sequence Labeling (BIG-bench) | polish_sequence_labeling | domain | 0 |
| PortugueseBench | portuguese_bench | composite | 0 |
| Pre-Flight | pre_flight | knowledge | 0 |
| Presuppositions as NLI (BIG-bench) | presuppositions_as_nli | reasoning | 0 |
| PRiSM | prism | multimodal | 0 |
| PRISM-Bench | prism_bench | multimodal | 0 |
| PrivacyDetection | shc_privacy | domain | 0 |
| ProcessBench | processbench | math | 0 |
| PromptBench | promptbench | safety | 0 |
| ProofBench | proofbench | math | 0 |
| PROST | prost | reasoning | 0 |
| Protein Interacting Sites (BIG-bench) | protein_interacting_sites | domain | 0 |
| ProteinLMBench | proteinlmbench | domain | 0 |
| ProxySender | shc_proxy | domain | 0 |
| PubMedQA | pubmedqa | domain | 4 |
| Putnam-AXIOM | putnam_axiom | math | 0 |
| PY150 (OpenCompass line completion) | py150 | coding | 0 |
| Python Program Synthesis (BIG-bench) | program_synthesis | coding | 0 |
| Python Programming Challenge | python_programming_challenge | coding | 0 |
| QA WikiData | qa_wikidata | knowledge | 0 |
| QA4MRE | qa4mre | reasoning | 0 |
| qabench | qabench | knowledge | 0 |
| QASPER | qasper | long-context | 0 |
| QASPER-cut | qaspercut | long-context | 0 |
| QuAC (Question Answering in Context) | quac | reasoning | 0 |
| QuALITY | quality | long-context | 0 |
| Question Selection (BIG-bench) | question_selection | reasoning | 0 |
| Question-Answer Creation | question_answer_creation | generation | 0 |
| R-Bench (Reasoning Bench) | r_bench | reasoning | 0 |
| RACE | race | reasoning | 0 |
| RACE-H (inspect_evals) | race_h | reasoning | 0 |
| RaceBias (HELM race_based_med) | race_based_med | safety | 0 |
| RAFT (Real-world Annotated Few-shot Tasks) | raft | domain | 0 |
| RE-Bench | re_bench | agentic | 0 |
| Real or Fake Text (RoFT, BIG-bench) | real_or_fake_text | generation | 0 |
| RealToxicityPrompts | real_toxicity_prompts | safety | 0 |
| RealWorldQA | realworldqa | multimodal | 8 |
| Reasoning about Colored Objects (BIG-bench) | reasoning_about_colored_objects | reasoning | 0 |
| Reordering | undo_permutation | reasoning | 0 |
| Repeat Copy Logic (BIG-bench) | repeat_copy_logic | reasoning | 0 |
| Rephrase (BIG-bench) | rephrase | generation | 0 |
| Rhyming (BIG-bench) | rhyming | knowledge | 0 |
| RiddleSense (BIG-bench) | riddle_sense | reasoning | 0 |
| RO-Bench | ro_bench | multimodal | 0 |
| RO-N3WS | ro_n3ws | domain | 0 |
| RoleBench | rolebench | generation | 0 |
| Roots, Optimization and Games (BIG-bench) | roots_optimization_and_games | math | 0 |
| Ruin Names (BIG-bench) | ruin_names | reasoning | 0 |
| RULER | ruler | long-context | 0 |
| S2-TOMG-Bench | s2_tomg_bench | domain | 0 |
| S3Eval | s3eval | long-context | 0 |
| SAD (Situational Awareness Dataset) | sad | safety | 0 |
| SAGE | sage | agentic | 0 |
| Salient Translation Error Detection (BIG-bench) | salient_translation_error_detection | translation | 0 |
| scBench | scbench | agentic | 0 |
| SciBench | scibench | reasoning | 0 |
| ScienceQA | scienceqa | multimodal | 0 |
| Scientific Press Release (BIG-bench) | scientific_press_release | generation | 0 |
| SciEval | scieval | domain | 0 |
| SciKnowEval | sciknoweval | domain | 0 |
| SciQ | sciq | knowledge | 0 |
| SciReasoner | scireasoner | domain | 0 |
| SciReasoner 1.5 (OpenCompass) | scireasoner1_5 | domain | 0 |
| SCORE (Systematic COnsistency and Robustness Evaluation) | score | composite | 0 |
| ScreenSpot-Pro | screenspot_pro | agentic | 2 |
| ScreenSpot-Pro (with tools) | screenspot_pro_tools | agentic | 2 |
| SCROLLS (Standardized CompaRison Over Long Language Sequences) | scrolls | long-context | 0 |
| SE-Bench | se_bench | coding | 0 |
| SE-Eval | se_eval | multimodal | 0 |
| SEA-HELM (Southeast Asian Holistic Evaluation of Language Models) | seahelm | composite | 0 |
| SecQA | sec_qa | domain | 0 |
| SeedBench | seedbench | domain | 0 |
| Self Instruct (HELM) | self_instruct | instruction-following | 0 |
| self_awareness (BIG-bench) | self_awareness | reasoning | 0 |
| self_evaluation_courtroom (BIG-bench) | self_evaluation_courtroom | reasoning | 0 |
| self_evaluation_tutoring (BIG-bench Self Evaluation of Tutoring) | self_evaluation_tutoring | reasoning | 0 |
| semantic_parsing_in_context_sparc (BIG-bench SParC) | semantic_parsing_in_context_sparc | coding | 0 |
| semantic_parsing_spider (BIG-bench Spider) | semantic_parsing_spider | coding | 0 |
| Sentence Ambiguity | sentence_ambiguity | reasoning | 0 |
| SEvenLLM (SEvenLLM-Bench) | sevenllm | domain | 0 |
| Similarities Test for Abstraction | similarities_abstraction | reasoning | 0 |
| Simple Cooccurrence Bias | simple_cooccurrence_bias | safety | 0 |
| Simple Ethical Questions (BIG-bench) | simple_ethical_questions | safety | 0 |
| Simple Text Editing (BIG-bench) | simple_text_editing | instruction-following | 0 |
| simple_arithmetic (BIG-bench programmatic addition template) | simple_arithmetic | math | 0 |
| simple_arithmetic_json (BIG-bench JSON arithmetic template) | simple_arithmetic_json | math | 0 |
| simple_arithmetic_json_multiple_choice (BIG-bench JSON MC arithmetic template) | simple_arithmetic_json_multiple_choice | math | 0 |
| simple_arithmetic_json_subtasks (BIG-bench nested JSON arithmetic template) | simple_arithmetic_json_subtasks | math | 0 |
| simple_arithmetic_multiple_targets_json (BIG-bench multi-target JSON template) | simple_arithmetic_multiple_targets_json | math | 0 |
| SimpleQA | simpleqa | knowledge | 0 |
| SimpleSafetyTests | simple_safety_tests | safety | 0 |
| SkillsBench | skillsbench | agentic | 0 |
| SLM-Bench | slm_bench | composite | 0 |
| SLR-Bench (group) | slr_bench_group | reasoning | 0 |
| SMolInstruct | smolinstruct | domain | 0 |
| SNARKS (BIG-bench sarcasm contrast set) | snarks | reasoning | 0 |
| Social Bias from Sentence Probability (BIG-bench) | bias_from_probabilities | safety | 0 |
| Social IQa | siqa | reasoning | 0 |
| Social Support | social_support | safety | 0 |
| SOSBench | sosbench | safety | 0 |
| SPADE-Bench | spade_bench | safety | 0 |
| SpanishBench | spanish_bench | composite | 0 |
| Spelling Bee | spelling_bee | reasoning | 0 |
| Spider | spider | coding | 0 |
| Sports Understanding | sports_understanding | knowledge | 0 |
| SQuAD | squad | reasoning | 0 |
| SQuAD 2.0 (lm-evaluation-harness squadv2) | squadv2 | reasoning | 0 |
| SQuAD 2.0 (OpenCompass squad20) | squad20 | reasoning | 0 |
| SQuAD completion (Based / lm-eval) | squad_completion | reasoning | 0 |
| squad_shifts (BIG-bench SQuADShifts) | squad_shifts | reasoning | 0 |
| SRBench | srbench | math | 0 |
| STARR Patient Instructions (PatientInstruct) | starr_patient_instructions | domain | 0 |
| StereoSet | stereoset | safety | 0 |
| Story Cloze Test | storycloze | reasoning | 0 |
| Strange Stories | strange_stories | reasoning | 0 |
| StrategyQA | strategyqa | reasoning | 0 |
| StrongREJECT | strong_reject | safety | 0 |
| Subject-Verb Agreement | subject_verb_agreement | reasoning | 0 |
| Sudoku | sudoku | reasoning | 0 |
| Sufficient Information | sufficient_information | reasoning | 0 |
| SummEdits | summedits | reasoning | 0 |
| SummScreen | summscreen | generation | 0 |
| SUMOSum (HELM climate-claims summarization) | sumosum | domain | 0 |
| SuperCLUE-Agent | superclue_agent | agentic | 0 |
| SuperCLUE-Safety | superclue_safety | safety | 0 |
| SuperGLUE (Super General Language Understanding Evaluation benchmark) | super_glue | composite | 0 |
| SuperGLUE AX-b (Broad Coverage Diagnostics) | superglue_ax_b | reasoning | 0 |
| SuperGLUE AX-g (Winogender Schema Diagnostics) | superglue_ax_g | safety | 0 |
| SuperGLUE CB (CommitmentBank) | superglue_cb | reasoning | 0 |
| SuperGLUE COPA (Choice of Plausible Alternatives) | superglue_copa | reasoning | 0 |
| SuperGLUE MultiRC (Multi-Sentence Reading Comprehension) | superglue_multirc | reasoning | 0 |
| SuperGLUE ReCoRD (Reading Comprehension with Commonsense Reasoning Dataset) | superglue_record | reasoning | 0 |
| SuperGLUE RTE (Recognizing Textual Entailment) | superglue_rte | reasoning | 0 |
| SuperGLUE WiC (Word-in-Context) | superglue_wic | knowledge | 0 |
| SuperGLUE WSC (Winograd Schema Challenge, SuperGLUE recast) | superglue_wsc | reasoning | 0 |
| SuperGPQA | supergpqa | knowledge | 0 |
| SVAMP | svamp | math | 0 |
| SWAG (Situations With Adversarial Generations) | swag | reasoning | 0 |
| Swahili-English Proverbs | swahili_english_proverbs | translation | 0 |
| SWDE (lm-evaluation-harness zero-shot extraction task) | swde | long-context | 0 |
| SWE-agent | swe_agent | knowledge | 0 |
| SWE-AGI | swe_agi | coding | 0 |
| SWE-bench | swe_bench | coding | 0 |
| SWE-bench Agent | swe_bench_agent | coding | 61 |
| SWE-bench Extra | swe_bench_extra | coding | 0 |
| SWE-bench Lite | swe_bench_lite | coding | 0 |
| SWE-bench Multilingual | swe_bench_multilingual | coding | 2 |
| SWE-bench Multimodal | swe_bench_multimodal | coding | 2 |
| SWE-bench Pro | swe_bench_pro | coding | 5 |
| SWE-Bench ProMax | swe_bench_promax | coding | 0 |
| SWE-bench Science | swe_bench_science | coding | 0 |
| SWE-bench Verified | swe_bench_verified | coding | 120 |
| SWE-Bench-CL | swe_bench_cl | coding | 0 |
| swe-bench-dummy-test-dataset | swe_bench_dummy_test_dataset | knowledge | 0 |
| SWE-bench-java | swe_bench_java | coding | 0 |
| SWE-bench-Live | swe_bench_live | coding | 0 |
| SWE-Bench-Mutated | swe_bench_mutated | coding | 0 |
| SWE-Bench-Verified-O1-reasoning-high-results | swe_bench_verified_o1_reasoning_high_results | knowledge | 0 |
| SWE-EVO | swe_evo | coding | 0 |
| SWE-Explore | swe_explore | coding | 0 |
| SWE-Gym | swe_gym | coding | 0 |
| SWE-Lancer | swe_lancer | coding | 0 |
| SWE-NFI | swe_nfi | knowledge | 0 |
| SWE-PolyBench | swe_polybench | knowledge | 0 |
| SWE-rebench | swe_rebench | knowledge | 0 |
| SWE-Touch | swe_touch | knowledge | 0 |
| Swedish to German Proverbs | swedish_to_german_proverbs | translation | 0 |
| Swiss-Bench 003 | swiss_bench_003 | composite | 0 |
| Swiss-Bench SBP-002 | swiss_bench_sbp_002 | domain | 0 |
| Sycophancy Eval (inspect_evals, 'Are you sure?') | sycophancy | safety | 0 |
| Symbol Interpretation | symbol_interpretation | reasoning | 0 |
| Synthetic efficiency (HELM) | synthetic_efficiency | generation | 0 |
| Synthetic reasoning (HELM, abstract symbols) | synthetic_reasoning | reasoning | 0 |
| Synthetic Reasoning (Natural Language) | synthetic_reasoning_natural | reasoning | 0 |
| T-Eval | teval | agentic | 0 |
| TabMWP | tabmwp | math | 0 |
| Taboo | taboo | generation | 0 |
| TAC (Travel Agent Compassion) | tac | safety | 0 |
| TACO (Topics in Algorithmic COde generation) | taco | coding | 0 |
| TalkDown (BIG-bench) | talkdown | safety | 0 |
| TellMeWhy (BIG-bench) | tellmewhy | reasoning | 0 |
| Temporal Sequences (BIG-bench) | temporal_sequences | reasoning | 0 |
| Terminal-Bench | terminal_bench | agentic | 53 |
| Terminal-Bench 2.0 | terminal_bench_2 | agentic | 29 |
| Terminal-Bench 2.0 | terminal_bench_2_0 | agentic | 0 |
| Terminal-Bench 2.0 Verified | terminal_bench_2_verified | agentic | 0 |
| Terminal-Bench 3.0 | terminal_bench_3_0 | agentic | 0 |
| Terminal-Bench v2.1 | terminal_bench_v2_1 | agentic | 0 |
| Terminal-Bench-LILT | terminal_bench_lilt | agentic | 0 |
| Terminal-Bench-Science | terminal_bench_science | agentic | 0 |
| Text Navigation Game | text_navigation_game | reasoning | 0 |
| ThaiExam | thai_exam | knowledge | 0 |
| The FACTS Grounding Leaderboard | the_facts_grounding_leaderboard | generation | 0 |
| The FACTS Leaderboard | the_facts_leaderboard | composite | 0 |
| The Pile | the_pile | generation | 0 |
| The Pile (lm-eval BPB group) | pile | generation | 0 |
| TheAgentCompany | theagentcompany | agentic | 0 |
| TheoremQA | theoremqa | math | 0 |
| ThreeCB | threecb | agentic | 0 |
| TimeDial | timedial | reasoning | 0 |
| tinyBenchmarks | tinybenchmarks | composite | 0 |
| TMMLU+ | tmmluplus | knowledge | 0 |
| Topical-Chat | topical_chat | generation | 0 |
| ToxiGen | toxigen | safety | 71 |
| Tracking Shuffled Objects | tracking_shuffled_objects | reasoning | 0 |
| Training on Test Set | training_on_test_set | reasoning | 0 |
| Translation Tasks | translation | translation | 0 |
| TriviaQA | triviaqa | knowledge | 0 |
| TriviaQA RC | triviaqarc | knowledge | 0 |
| TruthfulQA | truthfulqa | safety | 50 |
| TruthfulQA-Multi | truthfulqa_multi | safety | 0 |
| TurBLiMP Core | turblimp_core | knowledge | 0 |
| TurkishMMLU | turkishmmlu | knowledge | 0 |
| TweetSentBR | tweetsentbr | domain | 0 |
| Twenty Questions | twenty_questions | agentic | 0 |
| TwitterAAE | twitter_aae | domain | 0 |
| TyDi QA | tydiqa | knowledge | 0 |
| Uganda Cultural and Cognitive Benchmark | uccb | domain | 0 |
| ULQA (Uyghur language eval group) | ulqa | composite | 0 |
| Uncheatable Eval | uncheatable_eval | generation | 0 |
| Understanding Fables | understanding_fables | reasoning | 0 |
| unit_conversion (BIG-bench) | unit_conversion | math | 0 |
| unit_interpretation (BIG-bench) | unit_interpretation | math | 0 |
| Unnatural In-Context Learning | unnatural_in_context_learning | reasoning | 0 |
| UnQover | unqover | safety | 0 |
| Unscramble | unscramble | reasoning | 0 |
| USACO | usaco | coding | 0 |
| USAMO 2026 | usamo_2026 | math | 5 |
| V*Bench | vstar_bench | multimodal | 0 |
| V-FAT | v_fat | multimodal | 0 |
| V-FiLLM | v_fillm | domain | 0 |
| Verb Tense (BIG-bench) | tense | reasoning | 0 |
| Verifiability Judgment | verifiability_judgment | knowledge | 0 |
| VGA-Bench | vga_bench | generation | 0 |
| VGA-BenchV2 | vga_benchv2 | generation | 0 |
| VIBE | vibe | knowledge | 0 |
| VIBE-Bench | vibe_bench | reasoning | 0 |
| Vibe-Eval | vibe_eval | multimodal | 0 |
| Vicuna Questions | vicuna | instruction-following | 0 |
| VimGolf Challenges (inspect_evals) | vimgolf_challenges | coding | 0 |
| VitaminC Fact Verification | vitaminc_fact_verification | knowledge | 0 |
| VQA-RAD | vqa_rad | multimodal | 0 |
| Web of Lies | web_of_lies | reasoning | 0 |
| WebQuestions | webqs | knowledge | 0 |
| What Is the Tao? | what_is_the_tao | knowledge | 0 |
| WikiBench (OpenCompass) | wikibench | knowledge | 0 |
| WikiText | wikitext | generation | 0 |
| WildBench | wildbench | human-preference | 45 |
| WinoGrande | winogrande | reasoning | 60 |
| Winogrande African Languages | winogrande_afr | knowledge | 0 |
| WinoWhy | winowhy | knowledge | 0 |
| WMDP (Weapons of Mass Destruction Proxy) | wmdp | safety | 0 |
| WMT 14 | wmt_14 | translation | 0 |
| WMT 2016 (Romanian-English, T5 prompt) | wmt2016 | translation | 0 |
| Word Problems on Sets and Graphs | word_problems_on_sets_and_graphs | reasoning | 0 |
| Word Sorting (BIG-bench) | word_sorting | reasoning | 0 |
| Word Unscrambling (BIG-bench) | word_unscrambling | reasoning | 0 |
| WorldSense | worldsense | reasoning | 0 |
| WritingBench | writingbench | generation | 0 |
| WSC273 | wsc273 | reasoning | 0 |
| XCOPA | xcopa | reasoning | 0 |
| Xiezhi | xiezhi | knowledge | 0 |
| XL-DocBench | xl_docbench | long-context | 0 |
| XL-Sum | xlsum | generation | 0 |
| XNLI (Cross-lingual Natural Language Inference) | xnli | reasoning | 0 |
| XNLIeu | xnli_eu | reasoning | 0 |
| XQuAD | xquad | knowledge | 0 |
| XSTest | xstest | safety | 0 |
| XStoryCloze | xstorycloze | reasoning | 0 |
| XSum | xsum | generation | 0 |
| XWinograd | xwinograd | reasoning | 0 |
| yes_no_black_white | yes_no_black_white | knowledge | 0 |
| ZebraLogic | zebralogic | knowledge | 0 |
| ZeroBench | zerobench | multimodal | 1 |
| ZhoBLiMP | zhoblimp | reasoning | 0 |
| τ-bench | bench_bench | agentic | 0 |
| τ-bench | tau_bench | agentic | 54 |
| τ²-bench | tau2 | agentic | 0 |
| ∞Bench (InfiniteBench) | infinitebench | long-context | 0 |
| ∞Bench: En.MC (English Multiple-Choice) | infinite_bench_en_mc | long-context | 0 |
| ∞Bench: En.QA (English Question Answering) | infinite_bench_en_qa | long-context | 0 |
| ∞Bench: En.Sum (English Summarisation) | infinite_bench_en_sum | long-context | 0 |