The benchmark catalogue

1106 pages. 7 hold active eligibility as of 2026-09-09.

Active catalogue 7

Verified against the catalogue contract as of 2026-09-09. These are the only benchmarks presented as current.

BenchmarkIDCategoryModels reporting
AA-Briefcaseaa_briefcaseagentic0
AA-LCR (Artificial Analysis Long Context Reasoning)aa_lcrlong-context0
AutomationBench-AAautomationbench_aaagentic0
CritPt (Complex Research using Integrated Thinking — Physics Test)critptreasoning0
GDP.pdf-AAgdp_pdf_aalong-context0
GDPval-AAgdpval_aaagentic0
SciCodescicodecoding0

Unverified 7

Assessed, but missing evidence the contract requires, or not yet approved by a reviewer. Not a claim of staleness.

BenchmarkIDCategoryModels reporting
AutomationBenchautomationbenchagentic0
LiveBenchlivebenchcomposite0
OpenBookQAopenbookqareasoning0
OpenML Benchmarkopenml_benchmarkknowledge0
OpenML Benchmarksopenml_benchmarksknowledge0
SWE-Togetherswe_togetherknowledge0
Terminal-Bench v4.0terminal_bench_v4_0agentic0

Aliases 1

Alternate identifiers that resolve to a canonical benchmark.

BenchmarkIDCategoryModels reporting
OpenBookQAobqaknowledge0

Unassessed discovery leads 1091

Pages nobody has yet assessed against the contract. Reported separately from evaluated dispositions, as the spec requires.

BenchmarkIDCategoryModels reporting
AA-Omniscienceartificialanalysis_aa_omniscience_publicknowledge0
AbstentionBenchabstention_benchsafety0
Abstract Narrative Understanding (BIG-bench)abstract_narrative_understandingreasoning0
Abstraction and Reasoning Corpus (BIG-bench wrapping)abstraction_and_reasoning_corpusreasoning0
ACI-Benchaci_benchdomain0
ACLUE (Ancient Chinese Language Understanding Evaluation)aclueknowledge0
ACPBenchacp_benchreasoning0
ACPBench-Hardacp_bench_hardreasoning0
ADHD-Behaviorshc_ptbmdomain0
ADHD-MedEffectsshc_seidomain0
AdvancedIFadvancedifinstruction-following0
AdvGLUE (Adversarial GLUE)adv_gluecomposite0
Agent Memory Benchmark (AMB)agent_memory_benchmark_ambknowledge0
Agent Memory Leaderboardagent_memory_leaderboardknowledge0
agent-memory-benchagent_memory_benchknowledge0
AgentBenchagent_benchagentic0
AgentDojoagentdojoagentic0
AgentHarmagentharmsafety0
Agentic misalignmentagentic_misalignmentsafety0
AgentThreatBenchagent_threat_benchsafety0
AGIEvalagievalcomposite0
AI Agent Benchmark Leaderboardsai_agent_benchmark_leaderboardsknowledge0
AI Agent Benchmark Resultsai_agent_benchmark_resultsknowledge0
ai-agent-benchmarkai_agent_benchmarkknowledge0
AI2D (AI2 Diagrams)ai2dmultimodal39
Aider Polyglot Benchmarkaider_polyglotcoding116
AIME (American Invitational Mathematics Examination)aimemath0
AIME 2024aime_2024math0
AIME 2025aime_2025math29
AIME 2026aime_2026math0
AIR-Bench 2024air_benchsafety0
AIR-Bench 2024air_bench_2024knowledge0
AIR-BENCH Liveair_bench_liveknowledge0
AIR-Bench-Datasetair_bench_datasetknowledge0
AlGhafaalghafacomposite0
Alignment of Simplicity Priors for Turing-Complete Concept Learningsimp_turing_conceptreasoning0
AlpacaEvalalpaca_evalhuman-preference56
ALRAGEalrageknowledge0
Anachronismsanachronismsknowledge0
Analogical Similarityanalogical_similarityreasoning0
Analytic Entailmentanalytic_entailmentreasoning0
ANIMA (Animal Norms In Moral Assessment)animasafety0
ANLI (Adversarial NLI)anlireasoning0
Anthropic HH-RLHF (HELM Instruct)anthropic_hh_rlhfinstruction-following0
Anthropic Red Team (HELM)anthropic_red_teamsafety0
APE (Attempt to Persuade Eval)apesafety0
APPS (Automated Programming Progress Standard)appscoding0
ArabCulturearabculturereasoning0
Arabic Content Generation (HELM Arabic Enterprise)arabic_content_generationgeneration0
Arabic EXAMSarabic_examsknowledge0
Arabic Finance (HELM Arabic Enterprise)arabic_financedomain0
Arabic Legal (HELM Arabic Enterprise)arabic_legaldomain0
ArabicMMLUarabic_mmluknowledge0
AraDiCEaradicecomposite0
AraTrustaratrustsafety0
ARC (AI2 Reasoning Challenge)arcreasoning0
ARC Prize Public Evaluationarc_prize_public_evaluationreasoning0
ARC-AGI-2arc_agi_2reasoning1
ARC-Challengearc_challengereasoning50
ARC-Easyarc_easyreasoning0
Arena Elo (Chatbot Arena / LMArena)arena_elohuman-preference0
Arena Elo — Codingarena_elo_codinghuman-preference210
Arena Elo — Hard Promptsarena_elo_hard_promptshuman-preference103
Arena Elo — Matharena_elo_mathhuman-preference209
Arena Elo — Overall (Text)arena_elo_overallhuman-preference211
Arena Elo — Style Controlarena_elo_style_controlhuman-preference103
Arena Elo — Visionarena_elo_visionhuman-preference55
Arithmetic (GPT-3 synthetic arithmetic tasks)arithmeticmath0
Artificial Analysisartificial_analysiscomposite0
Artificial Analysis Intelligence Indexartificial_analysis_quality_indexcomposite130
Artificial Analysis Output Speedartificial_analysis_speed_indexcomposite129
ArtificialAnalysis/AA-Briefcase-Liteartificialanalysis_aa_briefcase_liteknowledge0
ArtificialAnalysis/AA-LCRartificialanalysis_aa_lcrknowledge0
ArxivRollBencharxivrollbenchreasoning0
ASCII MNISTmnist_asciimultimodal0
ASCII Word Recognition (BIG-bench)ascii_word_recognitionreasoning0
ASDivasdivmath0
AssistantBenchassistant_benchagentic0
ASTRA-benchastra_benchagentic0
ASTRA-QAastra_qaknowledge0
ATLAS (AGI-Oriented Testbed for Logical Application in Science)atlasreasoning0
Authorship Verification (BIG-bench)authorship_verificationreasoning0
Auto Categorizationauto_categorizationknowledge0
Auto Debuggingauto_debuggingcoding0
AutoBencher Capabilitiesautobencher_capabilitiescomposite0
AutoBencher Safetyautobencher_safetysafety0
bAbI (Question-Answering Tasks)babi_qareasoning0
BABILongbabilonglong-context0
Bangla BoolQAbangla_boolqareasoning0
Bangla CommonsenseQAbangla_commonsenseqareasoning0
Bangla MMLUbangla_mmluknowledge0
Bangla OpenBookQAbangla_openbookqareasoning0
Bangla PIQAbangla_piqareasoning0
BANKING77banking77domain0
BasqueBenchbasque_benchcomposite0
BasqueGLUEbasquegluecomposite0
BBQ (Bias Benchmark for QA)bbqsafety66
BBQ-Lite (Bias Benchmark for QA, BIG-bench)bbq_litesafety0
BBQ-Lite JSON (BIG-bench)bbq_lite_jsonsafety0
BEIRbeirembedding69
Belebelebelebeleknowledge0
Bench-CoEbench_coecomposite0
Bench-MFGbench_mfgreasoning0
Benchmark Contaminationbenchmark_contaminationknowledge0
benchmark-bcplusbenchmark_bcplusknowledge0
benchmark-researchbenchmark_researchknowledge0
benchmark_resultsbenchmark_resultsknowledge0
Benchmarking the Benchmarksbenchmarking_the_benchmarksreasoning0
Benchmarking the Domain Gapbenchmarking_the_domain_gapdomain0
Benchmarking the Residualbenchmarking_the_residuallong-context0
BertaQAbertaqaknowledge0
Beyond BLEUbeyond_bleuknowledge0
Beyond FLOPsbeyond_flopsknowledge0
Beyond Leaderboardsbeyond_leaderboardsknowledge0
Beyond MSEbeyond_mseknowledge0
BeyondAIMEbeyondaimemath0
BFCL (Berkeley Function-Calling Leaderboard)bfclagentic0
BHS (Basque, Hindi, Swahili syntactic evaluation)bhsreasoning0
BIG-bench (Beyond the Imitation Game Benchmark)big_benchcomposite0
BIG-Bench Extra Hard (BBEH)bbehreasoning0
BIG-Bench Hardbbhreasoning221
BigCodeBenchbigcodebenchcoding0
BIRDbirdknowledge0
BIRD-CRITICbird_criticknowledge0
BIRD-Historybird_historyknowledge0
BIRD-INTERACTbird_interactknowledge0
BIRD-SQL (HELM bird_sql / BIRD Dev)bird_sqlcoding0
BLiMP (Benchmark of Linguistic Minimal Pairs)blimpknowledge0
BLiMP-NL (Benchmark of Linguistic Minimal Pairs for Dutch)blimp_nlknowledge0
BLUEX (Brazilian Leading Universities Entrance eXams)bluexknowledge0
BMT-Statusshc_bmtdomain0
BOLD (Bias in Open-Ended Language Generation Dataset)boldsafety0
Boolean Expressions (BIG-bench)boolean_expressionsreasoning0
BoolQboolqreasoning0
Bridging Anaphora Resolution as Question Answering (BARQA-ISNotes)bridging_anaphora_resolution_barqareasoning0
BrowseCompbrowsecompagentic3
BuySideFinBenchbuysidefinbenchdomain0
C-Evalcevalknowledge0
CaBBQ (Catalan Bias Benchmark for Question Answering)cabbqsafety0
CaLM (Causal Evaluation of Language Models)calmreasoning0
CARDBiomedBenchcardbiomedbenchdomain0
CareQAcareqadomain0
CaseHOLD (Case Holdings On Legal Decisions)caseholddomain0
CatalanBenchcatalan_benchcomposite0
Causal Judgment (BIG-bench)causal_judgmentreasoning0
Cause and Effectcause_and_effectreasoning0
CDI-QAshc_cdidomain0
CHARM (Benchmarking Chinese Commonsense Reasoning of LLMs)charmreasoning0
ChartQAchartqamultimodal64
CharXivcharxivmultimodal0
CharXiv Reasoningcharxiv_reasoningmultimodal3
CharXiv Reasoning (with tool use)charxiv_reasoning_toolsmultimodal2
Checkmate in Onecheckmate_in_onereasoning0
Chem Examchem_examdomain0
ChemBenchchembenchdomain0
Chess State Trackingchess_state_trackingreasoning0
Chinese Remainder Theoremchinese_remainder_theoremmath0
Chinese SimpleQAchinese_simpleqaknowledge0
CIBenchcibenchagentic0
CIFAR-10 Classification (BIG-bench encodings)cifar10_classificationmultimodal0
CIMCQAci_mcqadomain0
CivilComments (HELM)civil_commentssafety0
CL-benchcl_bench_a_benchmark_for_context_learninglong-context0
CL-benchclbenchreasoning0
CL-bench Lifecl_bench_lifelong-context0
ClassEvalclass_evalcoding0
CLEAR (MedHELM)cleardomain0
CLEVAclevacomposite0
CLIcKclickknowledge0
ClimaQAclimaqadomain0
ClinicBenchclinicbenchdomain0
ClinicReferralshc_sequoiadomain0
ClozeTest-maxminclozetest_maxmincoding0
CLUE (Chinese Language Understanding Evaluation)cluecomposite0
CLUE: AFQMC (Ant Financial Question Matching Corpus)clue_afqmccomposite0
CLUE: C3 (free-form multiple-choice Chinese reading comprehension)clue_c3composite0
CLUE: CMNLI (Chinese Multi-Genre NLI)clue_cmnlicomposite0
CLUE: CMRC 2018 (Simplified Chinese span-extraction reading comprehension)clue_cmrccomposite0
CLUE: DRCD (Traditional Chinese span-extraction reading comprehension)clue_drcdcomposite0
CLUE: OCNLI (Original Chinese Natural Language Inference)clue_ocnlicomposite0
CMB (Comprehensive Medical Benchmark in Chinese)cmbdomain0
CMMLU (Chinese Massive Multitask Language Understanding)cmmluknowledge0
CMO fill-in-the-blankcmo_fibmath0
CMPhysBenchcmphysbenchreasoning0
CNN/DailyMail (lm-eval, See et al. v3.0.0)cnn_dailymail_abiseegeneration0
CoCo-Benchcoco_benchcoding0
CoCoNotcoconotsafety0
Code Line Descriptioncode_line_descriptioncoding0
CodeCompasscodecompasscoding0
CodeInsights Code Efficiencycodeinsights_code_efficiencycoding0
CodeInsights Correct Codecodeinsights_correct_codecoding0
CodeInsights Edge Casecodeinsights_edge_casecoding0
CodeInsights Student Codingcodeinsights_student_codingcoding0
CodeInsights Student Mistakecodeinsights_student_mistakecoding0
CodeIPI (Indirect Prompt Injection for Coding Agents)ipi_coding_agentsafety0
Codenamescodenamesreasoning0
CodeXGLUEcode_x_gluecoding0
Colorcolorknowledge0
Com2Sensecom2sensereasoning0
Common Morphemecommon_morphemeknowledge0
CommonsenseQAcommonsense_qareasoning0
CommonsenseQA-CNcommonsenseqa_cnreasoning0
CompassBench v1.1compassbench_20_v1_1composite0
CompassBench v1.1 (public)compassbench_20_v1_1_publiccomposite0
CompassBench v1.3compassbench_v1_3composite0
ComputeEvalcompute_evalcoding0
Conceptual Combinationsconceptual_combinationsreasoning0
Conlang Translationconlang_translationtranslation0
Context Definition Alignmentcontext_definition_alignmentreasoning0
Contextual Parametric Knowledge Conflictscontextual_parametric_knowledge_conflictsknowledge0
ConvFinQACalcconv_fin_qa_calcdomain0
Convince Meconvincemesafety0
COPAL-ID (Choice of Plausible Alternatives — Local Nuances, Indonesia)copal_idreasoning0
Copyright (HELM memorisation / extraction)copyrightsafety0
CoQA (Conversational Question Answering Challenge)coqareasoning0
CORE-Benchcore_benchagentic0
CORE-Benchcore_bench_computational_reproducibility_agent_benchmarkagentic0
COVIDDialog (HELM English medical dialogue)covid_dialogdomain0
Crash Blossoms (BIG-bench)crash_blossomreasoning0
CRASS (BIG-bench crass_ai)crass_aireasoning0
CrowS-Pairscrows_pairssafety0
CrowS-Pairs-CNcrowspairs_cnsafety0
CRUXEvalcruxevalcoding0
Cryobiology Spanishcryobiology_spanishknowledge0
Cryptonitecryptonitereasoning0
CS Algorithmscs_algorithmsreasoning0
CSAT-QAcsatqadomain0
CTI-REALMcti_realmagentic0
CTI-to-MITREcti_to_mitredomain0
CValuescvaluessafety0
CVE-Benchcve_benchagentic0
Cybenchcybenchagentic0
CyberGymcybergymagentic0
CyberMetriccybermetricdomain0
CyberSecEval 2cyberseceval_2safety0
CyberSecEval 3cyberseceval_3safety0
CyberSecEval 4cyberseceval_4safety0
Cycled Letters (BIG-bench)cycled_lettersreasoning0
CzechBankQAczech_bank_qacoding0
DarijaBenchdarija_benchcomposite0
DarijaHellaSwagdarijahellaswagreasoning0
DarijaMMLUdarijammluknowledge0
Dark Humor Detection (BIG-bench)dark_humor_detectionreasoning0
Data imputation (HELM)entity_data_imputationreasoning0
Data Wrangling (BIG-bench)mult_data_wranglingreasoning0
Date Understandingdate_understandingreasoning0
DecodingTrust Adversarial Demonstrationsdecodingtrust_adv_demonstrationsafety0
DecodingTrust Adversarial Robustness (AdvGLUE++)decodingtrust_adv_robustnesssafety0
DecodingTrust Fairnessdecodingtrust_fairnesssafety0
DecodingTrust Machine Ethicsdecodingtrust_machine_ethicssafety0
DecodingTrust OoD Robustnessdecodingtrust_ood_robustnesssafety0
DecodingTrust Privacydecodingtrust_privacysafety0
DecodingTrust Stereotype Biasdecodingtrust_stereotype_biassafety0
DecodingTrust Toxicity Promptsdecodingtrust_toxicity_promptssafety0
DeepMind MRCR v2deepmind_mrcr_v2long-context0
DeepSearchQAdeepsearchqaagentic1
Dingo (OpenCompass wrap)dingogeneration0
Disambiguation QAdisambiguation_qareasoning0
DischargeMe (MedHELM)dischargemedomain0
Discourse Marker Predictiondiscourse_marker_predictionknowledge0
Discrim-Evaldiscrim_evalsafety0
DISFL-QAdisfl_qareasoning0
Disinformation (HELM)disinformationsafety0
Diverse Social Biasdiverse_social_biassafety0
DocVQAdocvqamultimodal67
Dr. Benchdr_benchagentic0
Dr. DocBenchdr_docbenchmultimodal0
DROPdropreasoning0
DS-1000ds1000coding0
Dyck Language (HELM)dyck_languagereasoning0
Dyck Languages (BIG-bench)dyck_languagesreasoning0
Dynamic Countingdynamic_countingreasoning0
Earth-Silverearth_silverdomain0
ECHR Judgment Classification (HELM)echr_judgment_classificationdomain0
EESE (Ever-Evolving Science Exam)eeseknowledge0
EgyHellaSwagegyhellaswagreasoning0
EgyMMLUegymmluknowledge0
EHRSHOT (HELM ehrshot)ehrshotdomain0
EHRSQL (HELM ehr_sql / eICU)ehr_sqlcoding0
Elementary Math QAelementary_math_qamath0
Emoji Movieemoji_movieknowledge0
Emojis Emotion Predictionemojis_emotion_predictionreasoning0
Empirical Judgmentsempirical_judgmentsreasoning0
ENEM Challengeenem_challengeknowledge0
English Proverbsenglish_proverbsreasoning0
English to Russian Proverbsenglish_russian_proverbsreasoning0
ENT-Referralshc_entdomain0
Entailed Polarityentailed_polarityreasoning0
Entailed Polarity in Hindientailed_polarity_hindireasoning0
Entity matching (HELM)entity_matchingreasoning0
Epistemic Reasoningepistemic_reasoningreasoning0
EQ-Bencheq_benchreasoning0
EQ-Bench (Catalan)eq_bench_careasoning0
EQ-Bench (Spanish)eq_bench_esreasoning0
ES-MemEvales_memevallong-context0
EsBBQ (Spanish Bias Benchmark for Question Answering)esbbqsafety0
Estimating Risk of Suicidesuicide_risksafety0
ETHICS (lm-eval hendrycks_ethics)hendrycks_ethicssafety0
EusExamseus_examsknowledge0
EusProficiencyeus_proficiencyknowledge0
EusReadingeus_readingknowledge0
EusTriviaeus_triviaknowledge0
Evalita-LLMevalita_llmcomposite0
Evaluating Information Essentialityevaluating_information_essentialityreasoning0
Evo-Benchevo_benchagentic0
EvoEvalevo_evalcoding0
EWoK (Elements of World Knowledge)ewokknowledge0
EXAMS (Multilingual)exams_multilingualknowledge0
Fact-Checking (BIG-bench)fact_checkerknowledge0
Factuality of Summaryfactuality_of_summarygeneration0
Fake Alignment (FINE)fake_alignmentsafety0
Fantasy Reasoningfantasy_reasoningreasoning0
FDA (BASED information extraction)fdadomain0
Few-shot NLG (BIG-bench)few_shot_nlggeneration0
FewCLUE (Chinese Few-shot Learning Evaluation Benchmark)fewcluecomposite0
FewCLUE: BUSTM (Dialogue Short Text Matching)fewclue_bustmreasoning0
FewCLUE: CHID (Chinese Idiom Cloze Test)fewclue_chidreasoning0
FewCLUE: CLUEWSC (Winograd Schema Coreference)fewclue_cluewscreasoning0
FewCLUE: CSL (Keyword Recognition)fewclue_cslreasoning0
FewCLUE: EPRSTMT (E-commerce Sentiment Analysis)fewclue_eprstmtreasoning0
FewCLUE: OCNLI-FC (Natural Language Inference)fewclue_ocnli_fcreasoning0
FewCLUE: TNEWS (Short News Classification)fewclue_tnewsreasoning0
Figure of Speech Detection (BIG-bench)figure_of_speech_detectionreasoning0
FinanceBenchfinancebenchdomain0
FinanceIQfinanceiqdomain0
Financial PhraseBankfinancial_phrasebankdomain0
FinBenchfinbenchdomain13
FinQAfin_qadomain0
FLD (Formal Logic Deduction)fldreasoning0
FLORES English-to-Chineseflores_en_zhtranslation15
FLORES English-to-Germanflores_en_detranslation15
FLORES English-to-Japaneseflores_en_jatranslation15
FLORES English-to-Spanishflores_en_estranslation15
FLORES-200florestranslation0
ForecastBench-Simforecastbench_simreasoning0
Forecasting Subquestions (BIG-bench)forecasting_subquestionsreasoning0
Formal Fallacies Syllogisms Negation (BIG-bench)formal_fallacies_syllogisms_negationreasoning0
FORTRESSfortresssafety0
FrenchBenchfrench_benchcomposite0
Frontier-CSfrontier_cscoding0
FrontierSciencefrontiersciencedomain0
FrontierScience — Research trackfrontierscience_researchdomain1
Full-Duplex-Benchfull_duplex_benchmultimodal0
Full-Duplex-Bench-v3full_duplex_bench_v3multimodal0
GAIAgaiaagentic0
GalicianBenchgalician_benchcomposite0
Game of 24game24math0
GAOKAO-Benchgaokaobenchknowledge0
GaoKaoMATHgaokao_mathmath0
GDM In-house CTFgdm_in_house_ctfagentic0
GDM InterCode CTFgdm_intercode_ctfagentic0
GDM Self-proliferationgdm_self_proliferationagentic0
GDM Self-reasoninggdm_self_reasoningagentic0
GDM Stealthgdm_stealthagentic0
GDPvalgdpvalagentic0
GEM (BIG-bench)gemgeneration0
GenAI-Benchgenai_benchmultimodal0
Gender Inclusive Sentences Germangender_inclusive_sentences_germangeneration0
Gender Sensitivity Test - Chinesegender_sensitivity_chinesesafety0
Gender Sensitivity Test - Englishgender_sensitivity_englishsafety0
General Knowledgegeneral_knowledgeknowledge0
General365general365reasoning0
Geometric Shapesgeometric_shapesreasoning0
Global-MMLUglobal_mmluknowledge0
GLUE (General Language Understanding Evaluation benchmark)gluecomposite0
GLUE: CoLA (Corpus of Linguistic Acceptability)glue_colacomposite0
GLUE: MRPC (Microsoft Research Paraphrase Corpus)glue_mrpccomposite0
GLUE: QQP (Quora Question Pairs)glue_qqpcomposite0
Goal-step WikiHow (BIG-bench)goal_step_wikihowreasoning0
Gold Commodity News (HELM)gold_commodity_newsdomain0
GovRepcrs (OpenCompass / GovReport CRS)govrepcrslong-context0
GPQAgpqareasoning0
GPQA Diamondgpqa_diamondreasoning356
Grammar / Best ChatGPT Prompts (HELM Instruct)grammarinstruction-following0
GraphWalksgraphwalkslong-context0
GraphWalks BFS (256K-1M context)graphwalks_bfs_256k_1mlong-context3
GraphWalks Parents (256K-1M context)graphwalks_parents_256k_1mlong-context1
GRE Reading Comprehension (BIG-bench)gre_reading_comprehensionknowledge0
GreekMMLUgreekmmluknowledge0
GroundCocoagroundcocoareasoning0
GSM-Hardgsm_hardmath0
GSM8Kgsm8kmath130
GSM8K contamination (OpenCompass PPL probe)gsm8k_contaminationmath0
GUI-CCgui_ccagentic0
HAE-RAE Bench (lm-eval haerae)haeraeknowledge0
HarmBenchharm_benchsafety0
HarmBench GCG-Transfer (GCG-T)harm_bench_gcg_transfersafety0
HEAD-QAheadqadomain0
HealthBenchhealthbenchdomain0
HealthBench Hardhealthbench_harddomain1
HealthQA-BR (HELM)healthqa_brdomain0
HEARThearthuman-preference0
HEART-Benchheart_benchhuman-preference0
HellaSwaghellaswagreasoning50
HELM MELT information retrieval (Vietnamese mMARCO and mRobust)melt_irknowledge0
HELM MELT knowledge (ZaloE2E and ViMMRC)melt_knowledgeknowledge0
HELM MELT synthetic reasoning (abstract symbols)melt_synthetic_reasoningreasoning0
HELM MELT synthetic reasoning (natural language)melt_srnreasoning0
HELM Safetyhelm_safetysafety71
HELM Summarizationsummarizationgeneration0
HHH Alignment (BIG-bench)hhh_alignmentsafety0
High Low Game (BIG-bench)high_low_gamereasoning0
Hindi Question Answering (BIG-bench)hindi_question_answeringknowledge0
Hindu Knowledge (BIG-bench)hindu_knowledgeknowledge0
Hinglish Toxicity (BIG-bench)hinglish_toxicitysafety0
Histoires Moraleshistoires_moralessafety0
HMMT 2026hmmt2026math0
HospiceReferralshc_gipdomain0
HRM8Khrm8kmath0
Human Organs and Senses (BIG-bench)human_organs_sensesknowledge0
HumanEvalhumanevalcoding200
HumanEval Prohumaneval_procoding0
HumanEval+humaneval_pluscoding0
HumanEval-CNhumaneval_cncoding0
HumanEval-Infillinghumaneval_infillingcoding0
HumanEval-Xhumanevalxcoding0
Humanity's Last Examhlereasoning6
Humanity's Last Exam (with tools)hle_toolsreasoning5
Hungarian National HS Finals Exam (Mathematics)hungarian_exammath0
Hyperbaton (BIG-bench)hyperbatoninstruction-following0
Icelandic WinoGrandeicelandic_winograndereasoning0
Identify Math Theorems (BIG-bench)identify_math_theoremsmath0
Identify Odd Metaphor (BIG-bench)identify_odd_metaphorreasoning0
IFBenchifbenchinstruction-following0
IFEvalifevalinstruction-following354
IFEval_ca (Catalan IFEval)ifeval_cainstruction-following0
IFEval_es (Spanish IFEval)ifeval_esinstruction-following0
IFEvalCodeifevalcodeinstruction-following0
IMDb (Large Movie Review Dataset)imdbreasoning0
IMDb PT-BR (HELM)imdb_ptbrreasoning0
Implicatures (BIG-bench)implicaturesreasoning0
Implicit Relations (BIG-bench)implicit_relationsreasoning0
Indic Cause and Effect (BIG-bench)indic_cause_and_effectreasoning0
Indic DiarBenchindic_diarbenchmultimodal0
INDIC-DIALECTindic_dialectcomposite0
IndicXNLIindicxnlireasoning0
Inference-PPLinference_pplgeneration0
InstrumentalEvalinstrumentalevalsafety0
Intent Recognitionintent_recognitiondomain0
InteractiveQA MMLUinteractive_qa_mmluknowledge0
International Corpus of Englishicegeneration0
InternSandboxinternsandboxreasoning0
Intersect Geometryintersect_geometrymath0
Inverse IFEvalinverseifevalinstruction-following0
Inverse Scaling Prizeinverse_scalingcomposite0
IPA Natural Language Inferenceinternational_phonetic_alphabet_nlireasoning0
IPA Transliterationinternational_phonetic_alphabet_transliteratetranslation0
IPhO 2025 Theoryipho_2025_theoryreasoning1
Irony Identification (BIG-bench)irony_identificationreasoning0
IWSLT 2017 (OpenCompass English-German)iwslt2017translation0
Japanese Leaderboard (lm-evaluation-harness)japanese_leaderboardcomposite0
jfinqajfinqadomain0
Jigsaw Multilingual Toxic Comment Classification (OpenCompass)jigsawmultilingualsafety0
JSONSchemaBenchjsonschema_benchgeneration0
K-Benchk_benchagentic0
K-MetBenchk_metbenchdomain0
Kanji ASCII Art (BIG-bench)kanji_asciireasoning0
Kannada Riddles (BIG-bench)kannadareasoning0
Kaoshi (OpenCompass)kaoshiknowledge0
KBL (Korean Benchmark for Legal Language Understanding)kbldomain0
KCLE (OpenCompass)kcleknowledge0
KernelBenchkernelbenchcoding0
Key/Value Mapskey_value_mapsreasoning0
KMMLU (Korean-MMLU)kmmluknowledge0
Known Unknownsknown_unknownsknowledge0
Koala (HELM Instruct)koalainstruction-following0
KoBEST (Korean Balanced Evaluation of Significant Tasks)kobestcomposite0
KOR-Benchkorbenchreasoning0
KorMedMCQAkormedmcqadomain0
KPI-EDGARkpi_edgardomain0
L-Evallevallong-context0
LAB-Benchlab_benchdomain0
LAB-Bench: FigQAlab_bench_figqadomain2
LAB-Bench: FigQA (with tools)lab_bench_figqa_toolsdomain2
LAMBADAlambadareasoning0
LAMBADA Clozelambada_clozereasoning0
LAMBADA multilingual (OpenAI MT)lambada_multilingualreasoning0
LAMBADA multilingual (Stable LM translations)lambada_multilingual_stablelmreasoning0
Language Games (BIG-bench)language_gamestranslation0
Language Identification (BIG-bench Lite)language_identificationknowledge0
LawBenchlawbenchdomain0
LCBench2023lcbenchcoding0
LCSTS (Large-scale Chinese Short Text Summarization)lcstsgeneration0
leaderboard-datasetleaderboard_datasetknowledge0
leaderboard-detailsleaderboard_detailsknowledge0
leaderboard-requestsleaderboard_requestsknowledge0
leaderboard-resultsleaderboard_resultsknowledge0
Legal Contract Summarization (HELM)legal_contract_summarizationdomain0
Legal Opinion Sentiment Classification (HELM)legal_opinion_sentiment_classificationdomain0
Legal summarization (HELM)legal_summarizationdomain0
LegalBenchlegalbenchdomain20
LegalSupportlegal_supportdomain0
LexGLUE (Legal General Language Understanding Evaluation)lex_gluedomain0
LEXTREMElextremedomain0
LIBRA (Long Input Benchmark for Russian Analysis)libralong-context0
LingOlylingolyreasoning0
Linguistic Mappingslinguistic_mappingsreasoning0
Linguistics Puzzleslinguistics_puzzlesreasoning0
List Functionslist_functionsreasoning0
LIT-RAGBenchlit_ragbenchreasoning0
LiveCodeBenchlive_code_benchcoding144
LiveCodeBench Prolivecodebench_procoding0
LiveMathBenchlivemathbenchmath0
LiveQA (TREC-2017 Medical Task)live_qadomain0
LiveReasonBenchlivereasonbenchreasoning0
LiveStemBenchlivestembenchdomain0
LLM Compressionllm_compressiongeneration0
LLM Statsllm_statsknowledge0
LLM-QBenchllm_qbenchknowledge0
LLM-SoccerArenallm_soccerarenaknowledge0
LM-SynEval (Targeted Syntactic Evaluation of Language Models)lm_synevalknowledge0
LMentrylm_entryreasoning0
Logic Grid Puzzlelogic_grid_puzzlereasoning0
Logical Argumentslogical_argsreasoning0
Logical Deductionlogical_deductionreasoning0
Logical Fallacy Detectionlogical_fallacy_detectionreasoning0
Logical Sequencelogical_sequencereasoning0
LogiQAlogiqareasoning0
LogiQA 2.0logiqa2reasoning0
Long Context Integrationlong_context_integrationlong-context0
LongBenchlongbenchlong-context0
LongBench v2longbenchv2long-context0
LongProclongproclong-context0
LSAT (Analytical Reasoning)lsat_qareasoning0
LV-Evallvevallong-context0
M3-BENCHm3_benchhuman-preference0
M3-DuplexBenchm3_duplexbenchgeneration0
MaCBenchmacbenchdomain0
MadinahQAmadinah_qadomain0
Make Me Paymake_me_paysafety0
MakeMeSaymakemesaysafety0
MAS-Benchmas_benchagentic0
MASBenchmas_orchestraagentic0
MASK (Model Alignment between Statements and Knowledge)masksafety0
Mastermath2024v1 (OpenCompass)mastermath2024v1math0
MastermindEvalmastermindreasoning0
Matbenchmatbenchdomain0
MATH (Mathematics Aptitude Test of Heuristics)mathmath0
MATH 401math401math0
MATH-500math_500math354
MathBenchmathbenchmath0
Mathematical Induction (BIG-bench)mathematical_inductionmath0
MathQAmathqamath0
MathVistamathvistamath66
Matrix Shapes (BIG-bench)matrixshapesmath0
MBPP (Mostly Basic Python Problems)mbppcoding0
MBPP Prombpp_procoding0
MBPP+mbpp_pluscoding0
MBPP-CNmbpp_cncoding0
MC-TACOmc_tacoreasoning0
MCP-Benchmcp_benchagentic0
MedAlignmedaligninstruction-following0
MedBenchmedbenchdomain0
Medbulletsmedbulletsdomain0
MedCalc-Benchmedcalc_benchdomain0
MedConceptsQAmed_concepts_qadomain0
MedConfInfoshc_confdomain0
MedDialogmed_dialogdomain0
MEDECmedecdomain0
MedHallumedhallusafety0
MedHELM Configurablemedhelm_configurabledomain0
Medical Questions Russianmedical_questions_russiandomain0
MedicationQAmedication_qadomain0
MEDIQA (HELM)medi_qadomain0
MEDIQA 2019 QA (lm-eval)mediqa_qa2019domain0
MedMCQAmedmcqadomain4
MedQAmedqadomain51
MedText (lm-eval)medtextdomain0
MedXpertQA MMmedxpertqa_multimodaldomain1
MedXpertQA Textmedxpertqadomain0
MELT translation (HELM Vietnamese OPUS-100 and PhoMT)melt_translationtranslation0
MentalHealth (MedHELM)mental_healthdomain0
MeQSummeqsumgeneration0
metabenchmetabenchcomposite0
Metaphor Boolean (BIG-bench)metaphor_booleanreasoning0
Metaphor Understanding (BIG-bench)metaphor_understandingreasoning0
MGSM (Multilingual Grade School Math)mgsmmath45
MIMIC-BHC (MedHELM)mimic_bhcdomain0
MIMIC-III Report Summarization (lm-eval)mimic_repsumdomain0
MIMIC-IV Billing Code (MedHELM)mimiciv_billing_codedomain0
MIMIC-RRS (MedHELM)mimic_rrsdomain0
Mind2Webmind2webagentic0
Mind2Web-SCmind2web_scsafety0
MINT (Medical Incremental N-Turn Benchmark)benchmarking_multi_turn_medical_diagnosisdomain0
Minute Mysteries QAminute_mysteries_qareasoning0
MIRACLmiraclembedding90
Misconceptionsmisconceptionsknowledge0
Misconceptions (Russian)misconceptions_russianknowledge0
MLE-benchmle_benchagentic0
MLQAmlqaknowledge0
MLRC-Benchmlrc_benchagentic0
MMBenchmmbenchmultimodal0
MMIUmmiumultimodal0
MMLU (Massive Multitask Language Understanding)mmluknowledge0
MMLU clinical African languages (HELM)mmlu_clinical_afrknowledge0
MMLU-CFmmlu_cfknowledge0
MMLU-Prommlu_proknowledge354
MMLU-Pro+mmlu_pro_plusreasoning0
MMLU-ProXmmlu_proxknowledge0
MMLU-Reduxmmlu_reduxknowledge0
MMLU-SRmmlusrreasoning0
MMLU: Abstract Algebrammlu_abstract_algebraknowledge50
MMLU: Anatomymmlu_anatomyknowledge50
MMLU: Astronomymmlu_astronomyknowledge90
MMLU: Biology (subcategory)mmlu_biologyknowledge40
MMLU: Business Ethicsmmlu_business_ethicsknowledge90
MMLU: Chemistry (subcategory)mmlu_chemistryknowledge40
MMLU: Clinical Knowledgemmlu_clinical_knowledgeknowledge90
MMLU: College Biologymmlu_college_biologyknowledge50
MMLU: College Chemistrymmlu_college_chemistryknowledge50
MMLU: College Computer Sciencemmlu_college_computer_scienceknowledge50
MMLU: College Mathematicsmmlu_college_mathematicsknowledge50
MMLU: College Medicinemmlu_college_medicineknowledge50
MMLU: College Physicsmmlu_college_physicsknowledge50
MMLU: Computer Science (subcategory)mmlu_computer_scienceknowledge40
MMLU: Computer Securitymmlu_computer_securityknowledge50
MMLU: Conceptual Physicsmmlu_conceptual_physicsknowledge50
MMLU: Econometricsmmlu_econometricsknowledge50
MMLU: Electrical Engineeringmmlu_electrical_engineeringknowledge50
MMLU: Elementary Mathematicsmmlu_elementary_mathematicsknowledge50
MMLU: Formal Logicmmlu_formal_logicknowledge50
MMLU: Global Factsmmlu_global_factsknowledge50
MMLU: High School Biologymmlu_high_school_biologyknowledge50
MMLU: High School Chemistrymmlu_high_school_chemistryknowledge50
MMLU: High School Computer Sciencemmlu_high_school_computer_scienceknowledge50
MMLU: High School European Historymmlu_high_school_european_historyknowledge50
MMLU: High School Geographymmlu_high_school_geographyknowledge50
MMLU: High School Government and Politicsmmlu_high_school_government_and_politicsknowledge50
MMLU: High School Macroeconomicsmmlu_high_school_macroeconomicsknowledge50
MMLU: High School Mathematicsmmlu_high_school_mathematicsknowledge50
MMLU: High School Microeconomicsmmlu_high_school_microeconomicsknowledge50
MMLU: High School Physicsmmlu_high_school_physicsknowledge50
MMLU: High School Psychologymmlu_high_school_psychologyknowledge50
MMLU: High School Statisticsmmlu_high_school_statisticsknowledge50
MMLU: High School US Historymmlu_high_school_us_historyknowledge50
MMLU: High School World Historymmlu_high_school_world_historyknowledge50
MMLU: Human Agingmmlu_human_agingknowledge50
MMLU: Human Sexualitymmlu_human_sexualityknowledge50
MMLU: International Lawmmlu_international_lawknowledge50
MMLU: Jurisprudencemmlu_jurisprudenceknowledge90
MMLU: Logical Fallaciesmmlu_logical_fallaciesknowledge50
MMLU: Machine Learningmmlu_machine_learningknowledge50
MMLU: Managementmmlu_managementknowledge50
MMLU: Marketingmmlu_marketingknowledge50
MMLU: Medical Geneticsmmlu_medical_geneticsknowledge50
MMLU: Miscellaneousmmlu_miscellaneousknowledge50
MMLU: Moral Disputesmmlu_moral_disputesknowledge50
MMLU: Moral Scenariosmmlu_moral_scenariosknowledge50
MMLU: Nutritionmmlu_nutritionknowledge50
MMLU: Philosophymmlu_philosophyknowledge50
MMLU: Physics (unresolved key)mmlu_physicsknowledge40
MMLU: Prehistorymmlu_prehistoryknowledge50
MMLU: Professional Accountingmmlu_professional_accountingknowledge90
MMLU: Professional Lawmmlu_professional_lawknowledge90
MMLU: Professional Medicinemmlu_professional_medicineknowledge50
MMLU: Professional Psychologymmlu_professional_psychologyknowledge50
MMLU: Public Relationsmmlu_public_relationsknowledge50
MMLU: Security Studiesmmlu_security_studiesknowledge50
MMLU: Sociologymmlu_sociologyknowledge50
MMLU: US Foreign Policymmlu_us_foreign_policyknowledge50
MMLU: Virologymmlu_virologyknowledge50
MMLU: World Religionsmmlu_world_religionsknowledge50
MMLUArabic (AceGPT translated MMLU)mmluarabicknowledge0
MMMLU-litemmmlu_liteknowledge0
MMMUmmmumultimodal68
MMMU-Prommmu_promultimodal0
Model-Written Evaluationsmodel_written_evalssafety0
Modified Arithmeticmodified_arithmeticmath0
Mol-Instructions molecule-oriented tasksmolinstructions_chemdomain0
MolecularIQmolculariqdomain0
Moral Permissibilitymoral_permissibilityreasoning0
Moral Storiesmoral_storiessafety0
MORU (Moral Reasoning under Uncertainty)morusafety0
Movie Dialogue Same or Different (BIG-bench)movie_dialog_same_or_differentreasoning0
Movie Recommendation (BIG-bench)movie_recommendationknowledge0
MP-20 (OpenCompass)mp20domain0
MRAGmragdomain0
MRAG-Benchmrag_benchmultimodal0
MS MARCO (HELM passage ranking)msmarcoknowledge0
MT-Benchmt_benchhuman-preference46
MTEB (Massive Text Embedding Benchmark)mtebembedding0
MTEB Classificationmteb_classificationembedding94
MTEB Clusteringmteb_clusteringembedding94
MTEB Overall (leaderboard average)mteb_overallembedding96
MTEB Pair Classificationmteb_pair_classificationembedding93
MTEB Rerankingmteb_rerankingembedding93
MTEB Retrievalmteb_retrievalembedding96
MTEB STS (Semantic Textual Similarity)mteb_stsembedding93
MTEB Summarizationmteb_summarizationembedding93
MTEB-BRmteb_brembedding0
MTR-Benchmtr_benchreasoning0
MTR-Suitemtr_suitecomposite0
MTS-Dialog (lm-eval)mts_dialogdomain0
MTSamples Procedures (MedHELM)mtsamples_proceduresdomain0
MTSamples Replicate (MedHELM)mtsamples_replicatedomain0
MULTImultimultimodal0
MULTI-Benchmulti_benchhuman-preference0
Multi-IFmultiifinstruction-following0
Multi-SWE-benchmulti_swe_benchcoding0
MultiBLiMP 1.0multiblimpknowledge0
MultiEmo (BIG-bench)multiemoknowledge0
Multilingual MMLU (MMMLU)mmmluknowledge2
MultiPL-Emultipl_ecoding37
MultiPL-E: C#multipl_e_csharpcoding135
MultiPL-E: C++multipl_e_cppcoding37
MultiPL-E: Gomultipl_e_gocoding37
MultiPL-E: Javamultipl_e_javacoding37
MultiPL-E: JavaScriptmultipl_e_javascriptcoding37
MultiPL-E: Juliamultipl_e_juliacoding129
MultiPL-E: Kotlin (unconfirmed)multipl_e_kotlincoding135
MultiPL-E: Luamultipl_e_luacoding135
MultiPL-E: Perlmultipl_e_perlcoding129
MultiPL-E: PHPmultipl_e_phpcoding135
MultiPL-E: Pythonmultipl_e_pythoncoding37
MultiPL-E: Rmultipl_e_rcoding129
MultiPL-E: Rubymultipl_e_rubycoding135
MultiPL-E: Rustmultipl_e_rustcoding37
MultiPL-E: Scalamultipl_e_scalacoding135
MultiPL-E: Swiftmultipl_e_swiftcoding135
MultiPL-E: TypeScriptmultipl_e_typescriptcoding37
Multistep Arithmetic (BIG-bench)multistep_arithmeticmath0
Muslim-Violence Bias (BIG-bench)muslim_violence_biassafety0
MuSRmusrreasoning221
MuTualmutualreasoning0
MV-Benchmv_benchmultimodal0
MV-dVRKmv_dvrkmultimodal0
N2C2-CT Matching (HELM)n2c2_ct_matchingdomain0
NarrativeQAnarrativeqalong-context0
Natural Questions (HELM)natural_qaknowledge0
natural_instructions (BIG-bench Natural Instructions)natural_instructionsinstruction-following0
Navigate (BIG-bench)navigatereasoning0
NeedleBenchneedlebenchlong-context0
NeedleBench V2needlebench_v2long-context0
NEJMAI / nephSAP nephrology benchmarknejm_ai_benchmarkdomain0
NewsQAnewsqaknowledge0
NIAH (Inspect Evals Needle in a Haystack)niahlong-context0
Nonsense Words Grammar (BIG-bench)nonsense_words_grammarreasoning0
NorEvalnorevalcomposite0
NoteExtract (MedHELM chw_care_plan)chw_care_plandomain0
Novel Concepts (BIG-bench)novel_conceptsreasoning0
NoveltyBenchnovelty_benchgeneration0
NPHardEvalnphardevalreasoning0
NQ-CN (OpenCompass)nq_cnknowledge0
NQ-Opennq_openknowledge0
NYU-LLM-CTF/CTFTinynyu_llm_ctf_ctftinyknowledge0
NYU-LLM-CTF/NYU_CTF_Benchnyu_llm_ctf_nyu_ctf_benchknowledge0
O-NET (Inspect Evals)onetknowledge0
OAB Examsoab_examsdomain0
Object Counting (BIG-bench)object_countingreasoning0
OCRBenchocrbenchmultimodal21
Odd One Out (BIG-bench)odd_one_outreasoning0
OECD Integrityoecd_integrityknowledge0
OECD Outlookoecd_outlookknowledge0
OECD PISAoecd_pisaknowledge0
OECD Statisticsoecd_statisticsknowledge0
OJBenchojbenchcoding0
Okapi ARC multilingual (lm-eval)okapi_arc_multilingualreasoning0
Okapi multilingual HellaSwagokapi_hellaswag_multilingualreasoning0
Okapi multilingual MMLUokapi_mmlu_multilingualknowledge0
Okapi multilingual TruthfulQAokapi_truthfulqa_multilingualsafety0
OLAPH / MedLFQAolaphdomain0
OlymMATHolymmathmath0
OlympiadBencholympiadbenchmultimodal0
Omni-MATHomni_mathmath0
Open Arabic LLM Leaderboard — Complete configurationarabic_leaderboard_completecomposite0
Open Arabic LLM Leaderboard — Light configurationarabic_leaderboard_lightcomposite0
Open Assistant (HELM Instruct)open_assistantinstruction-following0
OpenAI MRCRopenai_mrcrlong-context0
OpenCompass biodata (biology-instruction)biodatadomain0
OpenCompass safety (Perspective toxicity)safetysafety0
OpenFinDataopenfindatadomain0
OpenML Explainopenml_explainknowledge0
OpenSWIopenswidomain0
Operatorsoperatorsmath0
OpinionQAopinions_qasafety0
OPT-BENCHopt_benchknowledge0
OPT-Engineopt_engineknowledge0
ORT (Out-of-Distribution Robustness Testing)llm_unlearning_should_be_form_independentsafety0
OSWorldosworldagentic3
P-MMEvalpmmevalcomposite0
Palomapalomageneration0
PaperBenchpaperbenchagentic0
Paragraph Segmentationparagraph_segmentationgeneration0
Paragraph-level Simplification of Medical Textsmed_paragraph_simplificationgeneration0
PARSINLU QAparsinlu_qaknowledge0
ParsiNLU Reading Comprehensionparsinlu_reading_comprehensionknowledge0
PAWSpawsreasoning0
PAWS-Xpaws_xreasoning0
Penguins in a Tablepenguins_in_a_tablereasoning0
PennyLane QML Benchmarkspennylane_qmldomain0
Periodic Elements (BIG-bench)periodic_elementsknowledge0
Persian Idioms (BIG-bench)persian_idiomsknowledge0
PersistBenchpersistbenchsafety0
Personality (Inspect Evals)personalitydomain0
PerspectiveGapperspectivegapagentic0
Phrase Relatedness (BIG-bench)phrase_relatednessknowledge0
PHYBenchphybenchreasoning0
Physical Intuition (BIG-bench)physical_intuitionreasoning0
PHYSICS (Benchmarking Foundation Models on University-Level Physics Problem Solving)physicsreasoning0
Physics GRE (Inflection-Benchmarks)physics_greknowledge0
physics_questions (BIG-bench)physics_questionsmath0
PI-LLMpi_llmlong-context0
Pile-10kpile_10kgeneration0
PIQApiqareasoning0
PISA-Benchpisamultimodal0
PJExampjexamknowledge0
Play Dialogue Same or Different (BIG-bench)play_dialog_same_or_differentreasoning0
PMC-Patientspmc_patientsembedding0
PolEmo 2.0polemo2domain0
Polish Sequence Labeling (BIG-bench)polish_sequence_labelingdomain0
PortugueseBenchportuguese_benchcomposite0
Pre-Flightpre_flightknowledge0
Presuppositions as NLI (BIG-bench)presuppositions_as_nlireasoning0
PRiSMprismmultimodal0
PRISM-Benchprism_benchmultimodal0
PrivacyDetectionshc_privacydomain0
ProcessBenchprocessbenchmath0
PromptBenchpromptbenchsafety0
ProofBenchproofbenchmath0
PROSTprostreasoning0
Protein Interacting Sites (BIG-bench)protein_interacting_sitesdomain0
ProteinLMBenchproteinlmbenchdomain0
ProxySendershc_proxydomain0
PubMedQApubmedqadomain4
Putnam-AXIOMputnam_axiommath0
PY150 (OpenCompass line completion)py150coding0
Python Program Synthesis (BIG-bench)program_synthesiscoding0
Python Programming Challengepython_programming_challengecoding0
QA WikiDataqa_wikidataknowledge0
QA4MREqa4mrereasoning0
qabenchqabenchknowledge0
QASPERqasperlong-context0
QASPER-cutqaspercutlong-context0
QuAC (Question Answering in Context)quacreasoning0
QuALITYqualitylong-context0
Question Selection (BIG-bench)question_selectionreasoning0
Question-Answer Creationquestion_answer_creationgeneration0
R-Bench (Reasoning Bench)r_benchreasoning0
RACEracereasoning0
RACE-H (inspect_evals)race_hreasoning0
RaceBias (HELM race_based_med)race_based_medsafety0
RAFT (Real-world Annotated Few-shot Tasks)raftdomain0
RE-Benchre_benchagentic0
Real or Fake Text (RoFT, BIG-bench)real_or_fake_textgeneration0
RealToxicityPromptsreal_toxicity_promptssafety0
RealWorldQArealworldqamultimodal8
Reasoning about Colored Objects (BIG-bench)reasoning_about_colored_objectsreasoning0
Reorderingundo_permutationreasoning0
Repeat Copy Logic (BIG-bench)repeat_copy_logicreasoning0
Rephrase (BIG-bench)rephrasegeneration0
Rhyming (BIG-bench)rhymingknowledge0
RiddleSense (BIG-bench)riddle_sensereasoning0
RO-Benchro_benchmultimodal0
RO-N3WSro_n3wsdomain0
RoleBenchrolebenchgeneration0
Roots, Optimization and Games (BIG-bench)roots_optimization_and_gamesmath0
Ruin Names (BIG-bench)ruin_namesreasoning0
RULERrulerlong-context0
S2-TOMG-Benchs2_tomg_benchdomain0
S3Evals3evallong-context0
SAD (Situational Awareness Dataset)sadsafety0
SAGEsageagentic0
Salient Translation Error Detection (BIG-bench)salient_translation_error_detectiontranslation0
scBenchscbenchagentic0
SciBenchscibenchreasoning0
ScienceQAscienceqamultimodal0
Scientific Press Release (BIG-bench)scientific_press_releasegeneration0
SciEvalscievaldomain0
SciKnowEvalsciknowevaldomain0
SciQsciqknowledge0
SciReasonerscireasonerdomain0
SciReasoner 1.5 (OpenCompass)scireasoner1_5domain0
SCORE (Systematic COnsistency and Robustness Evaluation)scorecomposite0
ScreenSpot-Proscreenspot_proagentic2
ScreenSpot-Pro (with tools)screenspot_pro_toolsagentic2
SCROLLS (Standardized CompaRison Over Long Language Sequences)scrollslong-context0
SE-Benchse_benchcoding0
SE-Evalse_evalmultimodal0
SEA-HELM (Southeast Asian Holistic Evaluation of Language Models)seahelmcomposite0
SecQAsec_qadomain0
SeedBenchseedbenchdomain0
Self Instruct (HELM)self_instructinstruction-following0
self_awareness (BIG-bench)self_awarenessreasoning0
self_evaluation_courtroom (BIG-bench)self_evaluation_courtroomreasoning0
self_evaluation_tutoring (BIG-bench Self Evaluation of Tutoring)self_evaluation_tutoringreasoning0
semantic_parsing_in_context_sparc (BIG-bench SParC)semantic_parsing_in_context_sparccoding0
semantic_parsing_spider (BIG-bench Spider)semantic_parsing_spidercoding0
Sentence Ambiguitysentence_ambiguityreasoning0
SEvenLLM (SEvenLLM-Bench)sevenllmdomain0
Similarities Test for Abstractionsimilarities_abstractionreasoning0
Simple Cooccurrence Biassimple_cooccurrence_biassafety0
Simple Ethical Questions (BIG-bench)simple_ethical_questionssafety0
Simple Text Editing (BIG-bench)simple_text_editinginstruction-following0
simple_arithmetic (BIG-bench programmatic addition template)simple_arithmeticmath0
simple_arithmetic_json (BIG-bench JSON arithmetic template)simple_arithmetic_jsonmath0
simple_arithmetic_json_multiple_choice (BIG-bench JSON MC arithmetic template)simple_arithmetic_json_multiple_choicemath0
simple_arithmetic_json_subtasks (BIG-bench nested JSON arithmetic template)simple_arithmetic_json_subtasksmath0
simple_arithmetic_multiple_targets_json (BIG-bench multi-target JSON template)simple_arithmetic_multiple_targets_jsonmath0
SimpleQAsimpleqaknowledge0
SimpleSafetyTestssimple_safety_testssafety0
SkillsBenchskillsbenchagentic0
SLM-Benchslm_benchcomposite0
SLR-Bench (group)slr_bench_groupreasoning0
SMolInstructsmolinstructdomain0
SNARKS (BIG-bench sarcasm contrast set)snarksreasoning0
Social Bias from Sentence Probability (BIG-bench)bias_from_probabilitiessafety0
Social IQasiqareasoning0
Social Supportsocial_supportsafety0
SOSBenchsosbenchsafety0
SPADE-Benchspade_benchsafety0
SpanishBenchspanish_benchcomposite0
Spelling Beespelling_beereasoning0
Spiderspidercoding0
Sports Understandingsports_understandingknowledge0
SQuADsquadreasoning0
SQuAD 2.0 (lm-evaluation-harness squadv2)squadv2reasoning0
SQuAD 2.0 (OpenCompass squad20)squad20reasoning0
SQuAD completion (Based / lm-eval)squad_completionreasoning0
squad_shifts (BIG-bench SQuADShifts)squad_shiftsreasoning0
SRBenchsrbenchmath0
STARR Patient Instructions (PatientInstruct)starr_patient_instructionsdomain0
StereoSetstereosetsafety0
Story Cloze Teststoryclozereasoning0
Strange Storiesstrange_storiesreasoning0
StrategyQAstrategyqareasoning0
StrongREJECTstrong_rejectsafety0
Subject-Verb Agreementsubject_verb_agreementreasoning0
Sudokusudokureasoning0
Sufficient Informationsufficient_informationreasoning0
SummEditssummeditsreasoning0
SummScreensummscreengeneration0
SUMOSum (HELM climate-claims summarization)sumosumdomain0
SuperCLUE-Agentsuperclue_agentagentic0
SuperCLUE-Safetysuperclue_safetysafety0
SuperGLUE (Super General Language Understanding Evaluation benchmark)super_gluecomposite0
SuperGLUE AX-b (Broad Coverage Diagnostics)superglue_ax_breasoning0
SuperGLUE AX-g (Winogender Schema Diagnostics)superglue_ax_gsafety0
SuperGLUE CB (CommitmentBank)superglue_cbreasoning0
SuperGLUE COPA (Choice of Plausible Alternatives)superglue_copareasoning0
SuperGLUE MultiRC (Multi-Sentence Reading Comprehension)superglue_multircreasoning0
SuperGLUE ReCoRD (Reading Comprehension with Commonsense Reasoning Dataset)superglue_recordreasoning0
SuperGLUE RTE (Recognizing Textual Entailment)superglue_rtereasoning0
SuperGLUE WiC (Word-in-Context)superglue_wicknowledge0
SuperGLUE WSC (Winograd Schema Challenge, SuperGLUE recast)superglue_wscreasoning0
SuperGPQAsupergpqaknowledge0
SVAMPsvampmath0
SWAG (Situations With Adversarial Generations)swagreasoning0
Swahili-English Proverbsswahili_english_proverbstranslation0
SWDE (lm-evaluation-harness zero-shot extraction task)swdelong-context0
SWE-agentswe_agentknowledge0
SWE-AGIswe_agicoding0
SWE-benchswe_benchcoding0
SWE-bench Agentswe_bench_agentcoding61
SWE-bench Extraswe_bench_extracoding0
SWE-bench Liteswe_bench_litecoding0
SWE-bench Multilingualswe_bench_multilingualcoding2
SWE-bench Multimodalswe_bench_multimodalcoding2
SWE-bench Proswe_bench_procoding5
SWE-Bench ProMaxswe_bench_promaxcoding0
SWE-bench Scienceswe_bench_sciencecoding0
SWE-bench Verifiedswe_bench_verifiedcoding120
SWE-Bench-CLswe_bench_clcoding0
swe-bench-dummy-test-datasetswe_bench_dummy_test_datasetknowledge0
SWE-bench-javaswe_bench_javacoding0
SWE-bench-Liveswe_bench_livecoding0
SWE-Bench-Mutatedswe_bench_mutatedcoding0
SWE-Bench-Verified-O1-reasoning-high-resultsswe_bench_verified_o1_reasoning_high_resultsknowledge0
SWE-EVOswe_evocoding0
SWE-Exploreswe_explorecoding0
SWE-Gymswe_gymcoding0
SWE-Lancerswe_lancercoding0
SWE-NFIswe_nfiknowledge0
SWE-PolyBenchswe_polybenchknowledge0
SWE-rebenchswe_rebenchknowledge0
SWE-Touchswe_touchknowledge0
Swedish to German Proverbsswedish_to_german_proverbstranslation0
Swiss-Bench 003swiss_bench_003composite0
Swiss-Bench SBP-002swiss_bench_sbp_002domain0
Sycophancy Eval (inspect_evals, 'Are you sure?')sycophancysafety0
Symbol Interpretationsymbol_interpretationreasoning0
Synthetic efficiency (HELM)synthetic_efficiencygeneration0
Synthetic reasoning (HELM, abstract symbols)synthetic_reasoningreasoning0
Synthetic Reasoning (Natural Language)synthetic_reasoning_naturalreasoning0
T-Evaltevalagentic0
TabMWPtabmwpmath0
Tabootaboogeneration0
TAC (Travel Agent Compassion)tacsafety0
TACO (Topics in Algorithmic COde generation)tacocoding0
TalkDown (BIG-bench)talkdownsafety0
TellMeWhy (BIG-bench)tellmewhyreasoning0
Temporal Sequences (BIG-bench)temporal_sequencesreasoning0
Terminal-Benchterminal_benchagentic53
Terminal-Bench 2.0terminal_bench_2agentic29
Terminal-Bench 2.0terminal_bench_2_0agentic0
Terminal-Bench 2.0 Verifiedterminal_bench_2_verifiedagentic0
Terminal-Bench 3.0terminal_bench_3_0agentic0
Terminal-Bench v2.1terminal_bench_v2_1agentic0
Terminal-Bench-LILTterminal_bench_liltagentic0
Terminal-Bench-Scienceterminal_bench_scienceagentic0
Text Navigation Gametext_navigation_gamereasoning0
ThaiExamthai_examknowledge0
The FACTS Grounding Leaderboardthe_facts_grounding_leaderboardgeneration0
The FACTS Leaderboardthe_facts_leaderboardcomposite0
The Pilethe_pilegeneration0
The Pile (lm-eval BPB group)pilegeneration0
TheAgentCompanytheagentcompanyagentic0
TheoremQAtheoremqamath0
ThreeCBthreecbagentic0
TimeDialtimedialreasoning0
tinyBenchmarkstinybenchmarkscomposite0
TMMLU+tmmluplusknowledge0
Topical-Chattopical_chatgeneration0
ToxiGentoxigensafety71
Tracking Shuffled Objectstracking_shuffled_objectsreasoning0
Training on Test Settraining_on_test_setreasoning0
Translation Taskstranslationtranslation0
TriviaQAtriviaqaknowledge0
TriviaQA RCtriviaqarcknowledge0
TruthfulQAtruthfulqasafety50
TruthfulQA-Multitruthfulqa_multisafety0
TurBLiMP Coreturblimp_coreknowledge0
TurkishMMLUturkishmmluknowledge0
TweetSentBRtweetsentbrdomain0
Twenty Questionstwenty_questionsagentic0
TwitterAAEtwitter_aaedomain0
TyDi QAtydiqaknowledge0
Uganda Cultural and Cognitive Benchmarkuccbdomain0
ULQA (Uyghur language eval group)ulqacomposite0
Uncheatable Evaluncheatable_evalgeneration0
Understanding Fablesunderstanding_fablesreasoning0
unit_conversion (BIG-bench)unit_conversionmath0
unit_interpretation (BIG-bench)unit_interpretationmath0
Unnatural In-Context Learningunnatural_in_context_learningreasoning0
UnQoverunqoversafety0
Unscrambleunscramblereasoning0
USACOusacocoding0
USAMO 2026usamo_2026math5
V*Benchvstar_benchmultimodal0
V-FATv_fatmultimodal0
V-FiLLMv_fillmdomain0
Verb Tense (BIG-bench)tensereasoning0
Verifiability Judgmentverifiability_judgmentknowledge0
VGA-Benchvga_benchgeneration0
VGA-BenchV2vga_benchv2generation0
VIBEvibeknowledge0
VIBE-Benchvibe_benchreasoning0
Vibe-Evalvibe_evalmultimodal0
Vicuna Questionsvicunainstruction-following0
VimGolf Challenges (inspect_evals)vimgolf_challengescoding0
VitaminC Fact Verificationvitaminc_fact_verificationknowledge0
VQA-RADvqa_radmultimodal0
Web of Liesweb_of_liesreasoning0
WebQuestionswebqsknowledge0
What Is the Tao?what_is_the_taoknowledge0
WikiBench (OpenCompass)wikibenchknowledge0
WikiTextwikitextgeneration0
WildBenchwildbenchhuman-preference45
WinoGrandewinograndereasoning60
Winogrande African Languageswinogrande_afrknowledge0
WinoWhywinowhyknowledge0
WMDP (Weapons of Mass Destruction Proxy)wmdpsafety0
WMT 14wmt_14translation0
WMT 2016 (Romanian-English, T5 prompt)wmt2016translation0
Word Problems on Sets and Graphsword_problems_on_sets_and_graphsreasoning0
Word Sorting (BIG-bench)word_sortingreasoning0
Word Unscrambling (BIG-bench)word_unscramblingreasoning0
WorldSenseworldsensereasoning0
WritingBenchwritingbenchgeneration0
WSC273wsc273reasoning0
XCOPAxcopareasoning0
Xiezhixiezhiknowledge0
XL-DocBenchxl_docbenchlong-context0
XL-Sumxlsumgeneration0
XNLI (Cross-lingual Natural Language Inference)xnlireasoning0
XNLIeuxnli_eureasoning0
XQuADxquadknowledge0
XSTestxstestsafety0
XStoryClozexstoryclozereasoning0
XSumxsumgeneration0
XWinogradxwinogradreasoning0
yes_no_black_whiteyes_no_black_whiteknowledge0
ZebraLogiczebralogicknowledge0
ZeroBenchzerobenchmultimodal1
ZhoBLiMPzhoblimpreasoning0
τ-benchbench_benchagentic0
τ-benchtau_benchagentic54
τ²-benchtau2agentic0
∞Bench (InfiniteBench)infinitebenchlong-context0
∞Bench: En.MC (English Multiple-Choice)infinite_bench_en_mclong-context0
∞Bench: En.QA (English Question Answering)infinite_bench_en_qalong-context0
∞Bench: En.Sum (English Summarisation)infinite_bench_en_sumlong-context0