K-Bench

K-Bench 01 scores nine frontier models on 178 private first-turn scientific requests from K-Dense Web using three LLM judges and an eight-dimension 0-10 rubric.

Also known as: K-Bench 01, K-Bench01

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categoryagentic
Subcategoryprivate scientific-agent evaluation on real first-turn K-Dense Web requests
Page statusactive
Metricpanel mean of holistic overall (0-10); 8-anchor is scientist-acceptable with minor edits
Directionhigher_is_better
Unitpoints
Dataset size178
PublisherK-Dense, Inc.

What it measures

K-Bench 01 measures what a frontier model does with a real scientific request when it has only a stock agent harness, web tools, and the files the user attached. Items are first-turn messages sampled from live K-Dense Web traffic, kept verbatim, with no reference answers. Judges score the transcript and the files left on disk, not prose alone. The skill is executed scientific work: methods, claims, artifacts, and honesty. It is not exam QA, not protocol editing, and not GPU-kernel coding.

Task format

One-shot agent run: the first user message plus attachments, inside an isolated Modal sandbox running stock pi 0.84.0 with shell, file tools, web search, fetch, and source check. No K-Dense production skills, no sub-agents, no retries, no follow-up user turns. Three blinded LLM judges then score each run against rubric v1.0.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub