K-Bench 01 scores nine frontier models on 178 private first-turn scientific requests from K-Dense Web using three LLM judges and an eight-dimension 0-10 rubric.
unassessed
| Category | agentic |
|---|---|
| Subcategory | private scientific-agent evaluation on real first-turn K-Dense Web requests |
| Page status | active |
| Metric | panel mean of holistic overall (0-10); 8-anchor is scientist-acceptable with minor edits |
| Direction | higher_is_better |
| Unit | points |
| Dataset size | 178 |
| Publisher | K-Dense, Inc. |
K-Bench 01 measures what a frontier model does with a real scientific request when it has only a stock agent harness, web tools, and the files the user attached. Items are first-turn messages sampled from live K-Dense Web traffic, kept verbatim, with no reference answers. Judges score the transcript and the files left on disk, not prose alone. The skill is executed scientific work: methods, claims, artifacts, and honesty. It is not exam QA, not protocol editing, and not GPU-kernel coding.
One-shot agent run: the first user message plus attachments, inside an isolated Modal sandbox running stock pi 0.84.0 with shell, file tools, web search, fetch, and source check. No K-Dense production skills, no sub-agents, no retries, no follow-up user turns. Three blinded LLM judges then score each run against rubric v1.0.
No model card in ModelSpec reports this benchmark yet.