K-MetBench

K-MetBench evaluates multimodal language models for Korean meteorological expertise across visual reasoning, logic, geo-cultural comprehension and domain analysis.

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
SubcategoryKorean meteorological expert evaluation
Page statusactive
Metrictask success rate
Directionhigher_is_better
Unit%
Dataset size0
PublisherDr. Bench authors

What it measures

Expert-level Korean weather forecasting capability across charts, rationales, local context and fine-grained domain analysis.

Task format

Expert diagnostic questions grounded in national qualification exams, including chart interpretation and Korean domain reasoning.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub