MaCBench

A 1,153-question, 34-subset benchmark testing whether vision-language models can extract, reason about and interpret chemistry and materials-science data from images paired with text.

Also known as: MaCBench: Probing the limitations of multimodal language models for chemistry and materials research

unassessed

This page is a discovery lead. Nobody has yet assessed it against the catalogue contract, so it carries no disposition. Absence of evidence here is not evidence of staleness.
Categorydomain
Subcategorymultimodal (vision-language) chemistry and materials science
Page statusactive
Metricaccuracy: regex/pattern match against the labelled option for multiple-choice questions, and a tolerance-based numeric match (mean absolute error against a per-question relative or absolute tolerance) for open-ended quantitative questions, aggregated into one overall accuracy plus per-subset accuracy and standard error
Directionhigher_is_better
Unit%
Dataset size1153
Dataset licenceMIT
PublisherLaboratory of Organic and Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena; multi-institution collaboration

What it measures

MaCBench evaluates vision-language models on real-world chemistry and materials-science tasks that require reading an image together with accompanying text: every question needs both modalities to answer. Its 34 subsets are grouped into three categories -- data extraction (hand-drawn molecules, organic-chemistry reaction schemas, chirality, isomers, reading tables and plots, US patent figures), in-silico and lab experiments (lab safety and equipment QA, crystal-structure analysis from CIF data: atomic species, density, symmetry, volume, crystal system) and data interpretation (AFM image analysis, adsorption-isotherm and Henry-constant comparisons for metal-organic frameworks, electronic-structure and XRD pattern reading). It is a sibling of this repository's `chembench` page, built by an overlapping author group specifically to extend that text-only chemistry benchmark into the multimodal setting.

Task format

A mix of multiple-choice and open-ended (numeric or short-text) questions, each pairing an image (a diagram, plot, spectrum, micrograph or hand-drawn structure) with a text question. An optional chain-of-thought mode can be enabled; by default samples are shuffled.

Models reporting this benchmark

No model card in ModelSpec reports this benchmark yet.

Data

This page as JSON · Edit on GitHub