A 1,153-question, 34-subset benchmark testing whether vision-language models can extract, reason about and interpret chemistry and materials-science data from images paired with text.
unassessed
| Category | domain |
|---|---|
| Subcategory | multimodal (vision-language) chemistry and materials science |
| Page status | active |
| Metric | accuracy: regex/pattern match against the labelled option for multiple-choice questions, and a tolerance-based numeric match (mean absolute error against a per-question relative or absolute tolerance) for open-ended quantitative questions, aggregated into one overall accuracy plus per-subset accuracy and standard error |
| Direction | higher_is_better |
| Unit | % |
| Dataset size | 1153 |
| Dataset licence | MIT |
| Publisher | Laboratory of Organic and Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena; multi-institution collaboration |
MaCBench evaluates vision-language models on real-world chemistry and materials-science tasks that require reading an image together with accompanying text: every question needs both modalities to answer. Its 34 subsets are grouped into three categories -- data extraction (hand-drawn molecules, organic-chemistry reaction schemas, chirality, isomers, reading tables and plots, US patent figures), in-silico and lab experiments (lab safety and equipment QA, crystal-structure analysis from CIF data: atomic species, density, symmetry, volume, crystal system) and data interpretation (AFM image analysis, adsorption-isotherm and Henry-constant comparisons for metal-organic frameworks, electronic-structure and XRD pattern reading). It is a sibling of this repository's `chembench` page, built by an overlapping author group specifically to extend that text-only chemistry benchmark into the multimodal setting.
A mix of multiple-choice and open-ended (numeric or short-text) questions, each pairing an image (a diagram, plot, spectrum, micrograph or hand-drawn structure) with a text question. An optional chain-of-thought mode can be enabled; by default samples are shuffled.
No model card in ModelSpec reports this benchmark yet.