ChemArgusAutomated Scientific Diagnostics for Lab-Grade Multimodal Chemistry Reasoning: A Frontier Benchmark
Chemistry evaluation has moved past multiple-choice probes toward open-ended expert problems, yet the benchmarks that define the current frontier still score answer endpoints alone. Such scoring conflates memorized facts with stepwise derivation, treats visual and textual inputs as interchangeable, and confounds long-horizon planning with stepwise execution. These blind spots prevent any diagnosis of where the reasoning chain breaks in a model. We present ChemArgus, a competition-level chemistry benchmark drawn from chemistry competition material, textbooks, and research papers, whose scoring replaces the answer endpoint with three measurement properties. Fine-grained scoring grades every subquestion point by point with partial credit, so a score lands at the level of the step rather than the outcome. Multimodal capability is priced on every problem containing a real image, paired with a controlled text description, so visual understanding is measured against text on identical tasks. Diagnostic attribution reads paired metrics that share one scoring unit and differ only in context construction, so a gap between two scores isolates a single cause, namely integration, propagation, or local reasoning. GPT-5.6-Sol scores 23.2% on unassisted whole-problem solving against 28.8% and 30.4% under the two sequential metrics, a gap measuring the integration burden. Through rubric partial credit and these cause-isolating score gaps, ChemArgus shows not only whether a model answers chemistry correctly but where its reasoning chain fails, so that every score it reports doubles as a diagnosis, a readout of the chain rather than a verdict, aiming to guide language models step by step on their path toward chemistry expertise, one corrected step at a time, each step priced by the rubric, point by point.
- Exam papers
- 23
- Problems
- 191
- Scorable subquestions
- 820
- Rubric points
- 2582
Claude Opus 5†
Anthropic
SC 59.4%OS 64.9%
Gemini 3.8 Flash§
Google DeepMind
SC 58.5%OS 60.3%
Grok 4.6§
xAI
SC 41.8%OS 45.5%
Qwen3.8-Max
Alibaba Cloud
SC 31.2%OS 34.0%
Kimi K3
Moonshot AI
SC 30.6%OS 31.8%
MiniMax M3
MiniMax
SC 27.5%OS 27.9%
GPT-5.6-Sol
OpenAI
SC 28.8%OS 30.4%
Claude Opus 5†
Anthropic
SC 61.7%OS 61.7%
Gemini 3.8 Flash§
Google DeepMind
SC 59.0%OS 59.7%
Qwen3.8-Max
Alibaba Cloud
SC 39.0%OS 41.9%
Kimi K3
Moonshot AI
SC 36.0%OS 37.5%
Grok 4.6§
xAI
SC 42.0%OS 46.0%
GPT-5.6-Sol
OpenAI
SC 35.6%OS 36.5%
MiniMax M3
MiniMax
SC 28.7%OS 29.1%
DeepSeek V4 Pro
DeepSeek
SC 25.0%OS 27.6%
GLM-5.2
Zhipu AI
SC 24.9%OS 25.0%
GLM-5.2
Zhipu AI
SC 52.2%OS 52.4%
Claude Opus 5†
Anthropic
SC 53.9%OS 54.7%
DeepSeek V4 Pro
DeepSeek
SC 36.8%OS 37.0%
Gemini 3.8 Flash§
Google DeepMind
SC 45.1%OS 49.7%
Kimi K3
Moonshot AI
SC 32.3%OS 32.4%
GPT-5.6-Sol
OpenAI
SC 33.6%OS 34.0%
Grok 4.6§
xAI
SC 29.0%OS 31.1%
Qwen3.8-Max
Alibaba Cloud
SC 29.2%OS 31.5%
MiniMax M3
MiniMax
SC 27.2%OS 27.3%
Scores are from the current evaluation release, normalized rubric percentages rounded to one decimal. Each percentage is computed on that run's own scored problem set, so denominators differ across runs.
Claude Opus 5 runs rest on a small problem prefix of 39 to 103 points after the high-scoring problem removal of scheme A (2026-09-21), so its scores are provisional until the full set completes.
Gemini 3.8 Flash and Grok 4.6 runs sit on a 12-paper 1136-point subset of the evaluation set. The reader is the Qwen3-VL image reader.
- Multimodal input
- Text only
- Text + reader
- GLM-5.2
- Gemini 3.8 Flash
- DeepSeek V4 Pro
- Qwen3.8-Max
- Kimi K3
- Grok 4.6
- GPT-5.6-Sol
- MiniMax M3
A score is a diagnosis, not a verdict.
Problems are built from first principles, normalized into one JSON unit, gated by experts and pilots, and graded by one dual-track scoring unit, so every published score traces to the exact rubric points that earned it.
- 01
Construct
Each problem is built from first principles, drawing on chemistry competition knowledge, textbooks, and research literature, and must satisfy chemical correctness, derivability from the statement, and sufficient discrimination between factual recall and genuine reasoning.
- 02
Clean
Every problem is normalized into one JSON unit, a chain of subquestions that share a header, carry rubrics and knowledge tags, and admit no extra fields, so two instantiations of the benchmark differ in content, never in structure.
- 03
Gate
Candidates pass expert review, adversarial model probing, and student pilots. Only problems that satisfy correctness, scorability, and meaningful challenge enter the final set.
- 04
Evaluate
Every run is graded by one dual-track scoring unit, deterministic chemistry-equivalence checks for formal answers and an audited independent judge for flexible ones, so every published score traces to the exact rubric points that earned it.
One chain. Three protocols.
The same subquestions, the same scoring unit, one protocol per access policy on the chain. The protocols differ only in what the model holds in context when it answers each subquestion, which is what lets a score gap between two protocols name a cause instead of confounding one.
- 01
Full problem
The model receives the complete problem, header, images, and all subquestions at once, and answers everything in one pass. Unassisted whole-problem solving; global planning and cross-subquestion integration are fully loaded.
- 02
Sequential carry
The model answers subquestions in order inside one session; each sees the model's own previous answers but never the standard results or the rubric. Sequential solving; self-consistency and error propagation become visible.
- 03
Oracle scaffolded
The model answers each subquestion in isolation, receiving the standard intermediate result of the preceding subquestion as a scaffold. Upstream correctness is granted rather than earned, and the score measures local capability on the subquestion itself.
One problem. Three input modes.
The problem, the rubric, and the scoring unit stay fixed. Only the delivery of the visual content changes, so the same content reaches a model twice, once as pixels and once as prose.
- 01
Multimodal input
The model receives real image blocks for every figure, images as data rather than decoration, never as a path or a placeholder string.
- 02
Text only
The model receives curated text descriptions with zero image blocks, written once per figure by the benchmark authors rather than generated by the model, so both inputs describe the same content from the same source.
- 03
Text + reader
A text model is paired with the Qwen3-VL image reader, which converts each image to text before the model sees the problem, pricing the delegation to an external vision module against the model's own eyes.
← Back to Leaderboard
