← Leaderboard

ChemArgusAutomated Scientific Diagnostics for Lab-Grade Multimodal Chemistry Reasoning: A Frontier Benchmark

Ranked by full problem·25 runs·9 models·Updated September 2026

Chemistry evaluation has moved past multiple-choice probes toward open-ended expert problems, yet the benchmarks that define the current frontier still score answer endpoints alone. Such scoring conflates memorized facts with stepwise derivation, treats visual and textual inputs as interchangeable, and confounds long-horizon planning with stepwise execution. These blind spots prevent any diagnosis of where the reasoning chain breaks in a model. We present ChemArgus, a competition-level chemistry benchmark drawn from chemistry competition material, textbooks, and research papers, whose scoring replaces the answer endpoint with three measurement properties. Fine-grained scoring grades every subquestion point by point with partial credit, so a score lands at the level of the step rather than the outcome. Multimodal capability is priced on every problem containing a real image, paired with a controlled text description, so visual understanding is measured against text on identical tasks. Diagnostic attribution reads paired metrics that share one scoring unit and differ only in context construction, so a gap between two scores isolates a single cause, namely integration, propagation, or local reasoning. GPT-5.6-Sol scores 23.2% on unassisted whole-problem solving against 28.8% and 30.4% under the two sequential metrics, a gap measuring the integration burden. Through rubric partial credit and these cause-isolating score gaps, ChemArgus shows not only whether a model answers chemistry correctly but where its reasoning chain fails, so that every score it reports doubles as a diagnosis, a readout of the chain rather than a verdict, aiming to guide language models step by step on their path toward chemistry expertise, one corrected step at a time, each step priced by the rubric, point by point.

PDFGitHubSoonHugging FaceSoon
Exam papers
23
Problems
191
Scorable subquestions
820
Rubric points
2582
#ModelFP
Multimodal input7 runs
1

Claude Opus 5†

Anthropic

SC 59.4%OS 64.9%

53.8%
2

Gemini 3.8 Flash§

Google DeepMind

SC 58.5%OS 60.3%

48.5%
3

Grok 4.6§

xAI

SC 41.8%OS 45.5%

33.8%
4

Qwen3.8-Max

Alibaba Cloud

SC 31.2%OS 34.0%

29.3%
5

Kimi K3

Moonshot AI

SC 30.6%OS 31.8%

27.5%
6

MiniMax M3

MiniMax

SC 27.5%OS 27.9%

27.3%
7

GPT-5.6-Sol

OpenAI

SC 28.8%OS 30.4%

23.2%
Text only9 runs
1

Claude Opus 5†

Anthropic

SC 61.7%OS 61.7%

61.5%
2

Gemini 3.8 Flash§

Google DeepMind

SC 59.0%OS 59.7%

46.5%
3

Qwen3.8-Max

Alibaba Cloud

SC 39.0%OS 41.9%

36.0%
4

Kimi K3

Moonshot AI

SC 36.0%OS 37.5%

34.5%
5

Grok 4.6§

xAI

SC 42.0%OS 46.0%

33.4%
6

GPT-5.6-Sol

OpenAI

SC 35.6%OS 36.5%

29.8%
7

MiniMax M3

MiniMax

SC 28.7%OS 29.1%

26.2%
8

DeepSeek V4 Pro

DeepSeek

SC 25.0%OS 27.6%

25.0%
9

GLM-5.2

Zhipu AI

SC 24.9%OS 25.0%

23.5%
Text + reader9 runs
1

GLM-5.2

Zhipu AI

SC 52.2%OS 52.4%

52.0%
2

Claude Opus 5†

Anthropic

SC 53.9%OS 54.7%

46.2%
3

DeepSeek V4 Pro

DeepSeek

SC 36.8%OS 37.0%

36.7%
4

Gemini 3.8 Flash§

Google DeepMind

SC 45.1%OS 49.7%

33.8%
5

Kimi K3

Moonshot AI

SC 32.3%OS 32.4%

32.3%
6

GPT-5.6-Sol

OpenAI

SC 33.6%OS 34.0%

31.1%
7

Grok 4.6§

xAI

SC 29.0%OS 31.1%

28.6%
8

Qwen3.8-Max

Alibaba Cloud

SC 29.2%OS 31.5%

27.9%
9

MiniMax M3

MiniMax

SC 27.2%OS 27.3%

27.2%

Scores are from the current evaluation release, normalized rubric percentages rounded to one decimal. Each percentage is computed on that run's own scored problem set, so denominators differ across runs.

Claude Opus 5 runs rest on a small problem prefix of 39 to 103 points after the high-scoring problem removal of scheme A (2026-09-21), so its scores are provisional until the full set completes.

Gemini 3.8 Flash and Grok 4.6 runs sit on a 12-paper 1136-point subset of the evaluation set. The reader is the Qwen3-VL image reader.

Rubric score by model and input modefull problem, % of each run’s own scored problem set
  • Multimodal inputReal image blocks
  • Text onlyCurated descriptions, zero image blocks
  • Text + readerQwen3-VL converts images to text
  1. 23.5
    52.0
    GLM-5.2
  2. 48.5
    46.5
    33.8
    Gemini 3.8 Flash
  3. 25.0
    36.7
    DeepSeek V4 Pro
  4. 29.3
    36.0
    27.9
    Qwen3.8-Max
  5. 27.5
    34.5
    32.3
    Kimi K3
  6. 33.8
    33.4
    28.6
    Grok 4.6
  7. 23.2
    29.8
    31.1
    GPT-5.6-Sol
  8. 27.3
    26.2
    27.2
    MiniMax M3

Claude Opus 5's provisional prefix results appear in the table only.

A score is a diagnosis, not a verdict.

Problems are built from first principles, normalized into one JSON unit, gated by experts and pilots, and graded by one dual-track scoring unit, so every published score traces to the exact rubric points that earned it.

  1. 01

    Construct

    Each problem is built from first principles, drawing on chemistry competition knowledge, textbooks, and research literature, and must satisfy chemical correctness, derivability from the statement, and sufficient discrimination between factual recall and genuine reasoning.

  2. 02

    Clean

    Every problem is normalized into one JSON unit, a chain of subquestions that share a header, carry rubrics and knowledge tags, and admit no extra fields, so two instantiations of the benchmark differ in content, never in structure.

  3. 03

    Gate

    Candidates pass expert review, adversarial model probing, and student pilots. Only problems that satisfy correctness, scorability, and meaningful challenge enter the final set.

  4. 04

    Evaluate

    Every run is graded by one dual-track scoring unit, deterministic chemistry-equivalence checks for formal answers and an audited independent judge for flexible ones, so every published score traces to the exact rubric points that earned it.

One chain. Three protocols.

The same subquestions, the same scoring unit, one protocol per access policy on the chain. The protocols differ only in what the model holds in context when it answers each subquestion, which is what lets a score gap between two protocols name a cause instead of confounding one.

  1. 01

    Full problem

    The model receives the complete problem, header, images, and all subquestions at once, and answers everything in one pass. Unassisted whole-problem solving; global planning and cross-subquestion integration are fully loaded.

  2. 02

    Sequential carry

    The model answers subquestions in order inside one session; each sees the model's own previous answers but never the standard results or the rubric. Sequential solving; self-consistency and error propagation become visible.

  3. 03

    Oracle scaffolded

    The model answers each subquestion in isolation, receiving the standard intermediate result of the preceding subquestion as a scaffold. Upstream correctness is granted rather than earned, and the score measures local capability on the subquestion itself.

One problem. Three input modes.

The problem, the rubric, and the scoring unit stay fixed. Only the delivery of the visual content changes, so the same content reaches a model twice, once as pixels and once as prose.

  1. 01

    Multimodal input

    The model receives real image blocks for every figure, images as data rather than decoration, never as a path or a placeholder string.

  2. 02

    Text only

    The model receives curated text descriptions with zero image blocks, written once per figure by the benchmark authors rather than generated by the model, so both inputs describe the same content from the same source.

  3. 03

    Text + reader

    A text model is paired with the Qwen3-VL image reader, which converts each image to text before the model sees the problem, pricing the delegation to an external vision module against the model's own eyes.


← Back to Leaderboard