← Leaderboard

SWE-PolyVisionBenchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering

Ranked by Best E2E·11 models·Updated September 2026

Multimodal coding agents increasingly receive screenshots and other visual artifacts when repairing software, but existing benchmarks largely measure whether visual input improves patch success without testing whether agents connect evidence across multiple images to infer a shared cause. We introduce SWE-PolyVision, a benchmark for cross-image abductive reasoning in repository-level software engineering. Each task requires an agent to align states, differences, invariants, or transitions across visual observations, formulate a repository-level hypothesis, and produce a verified patch. SWE-PolyVision contains 92 tasks from 36 open-source organizations and 405 model-readable visual inputs. Results show no consistent advantage from visual access, while controlled interventions test whether agents use the joint structure of the image set.

Tasks
92
Public tasks
48
Organizations
36
Visual inputs
405
#ModelBest E2E
1
MoonshotAI

Kimi-K3

Moonshot AI

Text 43.8%Native 39.6%Tool 50.0%

50.0%
2
Zhipu

GLM-5.3-Flash

Zhipu AI

Text 39.6%Native 39.6%Tool 25.0%

39.6%
3
OpenAI

GPT-5.6-Sol

OpenAI

Text 37.5%Native 25.0%Tool 35.4%

37.5%
4
DeepSeek

DeepSeek-V4-Flash

DeepSeek

Text 31.2%Native —Tool 18.8%

31.2%
5
AlibabaCloud

Qwen3.8-Max

Alibaba

Text 18.8%Native 29.2%Tool 20.8%

29.2%
6
Anthropic

Claude-Opus-5

Anthropic

Text 25.0%Native 20.8%Tool 27.1%

27.1%
7
Zhipu

GLM-5.3

Zhipu AI

Text 22.9%Native —Tool 20.8%

22.9%
8
DeepSeek

DeepSeek-V4-Pro

DeepSeek

Text 20.8%Native —Tool 20.8%

20.8%
9
Zhipu

GLM-5.2

Zhipu AI

Text 14.6%Native —Tool 16.7%

16.7%
10
Minimax

MiniMax-M2.7

MiniMax

Text 10.4%Native —Tool 14.6%

14.6%
11
Minimax

MiniMax-M3

MiniMax

Text 12.5%Native 12.5%Tool 10.4%

12.5%

Scores are from the current public evaluation release.

Verified repair rate by model and access modeshare of tasks solved
  • Text-onlyvisual evidence hidden
  • Native Visionimages supplied directly
  • Tool-mediated Visionimages queried on demand
  1. 43.8
    39.6
    50.0
    MoonshotAIKimi-K3
  2. 39.6
    39.6
    25.0
    ZhipuGLM-5.3-Flash
  3. 37.5
    25.0
    35.4
    OpenAIGPT-5.6-Sol
  4. 31.2
    —
    18.8
    DeepSeekDeepSeek-V4-Flash
  5. 18.8
    29.2
    20.8
    AlibabaCloudQwen3.8-Max
  6. 25.0
    20.8
    27.1
    AnthropicClaude-Opus-5
  7. 22.9
    —
    20.8
    ZhipuGLM-5.3
  8. 20.8
    —
    20.8
    DeepSeekDeepSeek-V4-Pro
  9. 14.6
    —
    16.7
    ZhipuGLM-5.2
  10. 10.4
    —
    14.6
    MinimaxMiniMax-M2.7
  11. 12.5
    12.5
    10.4
    MinimaxMiniMax-M3

One visual symptom is rarely the whole bug.

In a representative Apache ECharts issue, the agent must connect two screenshots: the legend looks correct in one state, then drifts after a chart update. The repair only becomes clear when the visual change is aligned with the repository state.

  1. 01

    Observe

    Legend alignment changes after an update.

  2. 02

    Align

    Compare the same component across states.

  3. 03

    Abduce

    Infer a shared layout or state-propagation cause.

  4. 04

    Verify

    Patch the repository and run the executable check.

One task. Three access modes.

The issue, repository, agent scaffold, budget, environment, and verifier stay fixed.

  1. 01

    Text-only

    Issue and repository, with visual evidence withheld.

  2. 02

    Native Vision

    The ordered evidence set is supplied to a multimodal model.

  3. 03

    Tool-mediated Vision

    A coding model queries a visual adapter for OCR and grounded descriptions.


← Back to Leaderboard