SWE-PolyVisionBenchmarking Cross-Image Abductive Reasoning for Repository-Level Software Engineering
Multimodal coding agents increasingly receive screenshots and other visual artifacts when repairing software, but existing benchmarks largely measure whether visual input improves patch success without testing whether agents connect evidence across multiple images to infer a shared cause. We introduce SWE-PolyVision, a benchmark for cross-image abductive reasoning in repository-level software engineering. Each task requires an agent to align states, differences, invariants, or transitions across visual observations, formulate a repository-level hypothesis, and produce a verified patch. SWE-PolyVision contains 92 tasks from 36 open-source organizations and 405 model-readable visual inputs. Results show no consistent advantage from visual access, while controlled interventions test whether agents use the joint structure of the image set.
- Tasks
- 92
- Public tasks
- 48
- Organizations
- 36
- Visual inputs
- 405
Kimi-K3
Moonshot AI
Text 43.8%Native 39.6%Tool 50.0%
GLM-5.3-Flash
Zhipu AI
Text 39.6%Native 39.6%Tool 25.0%
GPT-5.6-Sol
OpenAI
Text 37.5%Native 25.0%Tool 35.4%
DeepSeek-V4-Flash
DeepSeek
Text 31.2%Native —Tool 18.8%
Qwen3.8-Max
Alibaba
Text 18.8%Native 29.2%Tool 20.8%
Claude-Opus-5
Anthropic
Text 25.0%Native 20.8%Tool 27.1%
GLM-5.3
Zhipu AI
Text 22.9%Native —Tool 20.8%
DeepSeek-V4-Pro
DeepSeek
Text 20.8%Native —Tool 20.8%
GLM-5.2
Zhipu AI
Text 14.6%Native —Tool 16.7%
MiniMax-M2.7
MiniMax
Text 10.4%Native —Tool 14.6%
MiniMax-M3
MiniMax
Text 12.5%Native 12.5%Tool 10.4%
Scores are from the current public evaluation release.
- Text-only
- Native Vision
- Tool-mediated Vision
- Kimi-K3
- GLM-5.3-Flash
- GPT-5.6-Sol
- DeepSeek-V4-Flash
- Qwen3.8-Max
- Claude-Opus-5
- GLM-5.3
- DeepSeek-V4-Pro
- GLM-5.2
- MiniMax-M2.7
- MiniMax-M3
One visual symptom is rarely the whole bug.
In a representative Apache ECharts issue, the agent must connect two screenshots: the legend looks correct in one state, then drifts after a chart update. The repair only becomes clear when the visual change is aligned with the repository state.
- 01
Observe
Legend alignment changes after an update.
- 02
Align
Compare the same component across states.
- 03
Abduce
Infer a shared layout or state-propagation cause.
- 04
Verify
Patch the repository and run the executable check.
One task. Three access modes.
The issue, repository, agent scaffold, budget, environment, and verifier stay fixed.
- 01
Text-only
Issue and repository, with visual evidence withheld.
- 02
Native Vision
The ordered evidence set is supplied to a multimodal model.
- 03
Tool-mediated Vision
A coding model queries a visual adapter for OCR and grounded descriptions.
← Back to Leaderboard
