SWE-PrometheusMeasuring Engineering Governance Improvements in Real-World Repositories
Coding agents have made rapid progress on repository-level software engineering tasks, but existing evaluations usually begin with a known issue and a predefined success signal. We introduce SWE-Prometheus, a benchmark in which an agent receives a fixed snapshot of a real repository and a general governance objective, but no defect list, base score, or oracle patch. The agent must inspect the repository, prioritize risks, implement a retrofit, and verify its claims under a limited budget. The benchmark evaluates six dimensions: Tests & CI, Code Quality Gates, Documentation & Collaboration, Structure & Maintainability, Reproducible Environment, and Dependency & Security Health, using executable base/treated evidence, hidden characterization tests, and multi-teacher scoring.
- Repositories
- 60
- Public subset
- 22
- Models
- 10
- Governance dimensions
- 6
Kimi-K3
Moonshot AI
GLM-5.3-Flash
Zhipu AI
Claude-Opus-5
Anthropic
Qwen3.8-Max
Alibaba
GLM-5.2
Zhipu AI
DeepSeek-V4-Pro
DeepSeek
GLM-5.3
Zhipu AI
GPT-5.6-Sol
OpenAI
DeepSeek-V4-Flash
DeepSeek
MiniMax-M3
MiniMax
Scores are from the current public evaluation release.
- NGI mean
- Valid coverage
- Kimi-K3
- GLM-5.3-Flash
- Claude-Opus-5
- Qwen3.8-Max
- GLM-5.2
- DeepSeek-V4-Pro
- GLM-5.3
- GPT-5.6-Sol
- DeepSeek-V4-Flash
- MiniMax-M3
Useful governance begins with repository evidence.
In a representative task, the agent inspects the repository, identifies missing or ineffective engineering infrastructure, and proposes a targeted retrofit that can be checked from a clean environment.
- 01
Discover
Inspect the repository and collect baseline evidence.
- 02
Prioritize
Identify the highest-value governance gaps.
- 03
Retrofit
Implement targeted improvements without an oracle patch.
- 04
Verify
Rebuild from scratch and rerun behavior checks.
One snapshot. Three evidence stages.
The repository, agent scaffold, budget, environment, and verifier stay fixed.
- 01
Baseline evidence
Run tests, quality, documentation, environment, and dependency probes.
- 02
Agent retrofit
Let the coding agent diagnose and improve the repository under a fixed budget.
- 03
Clean verification
Rebuild the treated state and confirm valid evidence and preserved behavior.
← Back to Leaderboard
