← Leaderboard

SWE-PrometheusMeasuring Engineering Governance Improvements in Real-World Repositories

Ranked by NGI mean·10 models·Updated September 2026

Coding agents have made rapid progress on repository-level software engineering tasks, but existing evaluations usually begin with a known issue and a predefined success signal. We introduce SWE-Prometheus, a benchmark in which an agent receives a fixed snapshot of a real repository and a general governance objective, but no defect list, base score, or oracle patch. The agent must inspect the repository, prioritize risks, implement a retrofit, and verify its claims under a limited budget. The benchmark evaluates six dimensions: Tests & CI, Code Quality Gates, Documentation & Collaboration, Structure & Maintainability, Reproducible Environment, and Dependency & Security Health, using executable base/treated evidence, hidden characterization tests, and multi-teacher scoring.

Repositories
60
Public subset
22
Models
10
Governance dimensions
6
#ModelNGI meanValid coverage
1
MoonshotAI

Kimi-K3

Moonshot AI

0.576021/22
2
Zhipu

GLM-5.3-Flash

Zhipu AI

0.576017/22
3
Anthropic

Claude-Opus-5

Anthropic

0.529318/22
4
AlibabaCloud

Qwen3.8-Max

Alibaba

0.463018/22
5
Zhipu

GLM-5.2

Zhipu AI

0.443820/22
6
DeepSeek

DeepSeek-V4-Pro

DeepSeek

0.380318/22
7
Zhipu

GLM-5.3

Zhipu AI

0.319819/22
8
OpenAI

GPT-5.6-Sol

OpenAI

0.295621/22
9
DeepSeek

DeepSeek-V4-Flash

DeepSeek

0.208322/22
10
Minimax

MiniMax-M3

MiniMax

0.056822/22

Scores are from the current public evaluation release.

NGI and valid coverage by modelpublic shared subset
  • NGI meannormalized governance improvement
  • Valid coveragevalid runs out of 22
  1. 0.5760
    95.5%
    MoonshotAIKimi-K3
  2. 0.5760
    77.3%
    ZhipuGLM-5.3-Flash
  3. 0.5293
    81.8%
    AnthropicClaude-Opus-5
  4. 0.4630
    81.8%
    AlibabaCloudQwen3.8-Max
  5. 0.4438
    90.9%
    ZhipuGLM-5.2
  6. 0.3803
    81.8%
    DeepSeekDeepSeek-V4-Pro
  7. 0.3198
    86.4%
    ZhipuGLM-5.3
  8. 0.2956
    95.5%
    OpenAIGPT-5.6-Sol
  9. 0.2083
    100.0%
    DeepSeekDeepSeek-V4-Flash
  10. 0.0568
    100.0%
    MinimaxMiniMax-M3

Useful governance begins with repository evidence.

In a representative task, the agent inspects the repository, identifies missing or ineffective engineering infrastructure, and proposes a targeted retrofit that can be checked from a clean environment.

  1. 01

    Discover

    Inspect the repository and collect baseline evidence.

  2. 02

    Prioritize

    Identify the highest-value governance gaps.

  3. 03

    Retrofit

    Implement targeted improvements without an oracle patch.

  4. 04

    Verify

    Rebuild from scratch and rerun behavior checks.

One snapshot. Three evidence stages.

The repository, agent scaffold, budget, environment, and verifier stay fixed.

  1. 01

    Baseline evidence

    Run tests, quality, documentation, environment, and dependency probes.

  2. 02

    Agent retrofit

    Let the coding agent diagnose and improve the repository under a fixed budget.

  3. 03

    Clean verification

    Rebuild the treated state and confirm valid evidence and preserved behavior.


← Back to Leaderboard