HarnessHub
← Open Source
Open-source project

SWE-bench

SWE-bench/SWE-bench

Can language models resolve real-world GitHub issues?

SWE-bench is the canonical evaluation harness for coding agents: real GitHub issues, a Docker-based execution harness, and instance-level scoring. The benchmark the whole coding-agent leaderboard conversation is anchored to.

In the news

Related projects