SWE-bench
SWE-bench/SWE-bench
Can language models resolve real-world GitHub issues?
SWE-bench is the canonical evaluation harness for coding agents: real GitHub issues, a Docker-based execution harness, and instance-level scoring. The benchmark the whole coding-agent leaderboard conversation is anchored to.
In the news
Related projects
DSPy
stanfordnlp/dspy
Programming, not prompting: optimize the whole pipeline.
DSPy introduced programming models for LLM pipelines — declarative modules whose prompts and weights are optimized against metrics. Its optimizer-first philosophy reshaped how production harnesses treat prompts.
Langfuse
langfuse/langfuse
Open-source LLM engineering platform: traces, evals, prompts.
Langfuse is the open-source observability layer for LLM apps — tracing, prompt management, datasets, evals and cost tracking — deployable self-hosted or as a cloud.
Mastra
mastra-ai/mastra
The TypeScript framework for production agents.
Mastra is a full TS-native agent framework — workflows, tool calling, RAG, evals and a dev playground — that has become the default choice for JavaScript teams building agent backends.