Terminal-Bench
harbor-framework/terminal-bench
Hard, containerized terminal tasks, scored end to end.
Terminal-Bench is the terminal-task benchmark coding agents now cite next to SWE-bench: hard, containerized tasks scored end to end. Version 2.0 runs on the Harbor evaluation framework; the 1.0 task set lives on in the org's repos.
Related projects
DSPy
stanfordnlp/dspy
Programming, not prompting: optimize the whole pipeline.
DSPy introduced programming models for LLM pipelines — declarative modules whose prompts and weights are optimized against metrics. Its optimizer-first philosophy reshaped how production harnesses treat prompts.
Langfuse
langfuse/langfuse
Open-source LLM engineering platform: traces, evals, prompts.
Langfuse is the open-source observability layer for LLM apps — tracing, prompt management, datasets, evals and cost tracking — deployable self-hosted or as a cloud.
Mastra
mastra-ai/mastra
The TypeScript framework for production agents.
Mastra is a full TS-native agent framework — workflows, tool calling, RAG, evals and a dev playground — that has become the default choice for JavaScript teams building agent backends.