DeepEval
confident-ai/deepeval
The LLM evaluation framework — pytest for agents.
DeepEval grades LLM outputs with 20+ metrics (hallucination, faithfulness, task completion) inside pytest — unit-testing for prompts, RAG and agents, backed by Confident AI.
Related projects
DSPy
stanfordnlp/dspy
Programming, not prompting: optimize the whole pipeline.
DSPy introduced programming models for LLM pipelines — declarative modules whose prompts and weights are optimized against metrics. Its optimizer-first philosophy reshaped how production harnesses treat prompts.
Langfuse
langfuse/langfuse
Open-source LLM engineering platform: traces, evals, prompts.
Langfuse is the open-source observability layer for LLM apps — tracing, prompt management, datasets, evals and cost tracking — deployable self-hosted or as a cloud.
Mastra
mastra-ai/mastra
The TypeScript framework for production agents.
Mastra is a full TS-native agent framework — workflows, tool calling, RAG, evals and a dev playground — that has become the default choice for JavaScript teams building agent backends.