AgentOps
Agent observability, testing and session replay.
AgentOps instruments agents to record every step, cost and tool call — then lets teams replay sessions and build regression tests from real failures.
Open-source project
In the news
Related startups
LangChain
The de-facto standard toolkit for building LLM applications and agents.
LangChain makes the LangChain framework and LangGraph, the low-level orchestration standard for stateful, controllable agents, plus the LangSmith platform for tracing, evaluating and monitoring them in production.
Braintrust
Enterprise-grade evals, data and AI gateway.
Braintrust turns prompts and agent traces into testable datasets and CI, plus an inference gateway — helping teams measure whether harness changes actually improve agent quality.
Arize AI
Observability and evaluation for AI — makers of OSS Phoenix.
Arize traces and evaluates LLM and agent systems in production; their open-source Phoenix tracer is a common instrumentation choice for agent harnesses that need to be debuggable.