Haize Labs
Red-teaming and robustness for AI agents.
Haize Labs stress-tests agents with automated adversarial attacks — prompt injection, tool misuse, unsafe actions — and ships an SDK for continuous safety evaluation in CI.
In the news
Related startups
LangChain
The de-facto standard toolkit for building LLM applications and agents.
LangChain makes the LangChain framework and LangGraph, the low-level orchestration standard for stateful, controllable agents, plus the LangSmith platform for tracing, evaluating and monitoring them in production.
Braintrust
Enterprise-grade evals, data and AI gateway.
Braintrust turns prompts and agent traces into testable datasets and CI, plus an inference gateway — helping teams measure whether harness changes actually improve agent quality.
Arize AI
Observability and evaluation for AI — makers of OSS Phoenix.
Arize traces and evaluates LLM and agent systems in production; their open-source Phoenix tracer is a common instrumentation choice for agent harnesses that need to be debuggable.