Braintrust
Enterprise-grade evals, data and AI gateway.
Braintrust turns prompts and agent traces into testable datasets and CI, plus an inference gateway — helping teams measure whether harness changes actually improve agent quality.
In the news
HarnessTax: harness choice swings agent cost more than model choice
Arena.ai's HarnessTax study tests 21 model–harness combinations across Claude Code, Codex CLI and Pi, finding harness choice can significantly change cost even at similar task success rates. Weeks earlier, SWE-bench Pro analysis showed swapping harnesses moved pass@1 more than model upgrades (23%→52% on GLM-5.2) — and that harness rankings barely transfer across models.
ClickHouse acquires Langfuse
ClickHouse acquires the open-source LLM observability platform Langfuse as part of its $400M Series D at a $15B valuation — the first big consolidation in agent observability. The open-source core remains maintained; the deal signals that agent tracing is becoming a database-scale problem.
Related startups
LangChain
The de-facto standard toolkit for building LLM applications and agents.
LangChain makes the LangChain framework and LangGraph, the low-level orchestration standard for stateful, controllable agents, plus the LangSmith platform for tracing, evaluating and monitoring them in production.
Arize AI
Observability and evaluation for AI — makers of OSS Phoenix.
Arize traces and evaluates LLM and agent systems in production; their open-source Phoenix tracer is a common instrumentation choice for agent harnesses that need to be debuggable.
Galileo
Evaluation platform for LLMs and agentic systems.
Galileo (makers of the Luna evaluation models) scores agent runs for hallucination, tool-use errors and task completion — automated quality gates for harness engineering.