Relay
Self-improving coding agents that recover and verify.
Relay is a coding-agent harness where every run teaches the next: a queue with one worktree per task, models you already pay for (Claude, ChatGPT, Copilot or API keys), tests re-run by Relay, and lessons kept for the next lap. Free, local-first, BYOM.
In the news
EverMind open-sources Raven — 'the harness of harnesses'
EverMind AI releases Raven, a trusted, persistent, self-evolving multi-agent ecosystem built for recursive self-improvement — agents that modify their own harness under governance. The Show HN thread became a running debate on how much autonomy a harness should hand an agent over itself.
PrimeIntellect's Prime Agent: a self-improving RLM agent
PrimeIntellect open-sources Prime Agent, a recursive-language-model agent whose harness runs sub-agents as function calls inside a persistent IPython REPL. The launch post reports its harness taking Opus 5 from ~30% to 95.5% on ARC-AGI-3 — the sharpest evidence yet that harness engineering can beat model selection.
Related startups
LangChain
The de-facto standard toolkit for building LLM applications and agents.
LangChain makes the LangChain framework and LangGraph, the low-level orchestration standard for stateful, controllable agents, plus the LangSmith platform for tracing, evaluating and monitoring them in production.
Cognition
Devin, the AI software engineer, and the Windsurf IDE.
Cognition builds Devin, an autonomous AI software engineer that plans and executes engineering tasks end-to-end, and acquired the Windsurf agentic IDE in 2025. One of the defining startups of the coding-agent wave.
Braintrust
Enterprise-grade evals, data and AI gateway.
Braintrust turns prompts and agent traces into testable datasets and CI, plus an inference gateway — helping teams measure whether harness changes actually improve agent quality.