Research
Research & blog

AutoHarnessBench
Can frontier models automate agent harness R&D? A benchmark for models that improve task-specific harnesses under a fixed budget.
August 2026Preview
Privileged On-Policy Self-Distillation
Post-training method that builds privileged repair hints from failed rollouts and distills back into the policy.
July 2026Read
Meta-Reward: Reward Modeling as Harness Optimization
Optimizing LLM-judge evaluator harnesses to build better agent reward models.
May 2026Read
Meta-Agent: Continual Learning for Agents
An open-source library that automatically improves agent harnesses from production traces.
May 2026Read
Canvas Labs
How deploying Canvas led us to build the continual-improvement layer for production AI agents.
April 2026Read