Research
Research & blog

Privileged Self-Distillation
Post-training method that builds privileged repair hints from failed rollouts and distills back into the policy.
July 2026Read
Meta-Reward: Reward Modeling as Harness Optimization
Optimizing LLM-judge evaluator harnesses to build better agent reward models.
May 2026Read
Meta-Agent: Continual Learning for Agents
An open-source library that automatically improves agent harnesses from production traces.
May 2026Read
Canvas Labs
How deploying Canvas led us to build the continual-improvement layer for production AI agents.
April 2026Read