Capture trajectories
Keep actions, observations, tool calls, outcomes, and provenance together.
Use case / AI agents
Compare agent versions across tools, tasks, environments, and interventions without collapsing uncertainty into one score.
Who it is for
Agent behavior can change with tool availability, context, ordering, and environment state. Aggregate success rates rarely explain the first meaningful divergence.
Keep actions, observations, tool calls, outcomes, and provenance together.
Compare versions under matched tasks and conditions.
Identify when behavior separates and how strong the evidence is.
Turn consequential failures into repeatable evaluation assets.
What you leave with