scrollback — Saturday May 30

Turn messy agent failure traces into reproducible evals instead of hand-written benchmarks

The pipeline starts with raw traces, attributes the failure, isolates the earliest divergence, then shrinks to a minimal state you can turn into a targeted test. That loop beats static benchmarks because it captures actual trajectory failures your agents hit in production. Once you have the mechanism you can fuzz variants, which lines up with the kind of agent orchestration work that shows up in your Claude Code and Cursor sessions.

λux @novasarc01

i’m increasingly convinced that the best agent evals will come from mining real agent failure traces. my view is that every failed trace contains a potential eval but not in its raw form. raw traces are messy, long and too specific. the research problem is to distill them into clean reproducible tests. the pipeline i’m interested in is (which i'm currently working on):

failure trace → failure attribution → earliest divergence point → minimal reproducible state → targeted eval → regression suite

this turns trace data from passive observability into an active improvement loop. like can we extract the exact decision point where the agent should have behaved differently? and can we convert that into an eval that catches the same failure class in the future? i guess this matters because most agent failures are trajectory-level failures and not just output-level failures.

personally i think this is much more realistic than relying only on hand-written benchmarks (imo they should look more like failure memory systems). hand-written evals encode what we think agents will fail on. traces encode what agents actually failed on. also once you have the mechanism, you can mutate the trace into variants. that is basically fuzzing for agents.

May 29, 2026 View on X →

https://x.com/novasarc01/status/2060403234825785840