E
trajectoryrl caught a subtle regression in my agent's system prompt that only showed up on multi-turn conversations. single-turn evals would've missed it entirely. the multi-turn testing is the differentiator, that's where real agents actually break
