been quietly impressed with what trajectoryrl is doing. the whole premise, treating the agent instruction file itself as the thing you test and iterate against instead of just eyeballing outputs, feels obvious in hindsight but almost nobody builds tooling around it. i ran a week of our internal coding-agent prompts through it and it caught two regressions in the system prompt we would never have noticed by hand. unsexy layer, but this is the part that actually makes agents reliable.

BitFan
Public Service Atlas for Bittensor