Design RL environments, tasksets, and eval loops without rebuilding the setup stack from scratch. Models are getting good enough that the painful part of RL is no longer only training.