What it does
A Claude Code subagent for designing, training and deploying reinforcement learning agents. It reviews the environment, reward structure and agent design, then helps with algorithm choice, training stability, evaluation and safety constraints.
Use cases
- 01Define the state and action space for a new environment
- 02Review a reward function for unintended behaviour
- 03Pick a policy optimization algorithm for a task