reward and everything else changed again wont be the last time
new reward/penalty system learning phases with curriculum learning new training parameters cleanup of old code better logging while training multiple environments instead of robots (they could bumb into each other)
Training environment to make a walk model for the hexapod generated code that will be checked