new reward/penalty system learning phases with curriculum learning new training parameters cleanup of old code better logging while training multiple environments instead of robots (they could bumb into each other)
Training environment to make a walk model for the hexapod generated code that will be checked