robot starts balancing without changing phase and farms alive bonus fix ->external force and bigger height penalty
new reward/penalty system learning phases with curriculum learning new training parameters cleanup of old code better logging while training multiple environments instead of robots (they could bumb into each other)
outdated code from previous changes on env
Training environment to make a walk model for the hexapod generated code that will be checked