new reward/penalty system
learning phases with curriculum learning
new training parameters
cleanup of old code
better logging while training
multiple environments instead of robots (they could bumb into each other)
Robot into its own Class instead of lose Global Variables that cause circular imports
StateClass usage instead of the old RobotState.py
New Input Class for Controller and randome intputs