Risk-aware reinforcement-learning paper evaluates adaptive quadruped locomotion on Unitree Go2

Evidence notes
- The paper trains risk-conditioned quadruped locomotion policies using CVaR-constrained policy optimization and bandit-based online adaptation.
- The method is evaluated in simulation and on a Unitree Go2 robot in previously unseen terrain.
Company context
Develops quadruped and humanoid robots, with strength in dynamic locomotion, vertically integrated hardware, and relatively low-cost commercial deployment. Unitree is one of the clearest examples of a legged robotics company moving from research visibility into real productisation and broader market distribution.