Astribot releases SmoothRL for online learning on S1
Astribot released SmoothRL, an online reinforcement-learning framework that trains during asynchronous real-world execution on the Astribot S1.

Evidence notes
- SmoothRL separates committed, executed and discarded actions so policy gradients use only actions that physically ran.
- After 250 episodes on S1, reported success reached 94% for dynamic tossing, 83% for pen capping and 90% for box opening.
- The training loop supports VR takeover and residual joystick correction while the robot continues operating.
Company context
Astribot develops AI humanoid and mobile-manipulator systems that combine robot hardware, manipulation, teleoperation data collection, developer tooling and embodied models for research and emerging real-world applications.