Google DeepMind introduces RT-2 vision-language-action model
Google DeepMind introduced RT-2, a vision-language-action model that transfers web-scale vision-language learning into robot control.

Evidence notes
- RT-2 was presented as a robotics model trained on both web-scale vision-language data and robot action data.
- The release described improved generalization on emergent robotic skills compared with RT-1 and visual-language baselines.
- The event is sourced to Google DeepMind's official RT-2 research blog.
Company context
Google DeepMind is an AI research and development company whose robotics work spans learning, control, perception, simulation and embodied AI systems.