Google DeepMind introduces RT-2 vision-language-action model

Jul 28, 2023 · Research Publication · Google DeepMind · Foundation Models

Google DeepMind introduced RT-2, a vision-language-action model that transfers web-scale vision-language learning into robot control.

Google DeepMind company media
Company media · Google DeepMind
  • RT-2 was presented as a robotics model trained on both web-scale vision-language data and robot action data.
  • The release described improved generalization on emergent robotic skills compared with RT-1 and visual-language baselines.
  • The event is sourced to Google DeepMind's official RT-2 research blog.

Google DeepMind is an AI research and development company whose robotics work spans learning, control, perception, simulation and embodied AI systems.