Google's RT-2 Model Marks 'GPT-3 Moment' for Robotics
Google's release of RT-2 in July 2023 marked the 'GPT-3 moment' for robotics, according to researchers. This model, a multimodal large language model (LLM), was trained to directly generate robot actions and showed significant improvements in generalization over objects, scenes, and instructions.
The RT-2 team reported that their model exhibited emergent capabilities inherited from web-scale vision-language pretraining. For example, when prompted to 'move coke can to Taylor Swift,' the robot successfully grabbed the can and moved it toward Swift's photo, even though the model had never seen Taylor Swift in its data.
This breakthrough kicked off a robotics boom that has been underway since 2023. Big companies in both the US and China have invested heavily in robotics, and numerous startups have raised hundreds of millions of dollars in venture capital. The models powering most of these robots are based on the basic architecture Google pioneered with RT-2.
Karol Hausman, a member of the RT-2 team, left Google to become CEO of Physical Intelligence, a startup focused on developing physical intelligence. He stated that solving physical intelligence requires an organization dedicated solely to this purpose and not just as a secondary priority in another company.