Google DeepMind and Harvard Propose Visual-First Path to AGI
Researchers from Google DeepMind and Harvard have proposed a new path to artificial general intelligence (AGI), shifting focus from language-based models to visual learning. A white paper titled 'Visual General Intelligence: A White Paper' argues that AI systems should learn directly from images, videos, and geometric data to understand and interact with the physical world.
The paper, which grew out of discussions at the CVPR 2026 Visual General Intelligence Workshop, brings together over 21 researchers to outline a research agenda for visual general intelligence (VGI). The authors propose using generative video models and self-supervised learning to build intelligence through visual experience. This approach could provide a foundation for AGI that language alone cannot.
The paper also explores strategies for integrating multiple modalities, recognizing that vision and language are complementary channels rather than competitors. This framework builds on previous work from Google DeepMind, including their 'Levels of AGI' taxonomy and the more recent 'From AGI to ASI' publication. The VGI white paper is deliberately open-ended, designed to spark exploration and discussion.