Academics across six universities argue vision could be a path to general AI
Twenty-one authors proposed new benchmarks and training paradigms to pursue the idea, without presenting a working system that demonstrates it.
- Ideas & essays
- Models & capabilities
- Minor
Twenty-one researchers spanning six universities — including Oxford, Carnegie Mellon, Stanford, Princeton, NYU and Imperial College London — published a white paper arguing that general intelligence might be reached through large-scale visual experience rather than through the language-centred approach that has driven most recent frontier AI progress.
The paper is a position piece, not a report of a new system or result. Its central argument is that treating images, video and geometric information as a first-class training signal — rather than as inputs to be captioned or described in text — could let a model build the kind of grounded, predictive understanding of the physical world that language alone does not straightforwardly supply. The authors set out what they present as the open problems standing in the way: which visual input modalities matter, what benchmarks would actually measure progress toward this kind of intelligence, what training paradigms could exploit visual experience at scale, and how a vision-first system would eventually need to integrate with language rather than replace it.
The paper does not claim to have built or tested such a system, and offers no benchmark results of its own; its contribution is a research agenda and a set of proposed evaluation directions for others to take up. It joins an argument already underway in the field over whether scaling language models further is sufficient for general intelligence — including Yann LeCun’s departure from Meta to found AMI Labs around a vision-and-world-model-centred architecture — without itself taking a position on JEPA or any specific architecture.