Google unveils Project Astra, a universal AI assistant prototype
A prototype, not a product: Google showed a phone-camera assistant with conversational-speed responses but gave no release date beyond 'later this year'.
- Models & capabilities
- Major
At its I/O developer conference, Google DeepMind demonstrated Project Astra, described as its “vision for the future of AI assistants”: a system that processes live video and speech continuously, builds a running timeline of what it has seen and heard, and answers questions about its surroundings through a phone camera with close to conversational latency. In a single-take video, a tester pointed a phone at objects around an office and asked questions such as what neighbourhood she was in and where she had left her glasses, with the assistant answering from what it had observed moments earlier.
Google framed the low response latency as the main engineering achievement, reached by processing video and audio as a combined event stream and caching information rather than reprocessing each frame from scratch. Demis Hassabis described it as a step toward “a future where people could have an expert AI assistant by their side, through a phone or glasses” — the reference to glasses pointing at hardware ambitions the company had not yet detailed.
The announcement was explicitly a prototype demonstration rather than a launch: Google said Astra’s capabilities would arrive in Gemini’s app and web products “later in the year” without committing to a date, and gave no detail on memory persistence, privacy handling, or how the polished demo compared with real-world latency. It landed the same week as Gemini 1.5 Flash, updated Gemini 1.5 Pro, the Veo video model and Imagen 3, part of a broad push to reposition Gemini as an agent that acts on a user’s behalf rather than a chatbot that answers questions — a framing OpenAI and others would converge on over the following year.