Timeline

Microsoft Research shows VASA-1, real-time talking-face generation

Trained to render 512x512 video at up to 40 frames per second from one photo and an audio clip; Microsoft said it had no plans to release a demo or model.

  • Models & capabilities
  • Security & misuse
  • Minor

Microsoft Research published VASA-1, a model that turns a single still photo and a speech clip into a video of the subject appearing to talk, with lip movement, facial expression and head motion generated from the audio rather than filmed. The accompanying paper reported output at 512x512 resolution and up to 40 frames per second, fast enough for interactive, low-latency use, and was later accepted as an oral presentation at NeurIPS 2024.

The project page carried no public demo, code or downloadable weights, and Microsoft said it had “no plans” to release either an online demo or a product. The researchers wrote that their work targeted “positive applications” such as virtual avatars and “is not intended to create content that is used to mislead or deceive,” acknowledging in the same breath the technology’s obvious use for impersonation and disinformation.

VASA-1 arrived amid a run of realistic-face and voice-cloning systems whose creators withheld public access specifically because of misuse risk, a pattern distinct from the field’s usual practice of open publication or a paid API. Coverage at the time treated the gap between the paper’s demonstrated capability and its non-release as itself the story: a frontier lab showing what was technically possible for real-time deepfakes while declining to hand the tool to the public.