VideoPoet by Google

VideoPoet, by Google Research, represents a significant evolution in video generation, particularly in producing large, interesting, and hi...

Last verified:

Visit VideoPoet by Google

What is VideoPoet by Google?

VideoPoet is a large language model developed by Google Research for zero-shot video generation. It converts any autoregressive language model into a high-quality video generator by using pre-trained MAGVIT V2 video tokenizer and SoundStream audio tokenizer to transform images, video, and audio clips into discrete codes compatible with text-based language models. The model synthesizes and edits videos with high temporal consistency, producing state-of-the-art video generation with a wide range of large, interesting, and high-fidelity motions.

Key features include text-to-video generation creating high-motion variable length videos from text prompts, image-to-video generation animating any input image guided by text prompts, video editing with controllable motion changes like different dance styles, zero-shot stylization applying stylistic transformations based on text prompts, video inpainting and outpainting to fill missing video parts seamlessly, and video-to-audio generation producing matching audio for videos without text guidance. The model supports square and portrait video orientations for short-form content and enables zero-shot controllable camera motions like zoom out, dolly zoom, pan left, arc shot, crane shot, and FPV drone shot.

By default, VideoPoet outputs 2-second videos but can generate longer videos by repeatedly predicting 1 second of video output given a 1-second input clip, producing videos of any duration while maintaining strong object identity preservation. The model is designed for content creators, filmmakers, video editors, and researchers exploring innovative techniques in video creation and visual storytelling. It demonstrates emergent properties like visual narrative generation where prompts can be changed over time to tell visual stories, and the ability to compose styles and effects in text-to-video generation.

VideoPoet is free to use as a research project from Google Research. The project includes a research paper and blog, with additional examples available on dedicated pages for text-to-video, image-to-video, video editing, stylization, and inpainting. The model was publicly announced on December 19, 2023, and represents a simple modeling recipe showing that language models can synthesize and edit videos effectively.

VideoPoet by Google pricing

Pricing model: Free

Free - VideoPoet is a free research project by Google Research with no paid plans. The tool is available to use without cost as part of Google's research exploration in video generation.

VideoPoet by Google pros

  • Zero-shot video generation without task-specific training data
  • Converts any autoregressive LLM into video generator
  • High-quality temporal consistency in generated videos
  • State-of-the-art video generation capabilities
  • Produces high-fidelity large motions
  • Text-to-video with variable length output
  • Image-to-video animation from any input image
  • Controllable video editing with different motion styles
  • Zero-shot stylization with prompt adherence
  • Video-to-audio generation without text guidance
  • Supports square and portrait orientations for short-form content
  • Zero-shot controllable camera motions (zoom, pan, arc, FPV drone)
  • Video inpainting and outpainting capabilities
  • Strong object identity preservation in longer clips
  • Can generate videos of any duration through extension
  • Interactive video editing with candidate selection
  • Visual narrative generation by changing prompts over time
  • Composes styles and effects easily in prompts
  • MAGVIT V2 and SoundStream tokenizer integration
  • Decoder-only transformer architecture

VideoPoet by Google cons

  • Default output limited to 2-second videos
  • Longer videos require iterative extension process
  • Short input context for video generation
  • Research project not commercially deployed product
  • No official API mentioned for integration
  • Generation quality varies with complex prompts
  • No explicit resolution specifications provided
  • Audio generation only from video input without text guidance
  • May require significant computing resources
  • No user interface mentioned, primarily research demo

Frequently asked questions about VideoPoet by Google

What is VideoPoet?

VideoPoet is a large language model for zero-shot video generation developed by Google Research. It is a simple modeling method that can convert any autoregressive language model or large language model into a high-quality video generator using pre-trained MAGVIT V2 video tokenizer and SoundStream audio tokenizer.

How long are the videos VideoPoet generates?

By default, VideoPoet outputs 2-second videos. However, the model can generate longer videos by predicting 1 second of video output given a 1-second input video clip, and this process can be repeated indefinitely to produce videos of any duration.

What types of input can VideoPoet accept?

VideoPoet accepts text, images, videos, and audio as inputs. It processes multimodal inputs including images, videos, text, and audio through its decoder-only transformer architecture, converting them into discrete codes in a unified vocabulary.

What video generation tasks does VideoPoet support?

VideoPoet supports text-to-video, text-to-image, image-to-video, video frame continuation, video inpainting, video outpainting, video stylization, and video-to-audio. These tasks can also be composed together for additional zero-shot capabilities like text-to-audio.

Can VideoPoet generate audio for videos?

Yes, VideoPoet can output audio to match an input video without using any text as guidance (video-to-audio). It uses a SoundStream audio tokenizer to transform audio clips into discrete codes and can generate matching audio for videos.

What camera motions can VideoPoet control?

VideoPoet supports zero-shot controllable camera motions including zoom out, dolly zoom, pan left, arc shot, crane shot, and FPV drone shot. These emergent properties come from pre-training and can be specified in the text prompt.

Does VideoPoet preserve object identity in longer videos?

Yes, despite the short input context, VideoPoet shows strong object identity preservation not seen in prior works, as demonstrated in longer duration clips generated through the iterative extension process.

What video orientations does VideoPoet support?

The VideoPoet model supports generating videos in square orientation or portrait orientation to tailor generations towards short-form content like social media videos.

Can VideoPoet edit video motion styles?

Yes, the VideoPoet model can edit a subject to follow different motions such as dance styles. It can process the same input clip with different prompts to produce different motion styles like robot dancing, griddy, or freestyle.

Is VideoPoet free to use?

Yes, VideoPoet is free to use as it is a research project by Google Research focused on exploring innovative techniques in video creation and editing. There are no paid plans or pricing tiers mentioned.

Categories

Use cases

Browse all AI tools on NeedAnAI