At today's Google I/O 2026 developer conference, the much-leaked multimodal model Gemini Omni finally made its official debut. Designed for video generation and editing, it represents Google's consolidation of its top-tier AI media generation systems, promising to reshape the video creation landscape.
Three core highlights: from generation to conversational editing
According to the official demo, Gemini Omni offers full-spectrum generation and remixing: users can start from text, images, audio, existing video, or even hand-drawn sketches to produce any content. The standout feature is conversational editing—users modify videos via natural language in a chat interface, e.g., "change the camera angle," "switch to golden hour lighting," or "replace that object." The AI iterates while preserving character consistency and physical laws. Finally, high-fidelity physics simulation was demonstrated in complex scenes like a professor writing on a blackboard or two people eating pasta, showing strong text-to-video coherence.
Rollout schedule: paid access today, free on YouTube Shorts this week
Google outlined a phased release: effective immediately, subscribers of Google AI Plus, Pro, and Ultra can access Gemini Omni Flash via Gemini App and Flow by Google. This week, the feature will be free integrated into YouTube Shorts and YouTube Create App. Later, an API will open to developers and enterprises.
Industry watchers note that Gemini Omni likely extends Google's strongest video generation model Veo (e.g., Veo 3.1), but with a unified multimodal experience blending text, image, video, and audio. Generated videos include safety watermarks and are subject to strict content restrictions.

