Google's next-generation conversational video generation and editing model
Gemini Omni Flash (model ID: gemini-omni-flash-preview) is Google's next-generation video generation and editing model, designed for real-time, omnimodal interactions. It accepts text, image, and video inputs and outputs short video clips (currently up to 10 seconds at 720p). Available through Google AI Studio, the Gemini API, Vertex AI (Agent Platform), the Gemini app, and Google Flow, it is purpose-built for speed, interactive video workflows, and conversational video editing. It is a paid-tier-only model with no free tier, priced on output token consumption at approximately $0.10 per second of 720p video output.
Who it's for
developersvideo creatorsAI researchersproduct teams building video-generation features
Pricing · usage-based
checked 1d ago
Plan
Price
Includes
Standard (Paid API)
Free
~$0.10 per second of 720p video output · $17.50 per 1M output tokens (billed at 5,792 tokens/second) · $1.50 per 1M input tokens · No Batch API discount available · No free tier; paid Gemini API access required
AI-researched pricing — verify on the official site before subscribing.
Use it for
— Generating short AI videos from text or image prompts
— Conversational video editing via natural language instructions
— Integrating video generation into apps via the Gemini API
— Building interactive, low-latency video workflows in Google Flow
— Prototyping video AI features in Google AI Studio
Get the most out of it
01Use the model ID 'gemini-omni-flash-preview' when calling the API — it is not the same as other Flash models
02Budget carefully: at ~$0.10/second of output video, costs scale quickly for batch or high-volume use cases
03Note that Google uses your content to improve its products with this model, even on the paid tier — review data usage policies before sending sensitive content
04There is no Batch API discount for Gemini Omni Flash, unlike most other Gemini models, so plan your usage accordingly
05Test and prototype in Google AI Studio before integrating into production pipelines to understand output quality and token usage patterns
A powerful workflow for solo content creators to streamline their video production and enhancing their multimedia storytelling.
How the workflow runs
01gemini-omni — Use the multimodal capabilities to script and storyboard video content.
02gemini-omni-flash — Generate and edit video segments based on the script.
03mtt-transcriber-0-1-0 — Transcribe and translate the video for accessibility and wider audience reach.
Combining the text, image, audio, and video processing capabilities of these tools allows for rapid creation, editing, and accessibility of content, maximizing efficiency and audience engagement.