Gemini Omni Flash
ยอดนิยมGemini Omni Flash is the Google any-input-to-video model, turning text, image, audio, or video prompts into generated and edited clips with native audio, keyframe interpolation, and clip extension, generally available on the paid tier of the Gemini API.
ภาพรวม
Gemini Omni Flash is the next-generation video generation and editing model from Google DeepMind, generally available to developers on the paid tier of the Gemini API under the endpoint gemini-omni-1.1-flash. True to the Omni promise of creating anything from anything, it accepts text, images, video, or audio as input and produces video with synchronized native audio, covering generation, editing, keyframe interpolation, and clip extension in a single model. Developers use it to restyle footage, fill missing frames between keyframes, extend short clips, and convert stills or audio briefs into finished video without stitching separate tools. Pricing follows Gemini API token rules: input costs $1.50 per 1M tokens across any modality, output text is $9.00 per 1M tokens, and video output is $17.50 per 1M tokens, billed at 5,792 tokens per second of 720p video, an effective price of about $0.10 per second. There is no free API tier, so experimentation happens in Google AI Studio against a paid key, and Vertex AI adds provisioned throughput for production. Omni Flash complements Veo 3.1, the cinematic text-to-video flagship: Veo targets filmic generation quality while Omni Flash focuses on flexible input, fast editing loops, and interactive transformation, and both anchor the broader Gemini 3.8 generation announced in September 2026.
คุณสมบัติหลัก
ข้อดี
- +Accepts text, image, audio, or video as input
- +Editing, interpolation, and extension in one model
- +Predictable per second pricing for budgeting
ข้อเสีย
- -No free API tier for experimentation
- -Effective cost climbs quickly on long 720p renders
เหมาะสำหรับ
การผสานรวมและความเข้ากันได้
Frequently Asked Questions
What is Gemini Omni Flash?
Gemini Omni Flash is the Google any-input-to-video model on the paid Gemini API tier. It generates and edits video from text, image, audio, or video prompts, with keyframe interpolation, clip extension, and native audio built in.
How much does Gemini Omni Flash cost?
Input costs $1.50 per 1M tokens for any modality, output text is $9.00 per 1M, and video output is $17.50 per 1M tokens at 5,792 tokens per second of 720p, an effective price of about $0.10 per second. There is no free tier.
How does Gemini Omni Flash differ from Veo?
Veo 3.1 is the cinematic text-to-video flagship focused on filmic quality, while Omni Flash is built for flexible input and fast editing loops such as restyling footage, interpolating keyframes, and extending clips from any starting modality.
Can Gemini Omni Flash edit existing video?
Yes. Editing is a core capability alongside generation: you can restyle scenes, interpolate between keyframes, and extend clips, feeding video, images, audio, or text prompts as the starting point.
Is there a free way to try Gemini Omni Flash?
The API has no free tier for Omni models, but you can test prompts in Google AI Studio with a paid key before committing, and Google AI subscription plans include Gemini video features in the Gemini app.
What inputs does Gemini Omni Flash accept?
It accepts text, images, video, or audio and produces video with synchronized native audio, so one model handles generation, restyling, editing, keyframe interpolation, and clip extension. This removes the need to stitch separate tools for each step of a video pipeline.
How is Gemini Omni Flash different from Veo 3.1?
Veo 3.1 is the cinematic text to video flagship focused on filmic generation quality, while Omni Flash focuses on flexible inputs, fast editing loops, and interactive transformation. Teams often use Veo for hero shots and Omni Flash for iteration and pipeline automation, both on the Gemini 3.8 generation.
What resolutions does Gemini Omni Flash output?
Video output currently targets 720p, billed at 5,792 tokens per second, which works out to about $0.10 per second of video. Higher resolutions may arrive on later endpoints, but 720p with synchronized audio already covers social clips, previews, and most editing workflows.
Is there a free tier for the Gemini Omni Flash API?
No. The model is generally available only on the paid tier of the Gemini API, so experimentation happens in Google AI Studio against a paid key. Vertex AI adds provisioned throughput for production traffic, with input priced at $1.50 per 1M tokens across any modality.
Can Gemini Omni Flash extend a short clip into a longer video?
Yes. Clip extension is one of its core modes: you feed an existing clip and the model continues the motion and audio natively, which is useful for building longer b roll or stretching a 3 second shot into a fuller sequence. Keyframe interpolation works the same way between two stills or two clips.
รายละเอียดคะแนน
ภาษาที่รองรับ
ความเป็นส่วนตัวและความปลอดภัย
Processed under Google Cloud and Gemini API terms, paid tier data not used to improve products