Scope. First-party capability check for the two prepared tutorials. This note distinguishes the vendor API/model claims below from MITO product integration behavior, which must be verified in MITO before publication. Sources were checked 2026-09-08.
Gemini Omni Flash 1.1
Supported, externally verifiable claims
- Image and video references: Google documents image and video reference roles. A video can be used as a character or object reference; the guide also shows an image reference combined with a video reference. This supports wording that a phone clip can be provided as a reference, but not the stronger promise that it will exactly reproduce its movement or camera work. Google Gemini API: Generate and edit videos with Gemini Omni Flash
- Reference limits: Video references work best for likenesses; their audio is ignored. The current API supports at most three reference clips, each up to three seconds. It does not support reasoning across multiple videos. Use these constraints if the tutorial explains how to prepare a phone clip. Google Gemini API limitations
- Image-to-video motion direction: Google’s image-to-video example explicitly directs the model to use a drawing only as a guide for movement. Its guidance recommends high-resolution source images and specific descriptions of camera, subject, and environmental motion. Google Gemini API: Image to video generation
- Editing: The model accepts straightforward video-edit prompts, including changing lighting, adding an object, or changing sign text. Google recommends a short instruction plus “Keep everything else the same” when preserving the rest of the shot matters. Google Gemini API: Prompts for editing
- First/last-frame transitions: Two images plus a transition prompt can create a video that moves from the starting image to the ending image. Google Gemini API: First and last frame interpolation
- Extensions: Omni 1.1 Flash can append a prompted 10-second extension, up to 40 seconds total. Google says it uses the last 10 seconds as context to preserve video, motion, characters, and audio coherence. Uploaded-video extension is end-only; uploaded clips must be 10 seconds or shorter; it is unavailable in the EEA, Switzerland, and UK. Google Gemini API: Prompts for extending a video and extension constraints
- Audio and text: Google says the model attempts to generate an appropriate audio track by default, can take prompt direction for audio, and can render prompted text correctly and readably. These remain generative results, not guarantees; review output before relying on either. Google Gemini API: Prompting audio and text
Claims to avoid or qualify
- Do not say a phone clip is a reliable “motion transfer” or that its camera moves will be faithfully copied. Google only documents video reference as character/object reference and says such references work best with likenesses; it does not promise motion transfer.
- Do not imply a video reference’s sound is used: Google says it is ignored.
- Do not imply editing or extension is available in every location. In particular, uploaded-video extension has an EEA/Switzerland/UK restriction.
MiniMax H3 Max
Supported, externally verifiable claims
- Supported generation modes: H3 Max supports text-to-video plus image-to-video using a first frame, last frame, or both. MiniMax API overview and MiniMax model introduction
- Output constraints: MiniMax lists H3 Max at 480P or 768P, 5–15 seconds, with no 2K output. MiniMax API overview
Claims contradicted by current official documentation
- Reference input: The prepared H3 Max post says it can use approved image/video references and frames separately. MiniMax’s current API overview expressly says H3 Max does not support reference input. That workflow belongs to H3, not H3 Max. MiniMax API overview
- Audio reference or native sound: The prepared post says H3 Max supports an audio reference when paired with image/video. MiniMax documents those multimodal/audio-reference constraints for H3; it does not list audio or reference input for H3 Max. Do not publish that H3 Max claim without product-specific evidence that MITO uses a different backend capability. MiniMax H3 announcement, model variants
- Readable text: No first-party H3 Max source located supports a claim that it renders text accurately or legibly. Treat this as unsubstantiated; preserve the “review client-critical detail” advice but remove any capability promise.
Important naming distinction
MiniMax’s official materials make a material distinction between MiniMax H3 and MiniMax H3 Max. H3 is the omni-modal model with multimodal reference support, native stereo audio, 768P/2K, and 4–15-second output. H3 Max is the high-speed variant restricted to text-to-video and first/last-frame image-to-video, 480P/768P, and 5–15 seconds. MiniMax H3 announcement and MiniMax API overview
Publication implication
The Omni Flash 1.1 tutorial can cite Google’s official guide for references, editing, interpolation, and extensions, subject to the qualifiers above.
For the H3 Max tutorial, MITO product behavior is authoritative for what MITO customers can use. MITO’s current client constraints explicitly handle H3 Max image, video, and audio references, require a visual reference alongside audio, and prevent combining those references with start or end frames. Mara confirmed on 2026-09-08 that the MITO integration supports references. Do not cite MiniMax’s generic API overview as evidence against the MITO workflow; it appears incomplete or differently scoped. Until a public MITO help page documents the integration, keep H3 Max’s product-specific claims attributable to MITO rather than linking to the conflicting vendor page.