H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-31 04:43 PDT · no auto refresh

Exact Dialogue Stability 8

Four semantically accepted A1/A3 workloads, each repeated at two new seeds. This tests whether explicit short dialogue is more stable than requesting silence.

Generated 8/8 · raw QA 6/8
ref2va-sample-2026083190-000000

ref2va-sample-2026083103-000008

Make a realistic 5-second video of the teenage boy from the photo standing in the same covered walkway. Have him adjust his backpack straps, look at the camera, and say "I'm ready." Keep his clothing, mask, and the background the same.
character_subjectdirect ablation5sgeneratedexact_dialogueaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate a realistic 5-second video using the teenage boy from <Picture 1> standing in the covered school walkway. He raises both hands to adjust his backpack straps, looks directly toward the camera, and says in English: "I'm ready." No other dialogue, background chatter, or background music. Preserve his identity, face mask, clothing, and the walkway setting from <Picture 1>.
Final delivered output
ACCEPTED
The video faithfully follows the prompt instructions, maintaining character identity and scene consistency while accurately executing the motion and exact spoken dialogue.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 113.2s generation
ref2va-sample-2026083190-000001

ref2va-sample-2026083103-000008

Make a realistic 5-second video of the teenage boy from the photo standing in the same covered walkway. Have him adjust his backpack straps, look at the camera, and say "I'm ready." Keep his clothing, mask, and the background the same.
character_subjectdirect ablation5sgeneratedexact_dialogueaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate a realistic 5-second video using the teenage boy from <Picture 1> standing in the covered school walkway. He raises both hands to adjust his backpack straps, looks directly toward the camera, and says in English: "I'm ready." No other dialogue, background chatter, or background music. Preserve his identity, face mask, clothing, and the walkway setting from <Picture 1>.
Final delivered output
ACCEPTED
Accurate character animation and environment preservation matching all prompt instructions and dialogue constraints.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 98.5s generation
ref2va-sample-2026083190-000002

ref2va-sample-2026083103-000009

Make a 5-second anime clip of the character from <Picture 1> getting into a combat stance and shouting "Let's go!"
character_subjectdirect ablation5sgeneratedexact_dialogueaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
A 5-second anime video featuring the character from <Picture 1> on the outdoor industrial platform. She raises her clenched fists into a combat-ready pose, looks forward with a confident grin, and speaks clearly: <d>[English] Let's go!</d>. Afterwards, her mouth remains closed in a determined smile. Outdoor ambient breeze and rustling cloth only, no background music.
Final delivered output
ACCEPTED
The generated video accurately realizes the requested character animation, combat stance, and dialogue timing with strong consistency.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 61.9s generation
ref2va-sample-2026083190-000003

ref2va-sample-2026083103-000009

Make a 5-second anime clip of the character from <Picture 1> getting into a combat stance and shouting "Let's go!"
character_subjectdirect ablation5sgeneratedexact_dialogueaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
A 5-second anime video featuring the character from <Picture 1> on the outdoor industrial platform. She raises her clenched fists into a combat-ready pose, looks forward with a confident grin, and speaks clearly: <d>[English] Let's go!</d>. Afterwards, her mouth remains closed in a determined smile. Outdoor ambient breeze and rustling cloth only, no background music.
Final delivered output
ACCEPTED
The generated video accurately depicts the character performing the requested combat pose and speaking the exact dialogue.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 54.2s generation
ref2va-sample-2026083190-000004

ref2va-sample-2026083103-000010

Create a 5-second realistic video set in the bookstore from <Picture 3>. The woman from <Picture 1> stands next to the woman from <Picture 2> looking at the empty bookshelves. The woman from <Picture 1> says, "These shelves are empty." while the woman from <Picture 2> quietly looks at the shelf without speaking.
multi_subject_interactiondirect ablation5sgeneratedexact_dialoguerejected
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
In the empty bookstore from <Picture 3>, the woman from <Picture 1> stands beside the woman from <Picture 2> in front of the labeled wooden bookshelves. The woman from <Picture 1> looks at the empty shelves and says clearly, "These shelves are empty." before closing her mouth. The woman from <Picture 2> looks at the shelves and remains silent with her mouth closed throughout. Realistic lighting and quiet indoor bookstore room tone, with no other dialogue, voices, or subtitles.
Final delivered output
REJECTED
Visual subjects and setting match well, but the speech dialogue fails into unintelligible words instead of the required text.
failure: unintelligible_speech
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 independent output QA · 125.9s generation
ref2va-sample-2026083190-000005

ref2va-sample-2026083103-000010

Create a 5-second realistic video set in the bookstore from <Picture 3>. The woman from <Picture 1> stands next to the woman from <Picture 2> looking at the empty bookshelves. The woman from <Picture 1> says, "These shelves are empty." while the woman from <Picture 2> quietly looks at the shelf without speaking.
multi_subject_interactiondirect ablation5sgeneratedexact_dialoguerejected
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
In the empty bookstore from <Picture 3>, the woman from <Picture 1> stands beside the woman from <Picture 2> in front of the labeled wooden bookshelves. The woman from <Picture 1> looks at the empty shelves and says clearly, "These shelves are empty." before closing her mouth. The woman from <Picture 2> looks at the shelves and remains silent with her mouth closed throughout. Realistic lighting and quiet indoor bookstore room tone, with no other dialogue, voices, or subtitles.
Final delivered output
REJECTED
Visual identities and environment match the references well, but the speech violates the exact dialogue policy with nonsensical generated audio.
failure: unintelligible_speech, unexpected_speech
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 3source_preservation: 3target_identity: 3source_leakage: 2motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 independent output QA · 125.8s generation
ref2va-sample-2026083190-000006

ref2va-sample-2026083103-000011

Create a 5-second realistic video set in the woods from <Picture 3>. Have the man from <Picture 1> and the woman from <Picture 2> stand together. The man looks over at her and says, "It's peaceful here." She smiles silently in response without speaking.
multi_subject_interactiondirect ablation5sgeneratedexact_dialogueaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
A realistic 5-second shot in the lush wooded environment from <Picture 3>. The man from <Picture 1> stands beside the woman from <Picture 2>. Looking toward her with a gentle smile, the man from <Picture 1> (S1) speaks clearly in English: <d>[English] It's peaceful here.</d> The woman from <Picture 2> maintains closed lips, smiling back warmly in silence. Soft ambient forest sounds with rustling leaves and distant birdsong accompany the scene; no non-diegetic music.
Final delivered output
ACCEPTED
The generated video accurately incorporates both subject identities and the environment reference, flawlessly executing the requested motion, dialogue, and interaction.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 127.7s generation
ref2va-sample-2026083190-000007

ref2va-sample-2026083103-000011

Create a 5-second realistic video set in the woods from <Picture 3>. Have the man from <Picture 1> and the woman from <Picture 2> stand together. The man looks over at her and says, "It's peaceful here." She smiles silently in response without speaking.
multi_subject_interactiondirect ablation5sgeneratedexact_dialogueaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
A realistic 5-second shot in the lush wooded environment from <Picture 3>. The man from <Picture 1> stands beside the woman from <Picture 2>. Looking toward her with a gentle smile, the man from <Picture 1> (S1) speaks clearly in English: <d>[English] It's peaceful here.</d> The woman from <Picture 2> maintains closed lips, smiling back warmly in silence. Soft ambient forest sounds with rustling leaves and distant birdsong accompany the scene; no non-diegetic music.
Final delivered output
ACCEPTED
The video accurately depicts both subjects in the specified environment with the correct dialogue, timing, and silent response.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 126.5s generation