H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-30 · no auto refresh

Round 2 · Matched Seeds

Compact contract vs minimal six-section; deterministic silence/source soundtrack.

Core 2 / 8 · expansion gate failed
ref2va-sample-2026083004-000000

ref2va-sample-2026082240-002931

Animate the woman from the photo talking and moving her head with the performance from the video.
motion_performance_transfercompact Context-IR approximation5ssilentrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 5s
Exact H3 prompt
TASK: Create a 5-second portrait video of the woman from <Picture 1>.
USE: Transfer only the head turns, facial expressions, and speaking-performance timing from <Video 1>.
KEEP: Her face, dark curly hair, white fur coat, white trousers, and the bright windowed setting from <Picture 1>.
EXCLUDE: The man, beard, dark clothing, room, camera look, and audio from <Video 1>. The delivered video is silent.
Final delivered output
REJECTED
The subject identity and speech performance are transferred effectively, but unrequested hard cuts and sudden framing shifts disrupt temporal continuity.
failure: unrequested_change
primary_transformation: 3source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 4artifact: 3
Gemini 3.7 independent output QA · 208.6s generation
ref2va-sample-2026083004-000001

ref2va-sample-2026082240-002931

Animate the woman from the photo talking and moving her head with the performance from the video.
motion_performance_transferContext-IR approximation5ssilentrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 5s
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman from <Picture 1>, including her face, dark curly hair, white fur coat, white trousers, and bright windowed setting.
<Video 1> supplies only head turns, facial expressions, and speaking-performance timing.

summary:
[reference generation] Create a 5-second portrait of <Subject 1> performing the facial and head motion from <Video 1>.

retention_analysis:
<Subject 1>: fully_preserved - retain identity, clothing, and the windowed setting from <Picture 1>.
<Video 1>: attribute_transfer - transfer only performance motion; exclude source identity, clothing, room, camera look, and audio.

detailed_description:
[Shot 1] Keep <Subject 1> in the bright windowed setting from <Picture 1>. Apply only the head movement and facial-expression sequence from <Video 1>.

overall_soundscape:
Silent delivery.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The performance and facial movement transfer succeeded with good target identity retention, though unrequested camera cutaways occur during the sequence.
failure: unrequested_change
primary_transformation: 3source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 3artifact: 3
Gemini 3.7 independent output QA · 165.4s generation
ref2va-sample-2026083004-000002

ref2va-sample-2026082240-009763

Make a video of the man from the picture doing the speaking and mug-holding vlog performance from the video.
motion_performance_transfercompact Context-IR approximation12ssilentrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 12s
Exact H3 prompt
TASK: Create a 12-second video of the man from <Picture 1> performing the speaking and mug-holding motion from <Video 1>.
USE: Transfer only head movement, facial timing, hand motion, and mug handling from <Video 1>.
KEEP: The man's face, grey striped shirt, dark trousers, and plain white-wall studio setting from <Picture 1>.
EXCLUDE: The woman, blonde hair, kitchen, source clothing, source camera look, and audio from <Video 1>. The delivered video is silent.
Final delivered output
REJECTED
Identity and setting match the reference image well, but an abrupt mid-video framing jump and awkward cropping compromise the motion transfer continuity.
failure: motion_transfer_failed, unrequested_change
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 2audio: 3artifact: 3
Gemini 3.7 independent output QA · 626.5s generation
ref2va-sample-2026083004-000003

ref2va-sample-2026082240-009763

Make a video of the man from the picture doing the speaking and mug-holding vlog performance from the video.
motion_performance_transferContext-IR approximation12ssilentaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 12s
Exact H3 prompt
subject_definitions:
<Subject 1> is the man from <Picture 1>, including his face, grey striped shirt, dark trousers, and plain white-wall studio setting.
<Video 1> supplies only speaking, head movement, hand motion, and mug handling.

summary:
[reference generation] Create a 12-second video of <Subject 1> performing the motion from <Video 1>.

retention_analysis:
<Subject 1>: fully_preserved - retain identity, clothing, and white-wall setting from <Picture 1>.
<Video 1>: attribute_transfer - transfer only performance motion; exclude the woman, kitchen, source clothing, camera look, and audio.

detailed_description:
[Shot 1] Keep <Subject 1> against the plain white wall from <Picture 1>. Apply only the speaking and mug-handling performance from <Video 1>.

overall_soundscape:
Silent delivery.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
Successful motion performance transfer applying the speaking and mug-holding actions to the target identity with strong identity retention and clean temporal coherence.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 627.4s generation
ref2va-sample-2026083004-000004

ref2va-sample-2026082240-001528

Replace the speaker in the video with the woman from the photo, keeping all her original hand gestures, speech, and the background.
source_preserving_editingcompact Context-IR approximation8ssource_soundtrackaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
<Audio 1>source_soundtrack
audio · 8.0s
Exact H3 prompt
TASK: In the 8-second <Video 1>, replace only the visible speaker with the woman from <Picture 1>.
USE: <Picture 1> only for the replacement speaker's identity and clothing.
KEEP: The source room, animal pictures, camera, framing, gestures, expression timing, and full timeline from <Video 1>.
EXCLUDE: The original speaker's visible identity. Do not change any other visual content. <Audio 1> will be reused unchanged after generation.
Final delivered output
ACCEPTED
The subject replacement successfully transfers the identity from the reference image while faithfully preserving the original motions, background, and audio track.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 314.8s generation
ref2va-sample-2026083004-000005

ref2va-sample-2026082240-001528

Replace the speaker in the video with the woman from the photo, keeping all her original hand gestures, speech, and the background.
source_preserving_editingContext-IR approximation8ssource_soundtrackrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
<Audio 1>source_soundtrack
audio · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the replacement speaker from <Picture 1>.
<Video 1> is the source clip to edit.
<Audio 1> is the synchronized source soundtrack reused after generation.

summary:
[video editing + reference generation] In the 8-second <Video 1>, replace only the visible speaker with <Subject 1>.

retention_analysis:
<Subject 1>: fully_preserved - use the identity and clothing from <Picture 1>.
<Video 1>: partially_preserved - retain room, animal pictures, camera, framing, gestures, expression timing, and full timeline; replace only visible speaker identity.
<Audio 1>: fully_copy - reuse as the complete synchronized soundtrack after generation.

detailed_description:
[Shot 1] Match <Video 1> from its first frame through its last frame. Replace only the speaker's visible identity with <Subject 1>. Do not alter any other visual content.

overall_soundscape:
Reuse <Audio 1> unchanged and synchronized after generation.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The identity replacement and motion preservation were executed smoothly, though the subject's clothing diverges from the provided reference image.
failure: unrequested_change
primary_transformation: 4source_preservation: 4target_identity: 3source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 315.3s generation
ref2va-sample-2026083004-000006

ref2va-sample-2026082240-001528

Replace the speaker in the video with the woman from the photo, keeping all her original hand gestures, speech, and the background.
source_preserving_editingcompact Context-IR approximation8ssource_soundtrackrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
<Audio 1>source_soundtrack
audio · 8.0s
Exact H3 prompt
TASK: In the 8-second <Video 1>, replace only the visible speaker with the woman from <Picture 1>.
USE: <Picture 1> only for the replacement speaker's identity and clothing.
KEEP: The source room, animal pictures, camera, framing, gestures, expression timing, and full timeline from <Video 1>.
EXCLUDE: The original speaker's visible identity. Do not change any other visual content. <Audio 1> will be reused unchanged after generation.
Final delivered output
REJECTED
The requested identity replacement did not occur; the output retains the original speaker completely unchanged.
failure: primary_transformation_missing, target_identity_lost, source_identity_leakage
primary_transformation: 2source_preservation: 4target_identity: 2source_leakage: 4motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 314.2s generation
ref2va-sample-2026083004-000007

ref2va-sample-2026082240-001528

Replace the speaker in the video with the woman from the photo, keeping all her original hand gestures, speech, and the background.
source_preserving_editingContext-IR approximation8ssource_soundtrackrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
<Audio 1>source_soundtrack
audio · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the replacement speaker from <Picture 1>.
<Video 1> is the source clip to edit.
<Audio 1> is the synchronized source soundtrack reused after generation.

summary:
[video editing + reference generation] In the 8-second <Video 1>, replace only the visible speaker with <Subject 1>.

retention_analysis:
<Subject 1>: fully_preserved - use the identity and clothing from <Picture 1>.
<Video 1>: partially_preserved - retain room, animal pictures, camera, framing, gestures, expression timing, and full timeline; replace only visible speaker identity.
<Audio 1>: fully_copy - reuse as the complete synchronized soundtrack after generation.

detailed_description:
[Shot 1] Match <Video 1> from its first frame through its last frame. Replace only the speaker's visible identity with <Subject 1>. Do not alter any other visual content.

overall_soundscape:
Reuse <Audio 1> unchanged and synchronized after generation.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The requested subject replacement completely failed, leaving the original source speaker unchanged throughout the video.
failure: primary_transformation_missing, target_identity_lost, source_identity_leakage
primary_transformation: 2source_preservation: 4target_identity: 2source_leakage: 4motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 314.4s generation
ref2va-sample-2026083004-000008

ref2va-sample-2026082240-001452

Convert this video into high-contrast black and white while preserving the original actions and timing.
source_preserving_editingcapability_control10sgeneratedrejected
<Video 1>source_edit
video · 10.0s
Exact H3 prompt
Edit <Video 1> only by converting every frame to high-contrast black and white. Preserve the complete source timeline, people, hutch, backyard, action, framing, and camera exactly.
Final delivered output
REJECTED
The model preserved the scene motion and structure but failed to apply the requested black-and-white visual stylization.
failure: primary_transformation_missing
primary_transformation: 2source_preservation: 4target_identity: 3source_leakage: 1motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 379.7s generation
ref2va-sample-2026083004-000009

ref2va-sample-2026082240-001601

Change the lighting in this video to high-contrast studio lighting with deep crushed blacks and crisp specular highlights on the glossy mouse, keeping the camera framing, hand movement, and original audio intact.
source_preserving_editingcapability_control5ssource_soundtrackrejected
<Video 1>source_edit
video · 5.0s
<Audio 1>source_soundtrack
audio · 5.0s
Exact H3 prompt
Edit <Video 1> only by applying a crisp balanced contrast grade. Start from the exact first source frame and preserve the complete mouse movement, hand movement, desk, framing, and timing. <Audio 1> will be reused unchanged after generation.
Final delivered output
REJECTED
The generated output fails the prompt requirements by omitting the first half of the source video timeline, altering the temporal synchronization, and failing to execute the requested high-contrast studio lighting grade.
failure: primary_transformation_missing, continuation_boundary_failed, duration_mismatch
primary_transformation: 1source_preservation: 2target_identity: 5source_leakage: 1motion_or_temporal: 2audio: 3artifact: 4
Gemini 3.7 independent output QA · 119.8s generation