H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-30 · no auto refresh

Round 1 · Prompt A/B

Direct vs minimal six-section, matched seeds, with final soundtrack delivery where applicable.

2 accepted / 10 after final delivery
ref2va-sample-2026083002-000000

ref2va-sample-2026082240-002931

Animate the woman from the image in a plain studio portrait shot, using only the speaking head motion and facial performance from the video. Keep her identity and clothing; do not copy the source person, room, camera, or audio.
motion_performance_transferdirect ablation8sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 6.673s
Exact H3 prompt
Create an 8-second plain studio portrait of the woman from <Picture 1>. Transfer only the head motion, facial expressions, and speaking performance timing from <Video 1>. Preserve her identity and clothing. Do not copy the source performer, room, camera, or audio.
Final delivered output
REJECTED
Target identity and motion were transferred, but the output suffered heavy leakage of the motion video's environment and clothing.
failure: source_environment_leakage, unrequested_change
primary_transformation: 3source_preservation: 2target_identity: 3source_leakage: 4motion_or_temporal: 3audio: 3artifact: 3
Gemini 3.7 independent output QA · 325.6s generation
ref2va-sample-2026083002-000001

ref2va-sample-2026082240-002931

Animate the woman from the image in a plain studio portrait shot, using only the speaking head motion and facial performance from the video. Keep her identity and clothing; do not copy the source person, room, camera, or audio.
motion_performance_transferContext-IR approximation8sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 6.673s
Exact H3 prompt
subject_definitions:
<Subject 1> is the target woman from <Picture 1>.
<Video 1> supplies only head motion, facial expressions, and speaking-performance timing.

summary:
[reference generation] Create a 8-second portrait of <Subject 1> performing the facial and head motion from <Video 1>.

retention_analysis:
<Subject 1>: fully_preserved - retain identity and clothing from <Picture 1>.
<Video 1>: attribute_transfer - transfer only performance motion; exclude source identity, clothing, room, camera, style, and audio.

detailed_description:
[Shot 1] One continuous medium portrait on a plain neutral studio background. <Subject 1> follows the head movement, facial-expression sequence, and speaking cadence from <Video 1>. Do not copy the source performer or setting.

overall_soundscape:
Quiet neutral room tone. Do not use audio from <Video 1>.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The identity and speaking performance were transferred with good visual quality, but the subject attire was altered and unexpected speech audio was introduced.
failure: unrequested_change, audio_missing_or_changed
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 2artifact: 4
Gemini 3.7 independent output QA · 293.7s generation
ref2va-sample-2026083002-000002

ref2va-sample-2026082240-009763

Animate the man from the image in a neutral indoor shot, transferring only the direct-to-camera speaking and mug-holding performance from the video. Do not copy the source woman, kitchen, clothing, camera, or audio.
motion_performance_transferdirect ablation12sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 13.9s
Exact H3 prompt
Create a 12-second neutral indoor video of the man from <Picture 1>. Transfer only the speaking, head movement, and mug-holding performance from <Video 1>. Preserve the man's identity and clothing. Do not copy the source performer, kitchen, camera, visual style, or audio.
Final delivered output
REJECTED
Motion transfer and subject identity preservation succeeded, but the background environment leaked directly from the motion reference video contrary to instructions.
failure: source_environment_leakage
primary_transformation: 3source_preservation: 4target_identity: 4source_leakage: 4motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 659.4s generation
ref2va-sample-2026083002-000003

ref2va-sample-2026082240-009763

Animate the man from the image in a neutral indoor shot, transferring only the direct-to-camera speaking and mug-holding performance from the video. Do not copy the source woman, kitchen, clothing, camera, or audio.
motion_performance_transferContext-IR approximation12sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 13.9s
Exact H3 prompt
subject_definitions:
<Subject 1> is the target man from <Picture 1>.
<Video 1> supplies only direct-to-camera speaking, head movement, and mug-holding performance.

summary:
[reference generation] Create a 12-second neutral indoor shot of <Subject 1> performing the motion from <Video 1>.

retention_analysis:
<Subject 1>: fully_preserved - retain identity and clothing from <Picture 1>.
<Video 1>: attribute_transfer - transfer only performance motion; exclude source identity, clothing, kitchen, camera, style, and audio.

detailed_description:
[Shot 1] One continuous medium shot in a simple neutral room. <Subject 1> looks toward the camera and follows the speaking, head movement, and mug-handling timing from <Video 1>. Do not copy the source performer or kitchen.

overall_soundscape:
Quiet neutral room tone. Do not use audio from <Video 1>.

non_diegetic_music:
N/A
Final delivered output
REJECTED
Motion transfer and identity retention are executed well, though unrequested speech audio was generated instead of neutral room tone.
failure: unrequested_change
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 662.9s generation
ref2va-sample-2026083002-000004

ref2va-sample-2026082240-001452

Convert this video to high-contrast black and white. Keep the source action, timing, framing, and scene unchanged.
source_preserving_editingdirect ablation10sgeneratedrejected
<Video 1>source_edit
video · 10.0s
Exact H3 prompt
Edit <Video 1> only by applying a high-contrast black-and-white grade. Keep every subject, action, object, background, camera frame, and timing unchanged.
Final delivered output
REJECTED
The primary transformation to convert the video to black and white was completely missed, leaving the output in full color despite preserving motion and framing.
failure: primary_transformation_missing
primary_transformation: 2source_preservation: 4target_identity: 3source_leakage: 3motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 379.3s generation
ref2va-sample-2026083002-000005

ref2va-sample-2026082240-001452

Convert this video to high-contrast black and white. Keep the source action, timing, framing, and scene unchanged.
source_preserving_editingContext-IR approximation10sgeneratedrejected
<Video 1>source_edit
video · 10.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the source clip to edit.

summary:
[video editing] The target is a 10-second edited version of <Video 1>. Apply only a high-contrast black-and-white grade.

retention_analysis:
<Video 1>: partially_preserved - preserve all subjects, objects, action, background, framing, and timing; change only color and contrast.

detailed_description:
[Shot 1] Match <Video 1> frame for frame. Remove all color and apply strong but readable monochrome contrast. Do not add, remove, replace, relight, recamera, or retime anything.

overall_soundscape:
N/A

non_diegetic_music:
N/A
Final delivered output
REJECTED
The requested monochrome high-contrast style transformation was completely omitted, leaving the video in original color.
failure: primary_transformation_missing
primary_transformation: 2source_preservation: 4target_identity: 3source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 380.4s generation
ref2va-sample-2026083002-000006

ref2va-sample-2026082240-001601

Give this video a crisp, balanced contrast grade without changing the mouse, hand movement, framing, timing, or original audio.
source_preserving_editingdirect ablation5sgeneratedrejected
<Video 1>source_edit
video · 5.0s
<Audio 1>source_soundtrack
audio · 5.0s
Exact H3 prompt
Edit <Video 1> only by applying a crisp, balanced contrast grade with clean highlights and readable shadow detail. Preserve the mouse, hand movement, framing, and timing exactly. Reuse <Audio 1> as the synchronized source soundtrack.
Final delivered output
REJECTED
The requested contrast grade is visible, but the video starts mid-clip rather than preserving the complete timeline and motion from the source.
failure: duration_mismatch, unrequested_change
primary_transformation: 3source_preservation: 2target_identity: 3source_leakage: 3motion_or_temporal: 3audio: 3artifact: 4
Gemini 3.7 independent output QA · 119.7s generation
ref2va-sample-2026083002-000007

ref2va-sample-2026082240-001601

Give this video a crisp, balanced contrast grade without changing the mouse, hand movement, framing, timing, or original audio.
source_preserving_editingContext-IR approximation5sgeneratedrejected
<Video 1>source_edit
video · 5.0s
<Audio 1>source_soundtrack
audio · 5.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the source clip to edit.
<Audio 1> is the synchronized source soundtrack.

summary:
[video editing + audio reuse] The target is a 5-second edited version of <Video 1>. Apply only a crisp, balanced contrast grade and reuse <Audio 1>.

retention_analysis:
<Video 1>: partially_preserved - preserve the mouse, hand movement, desk, framing, and timing; change only the color grade.
<Audio 1>: fully_copy - reuse as the complete synchronized soundtrack.

detailed_description:
[Shot 1] Match <Video 1> frame for frame. Add clean highlights and readable shadow contrast without crushing blacks, changing the lighting setup, or replacing any content.

overall_soundscape:
Reuse <Audio 1> unchanged and synchronized.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The requested contrast grade is moderately applied, but the output starts midway through the source clip rather than maintaining the original starting point.
failure: continuation_boundary_failed, unrequested_change
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 3artifact: 4
Gemini 3.7 independent output QA · 119.1s generation
ref2va-sample-2026083002-000008

ref2va-sample-2026082240-001528

Replace the speaker in the video with the woman from the image. Keep the original gestures, framing, room, timing, and dialogue.
source_preserving_editingdirect ablation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
<Audio 1>source_soundtrack
audio · 8.0s
Exact H3 prompt
In <Video 1>, replace only the visible speaker with the woman from <Picture 1>. Preserve the source gestures, framing, room, and timing. Reuse <Audio 1> as the synchronized source dialogue. Do not change anything else.
Final delivered output
ACCEPTED
The subject replacement successfully transfers the target identity while preserving the source video background, motion timing, and dialogue audio.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 317.8s generation
ref2va-sample-2026083002-000009

ref2va-sample-2026082240-001528

Replace the speaker in the video with the woman from the image. Keep the original gestures, framing, room, timing, and dialogue.
source_preserving_editingContext-IR approximation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
<Audio 1>source_soundtrack
audio · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the replacement speaker from <Picture 1>.
<Video 1> is the source clip to edit.
<Audio 1> is the synchronized source dialogue and room sound.

summary:
[video editing + reference generation + audio reuse] The target is a 8-second edited version of <Video 1>. Replace only the visible speaker with <Subject 1>.

retention_analysis:
<Subject 1>: fully_preserved - use the identity from <Picture 1>.
<Video 1>: partially_preserved - preserve gestures, framing, room, timing, and all non-identity content.
<Audio 1>: fully_copy - reuse as the complete synchronized soundtrack.

detailed_description:
[Shot 1] Match <Video 1> frame for frame. Replace only the speaker's visible identity with <Subject 1> while retaining the source pose, gestures, expression timing, background, and camera. Do not alter any other content.

overall_soundscape:
Reuse <Audio 1> unchanged and synchronized.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
Successful identity replacement preserving original motion, framing, and synchronized soundtrack.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 315.2s generation