H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-30 · no auto refresh

Repair 15 · Base-Compatible

Fifteen coherent real-workload combinations, three each for A1, A3, A4, B3, and D4. Direct prompts use only FastVideo-bound media labels; unsupported Subject labels are removed.

Generated 15 / 15 · QA 7/15 accepted
ref2va-sample-2026083500-000000

ref2va-sample-2026083032-000000

Make the woman from the picture talk and gesture with the exact same body movements as the woman in the video.
motion_performance_transferdirect ablation10sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 10.0s
Exact H3 prompt
Generate a 10-second video where the blonde woman from <Picture 1> in her kitchen performs the exact hand gestures, head motion, and speaking delivery of the presenter in <Video 1>, keeping her visual identity, wardrobe, and environment strictly grounded in <Picture 1>.
Final delivered output
REJECTED
The performance and motion transfer failed due to unprompted shot changes, inconsistent camera perspectives, and synthesized unintelligible speech.
prompt drift: audible speech lacks exact dialogue text or an audio reference
failure: motion_transfer_failed, unintelligible_speech, unrequested_change
media contract: 10.176s actual / 10.0s target · audio present · pass
primary_transformation: 2source_preservation: 2target_identity: 3source_leakage: 2motion_or_temporal: 2audio: 2artifact: 2
Gemini 3.7 independent output QA · 409.5s generation
ref2va-sample-2026083500-000001

ref2va-sample-2026083032-000001

Continue the video from where the woman finishes applying the white powder base, showing her inspecting her face in the mirror and picking up a detail brush to start the next makeup step.
continuation_next_shotdirect ablation10sgeneratedaccepted
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
Continue seamlessly from the end of <Video 1> for 10 seconds. The woman, with her face covered in white base makeup, pauses to inspect her reflection in the mirror, slightly turning her head left and right to check the coverage. She then reaches down out of frame, picks up a thin detail makeup brush, and brings it up toward her eye area to begin the next stage of her application.
Final delivered output
ACCEPTED
The continuation accurately follows the prompt instructions, maintaining subject appearance, environment, and plausible makeup actions without notable artifacts.
failure: none
media contract: 10.176s actual / 10.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 417.0s generation
ref2va-sample-2026083500-000002

ref2va-sample-2026083032-000003

Make an anime video of the girl from the picture dancing and posing on a concert stage.
character_subjectdirect ablation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate an 8-second anime video featuring the blonde idol girl from <Picture 1> performing on a brightly lit concert stage. She performs an energetic dance routine, smiling brightly, gesturing towards the crowd, and giving a playful wink as colorful stage lights flare in the background.
Final delivered output
ACCEPTED
The generated video successfully animates the reference character performing on a concert stage with high fidelity to the prompt and character design.
failure: none
media contract: 8.032s actual / 8.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 122.3s generation
ref2va-sample-2026083500-000003

ref2va-sample-2026083032-000004

Continue the video from where it ends, showing the person continuing to paint the leaf design with blue paint on the paper.
continuation_next_shotdirect ablation10sgeneratedrejected
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
Continue seamlessly from the final frame of <Video 1>, showing the person using the paintbrush to apply thick blue acrylic paint across the remaining areas of the leaf stencil on the white paper.
Final delivered output
REJECTED
Continuation failed as the model replayed earlier source actions rather than continuing past the final frame, alongside losing source voiceover.
failure: continuation_boundary_failed, audio_missing_or_changed
media contract: 10.176s actual / 10.0s target · audio present · pass
primary_transformation: 2source_preservation: 3target_identity: 3source_leakage: 2motion_or_temporal: 2audio: 2artifact: 3
Gemini 3.7 independent output QA · 300.0s generation
ref2va-sample-2026083500-000004

ref2va-sample-2026083032-000005

Create a sleek 12-second commercial showcasing the pale green cosmetic tube from the reference photo in a clean, high-end studio environment with dynamic lighting and macro angles.
product_objectdirect ablation12sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
Generate a 12-second commercial showcasing the cosmetic squeeze tube from <Picture 1>. Place the pale green tube with its crystalline cap on a minimalist stone pedestal surrounded by gentle water ripples and warm morning sunlight, filmed with slow macro pans, crisp focus pulls, and elegant studio reflections.
Final delivered output
ACCEPTED
The model accurately reconstructed the reference product and placed it in the requested commercial environment with high fidelity and smooth motion.
failure: none
media contract: 12.288s actual / 12.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 200.1s generation
ref2va-sample-2026083500-000005

ref2va-sample-2026083032-000008

Animate the man in Picture 1 to speak and gesture with the exact head movements, facial expressions, and broadcast delivery of the presenter in Video 1, keeping his appearance and studio setting identical.
motion_performance_transferdirect ablation5sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 5.0s
Exact H3 prompt
Generate a 5-second video animating the man from <Picture 1> seated in his studio, transferring the speaking performance, head movements, facial expressions, and hand gesture timing from <Video 1> while completely preserving the man's identity, clothing, headphones, microphone, and memorabilia background from <Picture 1>.
Final delivered output
REJECTED
Visual identity and performance timing transfer well from the references, but the generated audio consists of unintelligible pseudo-speech, requiring rejection.
prompt drift: audible speech lacks exact dialogue text or an audio reference
failure: unintelligible_speech
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 3source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 2artifact: 4
Gemini 3.7 independent output QA · 162.5s generation
ref2va-sample-2026083500-000006

ref2va-sample-2026083032-000009

Create a clean 5-second commercial studio showcase video rotating around the cosmetic tube shown in <Picture 1>.
product_objectdirect ablation5sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
A 5-second commercial product showcase of the cosmetic tube from <Picture 1>. The tube stands gracefully on a minimalist studio podium against a soft pastel background. The camera performs a smooth, slow orbiting shot around the tube, highlighting its label artwork and sleek plastic packaging with elegant studio lighting.
Final delivered output
ACCEPTED
The model successfully creates a commercial studio showcase revolving around the cosmetic tube from Picture 1 with consistent product identity and clean lighting.
failure: none
media contract: 5.216s actual / 5.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 66.6s generation
ref2va-sample-2026083500-000007

ref2va-sample-2026083032-000010

Continue the video to show the cook wrapping the bacon weave around the stuffed meatloaf.
continuation_next_shotdirect ablation10sgeneratedrejected
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
Continue directly from the final frame of <Video 1> as the cook uses the parchment paper and both hands to tightly roll and wrap the bacon weave around the stuffed meatloaf roll on the checkered table.
Final delivered output
REJECTED
The continuation performs the requested wrapping action with good identity and scene preservation, but exhibits a temporal continuity mismatch at the starting boundary.
failure: continuation_boundary_failed
media contract: 10.176s actual / 10.0s target · audio present · pass
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 3artifact: 4
Gemini 3.7 independent output QA · 373.6s generation
ref2va-sample-2026083500-000008

ref2va-sample-2026083032-000011

Create an 8-second video of the woman from the photo walking along a fashion runway stage under soft spotlights, pausing to pose with a hand on her hip and a confident smile, then turning smoothly.
character_subjectdirect ablation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
A full-body shot of the woman from <Picture 1> walking gracefully along a sleek fashion runway stage under warm spotlights. She pauses near the front of the stage, rests a hand on her hip, turns slightly toward the camera with a subtle, confident smile, and then smoothly turns around to walk back.
Final delivered output
ACCEPTED
The generated video accurately translates the reference image into the requested runway walk and turn sequence with high visual fidelity and consistent temporal motion.
failure: none
media contract: 8.032s actual / 8.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 119.4s generation
ref2va-sample-2026083500-000009

ref2va-sample-2026083032-000012

Make a sleek 10-second commercial product showcase video featuring the hand cream tube from the picture in an elegant studio setting with a smooth orbiting camera.
product_objectdirect ablation10sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
Create a 10-second commercial product showcase video featuring the hand cream tube shown in <Picture 1>. The tube stands upright on a smooth matte pedestal in a minimalist cosmetic studio setting with soft warm lighting. The camera performs a continuous, elegant 360-degree orbit around the product, highlighting its packaging details and sleek design in crisp macro focus.
Final delivered output
ACCEPTED
The generated video successfully showcases the product from the reference image with faithful identity retention, smooth camera motion, and high rendering quality.
failure: none
media contract: 10.176s actual / 10.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 149.1s generation
ref2va-sample-2026083500-000011

ref2va-sample-2026083032-000015

Create a video of the woman from Picture 1 and the man from Picture 2 meeting and interacting near the counter inside the room from Picture 3.
multi_subject_interactiondirect ablation12sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
In the modern reception area from <Picture 3>, the woman from <Picture 1> stands by the white counter. The man from <Picture 2> walks over, greets her with a warm nod, and gestures toward the counter as both share a polite non-verbal interaction under bright interior lighting.
Final delivered output
REJECTED
Characters and environment match references well, but the video suffers from an erratic initial cut and unintelligible generated speech.
failure: unintelligible_speech, visible_artifacts, unrequested_change
media contract: 12.288s actual / 12.0s target · audio present · pass
primary_transformation: 3source_preservation: 3target_identity: 3source_leakage: 2motion_or_temporal: 3audio: 2artifact: 2
Gemini 3.7 independent output QA · 289.4s generation
ref2va-sample-2026083500-000012

ref2va-sample-2026083032-000016

Animate the warrior from Picture 1 in a dynamic fight sequence across a misty bamboo forest at twilight, aiming and firing his crossbow while leaping across branches.
character_subjectdirect ablation12sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate a 12-second stylized anime cinematic video of the warrior from <Picture 1> fighting in a misty bamboo forest at twilight. The warrior leaps between tall bamboo stalks, cocks his mechanical crossbow, takes aim with sharp focus, and fires glowing bolts into the mist as his black robes and ribbons billow in the wind, backed by dramatic taiko drums and weapon sound effects.
Final delivered output
ACCEPTED
The generated video successfully translates the reference character into a dynamic, stylized action sequence in a bamboo forest with high visual and thematic fidelity.
failure: none
media contract: 12.288s actual / 12.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 186.6s generation
ref2va-sample-2026083500-000013

ref2va-sample-2026083032-000017

Create an 8-second realistic video set in the reception area from Picture 3, where the woman from Picture 1 stands by the wooden desk and interacts non-verbally with the woman from Picture 2 as she walks up to the counter.
multi_subject_interactiondirect ablation8sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
Generate an 8-second realistic video set in the reception area from <Picture 3>. The woman from <Picture 1>, dressed in her hat, grey long-sleeve top, and jeans, stands behind the wooden reception desk. The woman from <Picture 2>, wearing her patterned sleeveless dress and ponytail, walks up to the counter. They exchange a polite smile, a greeting nod, and an inviting gesture toward the desk under the warm ambient lighting of <Picture 3>.
Final delivered output
REJECTED
While visual references and scene layout are integrated well, the video generates corrupted and unintelligible spoken audio instead of clean non-verbal interaction.
failure: unintelligible_speech
media contract: 8.032s actual / 8.0s target · audio present · pass
primary_transformation: 3source_preservation: 3target_identity: 3source_leakage: 2motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 independent output QA · 183.6s generation
ref2va-sample-2026083500-000014

ref2va-sample-2026083032-000018

Animate the long-haired man in the dark room from the image so that he talks directly to the camera, transferring the exact head movements, expressions, and speaking rhythm from the reference video while keeping his original face, outfit, and background.
motion_performance_transferdirect ablation8sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 8.0s
Exact H3 prompt
Generate an 8-second video animating <Picture 1> by applying the head motion, talking performance, and facial expressions from <Video 1>. Preserve the character identity, hair, black t-shirt, and dark room background from <Picture 1>, completely excluding the identity, clothing, and living room setting from <Video 1>.
Final delivered output
REJECTED
The generated video effectively transfers the motion performance onto the target identity, but the spoken audio is unintelligible gibberish, requiring a rejection.
prompt drift: audible speech lacks exact dialogue text or an audio reference
failure: unintelligible_speech
media contract: 8.032s actual / 8.0s target · audio present · pass
primary_transformation: 3source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 independent output QA · 320.1s generation
ref2va-sample-2026083500-000016

ref2va-sample-2026083032-000006

Create a 10-second realistic video showing the man from Picture 1 and the woman from Picture 2 interacting politely in the reception office from Picture 3 using natural non-verbal body language.
multi_subject_interactiondirect ablation10sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
In the reception office setting of <Picture 3>, the man from <Picture 1> stands beside the tall wooden front desk as the woman from <Picture 2> walks into the frame holding her black clutch. She pauses near the counter, and the two share an attentive, silent professional exchange with subtle nods and open-hand gestures under warm indoor lighting.
Final delivered output
REJECTED
Visual identities and environment from the input references are recreated reasonably well, but the output contains garbled, unintelligible spoken audio contrary to the silent exchange requirement.
failure: unintelligible_speech, unrequested_change
media contract: 10.176s actual / 10.0s target · audio present · pass
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 independent output QA · 232.6s generation