H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-30 · no auto refresh

Round 3 · Official Format A/B

Five baseline-success combinations with matched seeds: direct requests vs official six-section prompts. Gemini 3.7 prompt QA approved all five pairs.

Core 8 / 8 · pair coverage 4 / 4
ref2va-sample-2026083005-000000

ref2va-sample-2026082240-004277

Create an 8-second video of the woman in the picture walking through a sunny modern plaza, pausing next to a fountain, and smiling at the camera.
character_subjectdirect ablation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate an 8-second video featuring the woman from <Picture 1>. She walks gracefully across a sunlit modern plaza, stops near a sleek stone fountain, and turns to smile warmly toward the camera.
Final delivered output
ACCEPTED
The generated video accurately follows the prompt instructions, faithfully maintaining the subject identity and fulfilling all requested actions.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 161.9s generation
ref2va-sample-2026083005-000001

ref2va-sample-2026082240-004277

Create an 8-second video of the woman in the picture walking through a sunny modern plaza, pausing next to a fountain, and smiling at the camera.
character_subjectContext-IR approximation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman from <Picture 1>, retaining her facial identity, hairstyle, patterned dress, and footwear.

summary:
[reference generation] The target video follows <Subject 1> as she walks through a bright, modern urban plaza on a sunny afternoon, pauses beside a contemporary fountain, and looks toward the camera with a gentle, confident smile.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the woman's facial identity, hairstyle, outfit, and overall visual appearance from <Picture 1> are faithfully retained throughout.

detailed_description:
The target video is captured in a bright, elegant commercial cinematic style with warm afternoon sunlight and subtle handheld camera movement.
[Shot 1] The scene opens in a wide shot within an expansive modern urban plaza paved with sleek granite tiles and bordered by glass architecture. <Subject 1>, wearing her outfit from <Picture 1>, walks gracefully forward along the open walkway toward the camera. Warm sunlight casts soft shadows behind her as she takes poised, steady steps. A gentle breeze lightly moves her hair. The camera smoothly tracks backward while keeping her centered in a fluid medium-wide composition.
[Shot 2] At 00:04.500, the shot cuts to a medium close-up of <Subject 1> coming to a gentle halt near a sleek stone fountain. She turns slightly, letting her gaze sweep across the surrounding plaza before turning her eyes directly toward the lens. Her expression relaxes into a soft, genuine smile as specular highlights from the fountain water sparkle softly in the warm background bokeh. The camera slowly pushes in slightly on her face until the final frame.

overall_soundscape:
Gentle outdoor urban plaza ambience, soft flowing water sounds from the fountain, and rhythmic footsteps on stone pavement.

non_diegetic_music:
An airy, uplifting acoustic instrumental featuring soft guitar strumming and gentle piano accents, maintaining a warm and elegant atmosphere throughout.
Final delivered output
ACCEPTED
The generated video faithfully preserves the subject's identity and styling while executing all requested actions in the specified setting.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 116.3s generation
ref2va-sample-2026083005-000002

ref2va-sample-2026082240-007682

Make a 12-second commercial showcasing this baby cream tube on a bright pastel tabletop, showing someone opening the cap and dispensing a drop of cream.
product_objectdirect ablation12sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
Create a 12-second commercial video featuring the baby cream tube from <Picture 1>. Shot in a bright, soft pastel studio setting, a person picks up the tube to showcase the packaging, unscrews the cap to dispense a small drop of soothing cream onto a fingertip, and gently blends it onto the skin.
Final delivered output
ACCEPTED
The generated video successfully showcases the reference product in the requested commercial setting, executing all product interaction steps cleanly.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 199.3s generation
ref2va-sample-2026083005-000003

ref2va-sample-2026082240-007682

Make a 12-second commercial showcasing this baby cream tube on a bright pastel tabletop, showing someone opening the cap and dispensing a drop of cream.
product_objectContext-IR approximation12sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the product tube from <Picture 1>.

summary:
[reference generation] A 12-second commercial product presentation highlighting <Subject 1> on a clean, sunlit pastel tabletop, showing hands gently unscrewing its cap and dispensing a small drop of soothing cream.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the visual identity, form, labeling, and dimensions of the product tube from <Picture 1> are faithfully maintained across all shots.

detailed_description:
The video is styled as a soft, bright, high-end skincare commercial with clean studio lighting, pastel tones, and smooth cinematic focus pulls.
[Shot 1] A close-up establishes <Subject 1> resting horizontally across a soft beige textured mat surrounded by subtle botanical accents and soft daylight. A woman's manicured hand smoothly enters the frame from the right, picking up <Subject 1> and gently tilting it upward toward the lens to showcase the front packaging clearly as the camera executes a slight slow-push forward.
[Shot 2] At 00:04.500, the shot cuts to an extreme close-up of the user's hands twisting open the white screw cap of <Subject 1>. Soft physical clicks and friction sounds accompany the unscrewing action. Once the cap is removed, the hands gently squeeze the flexible body of <Subject 1>, dispensing a small, smooth pearl of white soothing cream onto an index fingertip.
[Shot 3] At 00:08.500, the shot cuts to a medium close-up from an over-the-shoulder angle. The fingertip dabs and softly blends the white cream onto the back of a hand in smooth, circular moisturizing motions, showing its light texture absorbing cleanly into the skin. <Subject 1> stands upright in soft focus in the background on the sunlit table as the shot glides gently to a halt.

overall_soundscape:
Subtle studio room tone, soft plastic cap unscrewing clicks, delicate cream dispensing texture sounds, and gentle skin-rubbing foley.

non_diegetic_music:
A gentle, warm, and uplifting acoustic guitar and ambient chime commercial track playing softly throughout the 12 seconds.
Final delivered output
ACCEPTED
The generated video accurately presents the referenced product tube across all shots with high visual fidelity, seamless motion, and appropriate audio accompaniment.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 193.1s generation
ref2va-sample-2026083005-000004

ref2va-sample-2026082240-000145

Continue the video from where it leaves off, showing the crafter finishing gluing the gold bead garland onto the centerpiece stand.
continuation_next_shotdirect ablation12sgeneratedaccepted
<Video 1>continuation_context
video · 5.0s
Exact H3 prompt
Continue directly from the final frame of <Video 1> as the crafter's hands (<Subject 2>) finish pressing the golden bead strand into the top slot of the gold centerpiece (<Subject 1>) with the craft stick, then rotate the stand slightly and glue the garland along the remaining arms.
Final delivered output
ACCEPTED
The generated continuation smoothly follows the prompt, preserving context, subject consistency, and continuous crafting actions.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 309.5s generation
ref2va-sample-2026083005-000005

ref2va-sample-2026082240-000145

Continue the video from where it leaves off, showing the crafter finishing gluing the gold bead garland onto the centerpiece stand.
continuation_next_shotContext-IR approximation12sgeneratedaccepted
<Video 1>continuation_context
video · 5.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the gold centerpiece stand and bead garland from <Video 1>.
<Subject 2> is the crafter's hands with purple glitter manicure from <Video 1>.
<Video 1> is the source video providing continuation context.

summary:
[video continuation] The video continues directly from the final frame of <Video 1>, showing <Subject 2> securing the golden bead garland onto <Subject 1> with glue and a craft stick.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the gold centerpiece stand and beaded garlands remain identical to <Video 1>.
<Subject 2> (appears in [Shot 1]): fully_preserved - the crafter's hands, movements, and purple glitter manicure continue seamlessly from <Video 1>.
<Video 1> (continuation source): fully_preserved - the tabletop environment, lighting, and crafting action resume without interruption.

detailed_description:
The scene maintains a clean, bright tabletop DIY tutorial aesthetic.
[Shot 1] Continuing seamlessly from the final frame of <Video 1>, the camera holds a steady medium close-up shot of <Subject 1> on the white table. <Subject 2> uses the wooden craft stick to press the golden bead garland firmly into the top adhesive joint of <Subject 1>. <Subject 2> then slightly turns <Subject 1>, dispenses a neat line of hot glue along the adjacent upper curve, and presses the next section of beads securely into place. The crafter gently wipes away stray glue threads, ensuring the garland hangs evenly across each arm of the centerpiece.

overall_soundscape:
Subtle craft room ambience with soft plastic bead clicks, glue gun triggers, and light wooden craft stick taps on the table.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The generated continuation naturally extends the crafting action with good visual and motion consistency.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 271.3s generation
ref2va-sample-2026083005-000006

ref2va-sample-2026082240-007247

Create a 12-second anime video where the girl in the pink gown descends the grand staircase to meet the girl in the green dress, who presents her with a four-leaf clover.
multi_subject_interactiondirect ablation12sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
Exact H3 prompt
Generate a 12-second anime video featuring <Subject 1> from <Picture 1> and <Subject 2> from <Picture 2>. <Subject 2> in her pink gown walks down the red-carpeted grand staircase to greet <Subject 1>, who approaches in her pale green dress and presents a four-leaf clover. Warm indoor palace lighting with drifting petals.
Final delivered output
ACCEPTED
The generated video accurately depicts both characters and executes the requested staircase meeting and clover handoff with consistent style and animation.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 247.9s generation
ref2va-sample-2026083005-000007

ref2va-sample-2026082240-007247

Create a 12-second anime video where the girl in the pink gown descends the grand staircase to meet the girl in the green dress, who presents her with a four-leaf clover.
multi_subject_interactionContext-IR approximation12sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the anime character in <Picture 1>, featuring short brown hair, purple eyes, a flower crown, and a ruffled pale green and white dress.
<Subject 2> is the anime character in <Picture 2>, featuring orange hair, blue eyes, and an ornate pink gown on a grand staircase.

summary:
[reference generation] A 12-second anime video showing <Subject 2> descending a grand staircase to greet <Subject 1>, who offers a four-leaf clover.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - character appearance and pale green outfit are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - character appearance, pink gown, and palace staircase setting are retained.

detailed_description:
The video is rendered in a vibrant anime style with soft lighting and drifting pink flower petals.
[Shot 1] A medium-wide shot shows <Subject 2> walking gracefully down the red-carpeted staircase. At the base, <Subject 1> stands smiling while holding out a four-leaf clover.
[Shot 2] At 00:06.000, the shot cuts to a medium two-shot at the foot of the stairs. <Subject 2> steps forward cheerfully and reaches out toward <Subject 1>, who beams as petals drift around them.

overall_soundscape:
Gentle footsteps on carpet, rustling fabric, and palace room acoustics.

non_diegetic_music:
An uplifting, light orchestral anime melody with flute, strings, and piano.
Final delivered output
ACCEPTED
The generated video accurately depicts both characters interacting on the grand staircase with strong identity consistency, smooth animation, and appropriate audio.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 239.6s generation
ref2va-sample-2026083005-000008

ref2va-sample-2026082240-002105

Change the reporter's red polo shirt in the video to navy blue, keeping everything else the same.
source_preserving_editingdirect ablation10ssource_soundtrackaccepted
<Video 1>source_edit
video · 10.0s
<Audio 1>source_soundtrack
audio · 10.0s
Exact H3 prompt
Edit <Video 1> to change <Subject 1>'s red polo shirt to a navy blue polo shirt, retaining his identity, motion, the background, and the synchronized speech from <Audio 1>.
Final delivered output
ACCEPTED
The requested modification to change the reporter's polo shirt from red to navy blue was executed cleanly with high fidelity to the source motion, identity, background, and audio track.
failure: none
primary_transformation: 5source_preservation: 5target_identity: 5source_leakage: 1motion_or_temporal: 5audio: 5artifact: 5
Gemini 3.7 independent output QA · 372.4s generation
ref2va-sample-2026083005-000009

ref2va-sample-2026082240-002105

Change the reporter's red polo shirt in the video to navy blue, keeping everything else the same.
source_preserving_editingContext-IR approximation10ssource_soundtrackaccepted
<Video 1>source_edit
video · 10.0s
<Audio 1>source_soundtrack
audio · 10.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the male reporter in <Video 1>.
<Video 1> is the source video for the target video edit.
<Audio 1> is the synchronized audio track of <Video 1>.

summary:
[video editing + audio reuse] The target video is an edited version of <Video 1>, replacing <Subject 1>'s red polo shirt with a navy blue polo shirt while retaining all motions, background, and audio.

retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - identity and actions preserved, shirt color changed to navy blue.
<Video 1> (appears in [Shot 1]): partially_preserved - overall frame, timing, and background retained with shirt edit.
<Audio 1>: fully_copy - reused 1:1 as the complete audio track.

detailed_description:
The target video maintains the realistic news broadcast look in natural daylight.
[Shot 1] Modifying <Video 1>, <Subject 1> (S1) wears a navy blue polo shirt instead of red. Standing in the stadium parking lot, he speaks to the camera with synchronized gestures matching <Audio 1>: <d>[Spanish] ...desde el Estadio Azteca por donde entrará la porra de los visitantes. Se calcula que habrá 3,000 efectivos los que se encargarán de garantizar la seguridad de los más de 80,000 asistentes esta noche al Estadio Azteca.</d> Camera framing and background remain fixed.

overall_soundscape:
Copied ambient noise and wind from <Audio 1> continue throughout.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The edit accurately transforms the polo shirt color to navy blue while maintaining subject motion, background environment, and audio sync with minimal artifacts.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 371.3s generation