H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-30 · no auto refresh

Promotion 20 · Matched A/B

Twenty semantically accepted real-workload combinations. Direct and Context-IR-style prompts use identical references, duration, seed, and optimized 4xGB200 runtime.

Generating 11 / 40 · QA 7/11 accepted · both accepted 2, direct only accepted 1, incomplete 1, neither accepted 1, structured only accepted 1 · task gates: A1 100% (1 pairs, insufficient evidence), B3 100% (1 pairs, insufficient evidence), D2 0% (1 pairs, insufficient evidence), D4 100% (2 pairs, insufficient evidence)
ref2va-sample-2026083400-000000

ref2va-sample-2026083032-000000

Animate the woman from the photo to talk and gesture with the exact body language from the video while keeping her in the same kitchen.
motion_performance_transferdirect ablation10sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 10.0s
Exact H3 prompt
Generate a 10-second video of the woman in <Picture 1> standing in her kitchen, animating her speech, head movements, and hand gestures to match the performance of <Video 1>.
Final delivered output
ACCEPTED
The performance and gesture transfer from Video 1 onto the subject and environment in Picture 1 is executed successfully with high visual fidelity.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 408.9s generation
ref2va-sample-2026083400-000001

ref2va-sample-2026083032-000000

Animate the woman from the photo to talk and gesture with the exact body language from the video while keeping her in the same kitchen.
motion_performance_transferContext-IR approximation10sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 10.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the blonde woman in <Picture 1> wearing a black jacket and patterned top, positioned in her modern kitchen behind the counter.
<Subject 2> is the speaking and gesturing performance from <Video 1>.

summary:
[reference generation] The target video animates <Subject 1> using the motion and gesturing performance of <Subject 2> from <Video 1>, preserving <Subject 1>'s appearance and kitchen background.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - visual identity, clothing, and kitchen setting from <Picture 1> are retained.
<Subject 2> (appears in [Shot 1]): attribute_transfer - gesture timing, head motion, and speaking cadence are transferred onto <Subject 1>.
<Video 1> (motion structure): attribute_transfer - controls the 10-second performance timing.

detailed_description:
The scene is captured in a realistic, clear indoor documentary style.
[Shot 1] A stationary medium shot shows <Subject 1> behind the kitchen counter with beverage containers visible in front. Transferring the performance of <Subject 2> from <Video 1>, she talks directly toward the camera, making emphatic hand gestures, raising her palms, and shifting posture naturally throughout the 10-second duration while maintaining her exact visual identity and environment.

overall_soundscape:
overall_soundscape: Clean, quiet indoor room ambiance with natural spoken acoustics.

non_diegetic_music:
non_diegetic_music: N/A
Final delivered output
REJECTED
The motion performance from Video 1 is successfully transferred onto the subject from Picture 1 while maintaining identity and scene continuity.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 3
Gemini 3.7 independent output QA · 406.1s generation
ref2va-sample-2026083400-000002

ref2va-sample-2026083032-000001

Continue the video from where it leaves off: have the woman put down the powder brush and pick up a finer detailing brush to start outlining her eyebrows and eye contours in the mirror.
continuation_next_shotdirect ablation10sgeneratedaccepted
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
Continue directly from the final frame of <Video 1>. In <Subject 2>, <Subject 1> sets down the powder applicator on the counter, examines her white face paint in the mirror, and picks up a slim detailing brush to begin working on her eyebrows and eye contours under consistent vanity lighting.
Final delivered output
ACCEPTED
The generated continuation faithfully preserves the subject, outfit, and environment while accurately executing the multi-step request to set down the powder brush and apply finer eyebrow and eye makeup with a slim detailing brush.
failure: none
primary_transformation: 5source_preservation: 5target_identity: 5source_leakage: 1motion_or_temporal: 5audio: 4artifact: 5
Gemini 3.7 independent output QA · 372.5s generation
ref2va-sample-2026083400-000003

ref2va-sample-2026083032-000001

Continue the video from where it leaves off: have the woman put down the powder brush and pick up a finer detailing brush to start outlining her eyebrows and eye contours in the mirror.
continuation_next_shotContext-IR approximation10sgeneratedaccepted
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman from <Video 1>, wearing a purple shirt with white base makeup applied to her face.
<Subject 2> is the bathroom vanity environment with large mirrors and glass block walling from <Video 1>.
<Video 1> is the continuation source video.

summary:
[video continuation] The target video continues seamlessly from the end of <Video 1>, showing <Subject 1> in <Subject 2> setting down her powder brush and beginning detail work with a fine cosmetic tool.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the woman's identity, purple shirt, pulled-back hair, and white base paint are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the bathroom interior, mirrors, lighting, and layout are retained.
<Video 1> (continuation source): fully_preserved - the action continues directly from the final frame and state of <Video 1>.

detailed_description:
The target video maintains the realistic domestic tutorial framing and soft indoor lighting of <Video 1>.
[Shot 1] Continuing seamlessly from the final state of <Video 1>, <Subject 1> stands before the mirror in <Subject 2> with her face covered in white base paint. She lowers her hand to set the large brush onto the vanity counter, leans forward slightly to inspect the evenness of the white coverage in the mirror, and picks up a slender detail brush. Looking steadily into the reflection, she begins carefully sketching along her brow line. The camera remains in a steady medium close-up throughout.

overall_soundscape:
Quiet indoor room tone with the subtle click of a cosmetic tool placed on the counter and faint fabric movement.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The generated video successfully continues the scene, preserving character appearance and environment while correctly carrying out the requested tool change and makeup detailing.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 independent output QA · 371.0s generation
ref2va-sample-2026083400-000004

ref2va-sample-2026083032-000002

Change the comedian's black suit jacket to a navy blue suit jacket while keeping his performance, stage setting, and camera angle identical to the original video.
source_preserving_editingdirect ablation5sgeneratedrejected
<Video 1>source_edit
video · 5.0s
Exact H3 prompt
In <Video 1>, change the comedian's black suit jacket to a navy blue suit jacket, preserving his facial identity, speaking performance, body movements, microphone handling, stage background, lighting, and camera framing.
Final delivered output
REJECTED
The suit jacket was successfully recolored to navy blue, but camera framing, motion performance, and original audio were not preserved.
failure: source_structure_lost, motion_transfer_failed, audio_missing_or_changed, unrequested_change
primary_transformation: 4source_preservation: 2target_identity: 3source_leakage: 2motion_or_temporal: 2audio: 2artifact: 4
Gemini 3.7 independent output QA · 115.5s generation
ref2va-sample-2026083400-000005

ref2va-sample-2026083032-000002

Change the comedian's black suit jacket to a navy blue suit jacket while keeping his performance, stage setting, and camera angle identical to the original video.
source_preserving_editingContext-IR approximation5sgeneratedrejected
<Video 1>source_edit
video · 5.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the source video for the target video edit, featuring a male stand-up comedian speaking into a microphone on stage.
<Subject 1> is the comedian's suit jacket from <Video 1>, edited from black to navy blue.

summary:
[video editing] The target video is an edited version of <Video 1>, modifying the comedian's black suit jacket to a navy blue suit jacket while fully preserving his identity, gestures, stage background, lighting, and camera framing.

retention_analysis:
<Video 1> (visual performance, stage environment, camera framing, and timing): fully_preserved - the performer's facial expressions, speech gestures, microphone handling, warm ambient stage lighting, background decor, and static medium shot are fully retained.
<Subject 1> (comedian's suit jacket): partially_preserved - the tailored cut, texture, and movement of the jacket are preserved, but its color is changed from solid black to navy blue.

detailed_description:
The target video preserves the visual style and stage presentation of <Video 1>.
[Shot 1] The medium shot from <Video 1> is maintained, showing the male stand-up comedian standing on stage against the warm, dimly lit background with decorative light fixtures. He holds a microphone in his hand while speaking. His suit jacket (<Subject 1>) is rendered in a clean navy blue color instead of black, keeping the original fabric texture, tailoring, lapels, and natural creasing as he shifts and gestures. All of his body movements, head orientation, facial expressions, and the static camera framing remain identical to <Video 1>.

overall_soundscape:
Ambient comedy club room tone and quiet stage acoustics.

non_diegetic_music:
N/A
Final delivered output
REJECTED
While the suit jacket color was successfully altered to navy blue, the output failed to retain the source camera angle, framing, motion, and audio synchronization.
failure: source_structure_lost, unrequested_change, audio_missing_or_changed
primary_transformation: 3source_preservation: 2target_identity: 3source_leakage: 2motion_or_temporal: 3audio: 2artifact: 4
Gemini 3.7 independent output QA · 115.8s generation
ref2va-sample-2026083400-000006

ref2va-sample-2026083032-000003

Create an 8-second animated video of the anime character in <Picture 1> performing a lively idol dance pose and greeting the camera on a brightly lit concert stage.
character_subjectdirect ablation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate an 8-second 2D anime video featuring <Subject 1>, the blonde anime idol from <Picture 1>, performing an energetic stage greeting and idol dance sequence on a concert stage illuminated by colorful spotlights.
Final delivered output
ACCEPTED
The generated video successfully animates the reference character performing on a concert stage with matching visual style and coherent motion.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 118.0s generation
ref2va-sample-2026083400-000007

ref2va-sample-2026083032-000003

Create an 8-second animated video of the anime character in <Picture 1> performing a lively idol dance pose and greeting the camera on a brightly lit concert stage.
character_subjectContext-IR approximation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the anime idol girl in <Picture 1>, featuring short blonde hair tied in a side ponytail with a pink accessory, an orange-amber eye, and a white, red plaid, and pink stage idol uniform.

summary:
[reference generation] The 8-second target video shows <Subject 1> performing a dynamic greeting and idol dance routine on an illuminated concert stage with colorful stage lights and soft particle effects.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the character's facial features, blonde hairstyle, eye color, and detailed stage idol costume from <Picture 1> are faithfully retained.

detailed_description:
The target video is rendered in a vibrant high-quality Japanese anime art style with clean line work and crisp cel-shaded lighting.
[Shot 1] A medium full shot opens on a vibrant concert stage bathed in pulsing pink and gold spotlights with gentle floating lens flare particles. <Subject 1>, the blonde anime girl in her stage costume from <Picture 1>, stands centered facing the audience. She winks with a radiant smile, raises her left hand in a cheerful greeting wave, and then transitions smoothly into a rhythmic dance bounce, swaying her hips slightly as her blonde side ponytail and tiered skirt sway naturally with her movement. The camera tracks slowly inward while rotating subtly to emphasize her energetic performance.
[Shot 2] At 00:04.500, the shot cuts to a dynamic close-up medium shot of <Subject 1>. She steps forward with an open smile, pointing affectionately toward the viewer with both hands forming a quick finger-heart gesture before finishing with a charming pose, her eyes sparkling under the warm stage spotlights.

overall_soundscape:
Ambient live concert atmosphere with soft stage reverberation and crowd excitement.

non_diegetic_music:
An upbeat, fast-tempo Japanese pop idol instrumental featuring energetic synth melodies, light brass accents, and a driving dance rhythm.
Final delivered output
ACCEPTED
The generated video successfully animates the reference character on a concert stage with accurate visual identity, lively choreography, and fitting background music.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 113.0s generation
ref2va-sample-2026083400-000008

ref2va-sample-2026083032-000004

Continue the video as the artist finishes painting around the leaf stencil and carefully peels it off the paper to reveal the design.
continuation_next_shotdirect ablation10sgeneratedrejected
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
Continue seamlessly from the end of <Video 1>. The artist carefully grasps the edge of the blue-painted leaf stencil and lifts it off the white paper, revealing the crisp botanical print silhouette against the rustic wooden table under steady natural lighting.
Final delivered output
REJECTED
The continuation partially fulfills the request by lifting the stencil, but it restarts earlier actions and exhibits an unrealistic disappearance of surrounding paint upon peeling.
failure: continuation_boundary_failed, unrequested_change
primary_transformation: 3source_preservation: 3target_identity: 3source_leakage: 1motion_or_temporal: 3audio: 3artifact: 3
Gemini 3.7 independent output QA · 292.5s generation
ref2va-sample-2026083400-000009

ref2va-sample-2026083032-000004

Continue the video as the artist finishes painting around the leaf stencil and carefully peels it off the paper to reveal the design.
continuation_next_shotContext-IR approximation10sgeneratedaccepted
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the continuation source video showing hands painting over a leaf stencil with bright blue acrylic paint on paper over a rustic wooden surface.

summary:
[video continuation] The target video continues directly from the final frame of <Video 1>, showing the artist completing the painting and peeling the stencil away from the paper.

retention_analysis:
<Video 1> (continuation anchor): fully_preserved - the visual composition, paper placement, rustic wooden surface, blue paint, and hands continue seamlessly from the end state of <Video 1>.

detailed_description:
The target video maintains the realistic top-down craft demonstration style with natural outdoor daylight.
[Shot 1] Continuing seamlessly from the final frame of <Video 1>, where the leaf stencil rests flat coated in rich cobalt blue paint on white paper, the artist's hand reaches in with steady fingers. Grasping the dry tip or stem edge of the leaf, they gently peel it upwards and away from the paper surface. As the stencil lifts, crisp outlines and subtle vein textures of the botanical silhouette are revealed in bold blue against the clean white paper. The artist sets the used stencil aside, revealing the completed painted print on the paper taped to the weathered wooden table. The top-down camera remains steady throughout.

overall_soundscape:
Subtle outdoor ambient room tone, soft rustling of paper, and the light peel of the stencil lifting from wet paint.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The continuation smoothly connects with the context video, successfully executing the painting completion and stencil removal sequence with consistent visual fidelity.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 1motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 293.6s generation
ref2va-sample-2026083400-000010

ref2va-sample-2026083032-000005

Create a 12-second commercial showcase video featuring the cosmetic tube from Picture 1, displaying it smoothly in an elegant studio environment.
product_objectdirect ablation12sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
Generate a 12-second high-end commercial showcase video for the cosmetic squeeze tube in <Picture 1>. Place <Subject 1> in a minimalist aesthetic studio setting with soft studio lighting and delicate water rippling reflections. The camera smoothly orbits around the product to display its light-green matte finish and crystalline cap from multiple angles, ending on a crisp hero shot.
Final delivered output
ACCEPTED
The model successfully creates a polished commercial video showcasing the cosmetic tube from Picture 1 in an elegant studio setup with high visual fidelity.
failure: none
media contract: 12.288s actual / 12.0s target · audio present · pass
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 independent output QA · 223.4s generation
ref2va-sample-2026083400-000011

ref2va-sample-2026083032-000005

Create a 12-second commercial showcase video featuring the cosmetic tube from Picture 1, displaying it smoothly in an elegant studio environment.
product_objectContext-IR approximation12sgeneratedpending
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the cosmetic squeeze tube in <Picture 1>, featuring a pastel green matte body and a translucent crystalline cap.

summary:
[reference generation] The target video is a 12-second premium commercial product showcase featuring <Subject 1> rotating smoothly on a sleek minimalist pedestal in an elegant studio environment with soft lighting and subtle water ripple caustics.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the pastel green tube body, rectangular shape, crimped top seam, and translucent textured cap are fully retained.

detailed_description:
The target video is filmed in a pristine, high-end commercial aesthetic with soft cinematic studio illumination and clean pastel color grading.
[Shot 1] The scene opens on a sleek cylindrical display pedestal set against a soft beige and warm cream minimalist studio background. <Subject 1>, the pastel green cosmetic squeeze tube with its translucent crystalline cap resting upright on the platform, stands centered in the frame. Gentle caustic light reflections from nearby rippling water dance across the surface of the tube. The camera performs a smooth, continuous orbital arc around <Subject 1>, gracefully showcasing its seamless matte green finish, clean crimped top edge, and sparkling cap texture. At 00:06.000, the camera slowly tightens into a closer three-quarter hero angle, capturing soft specular highlights gliding across the curved container surface. As the camera completes its elegant glide, subtle rising mist drifts past the base of the pedestal, settling into a stable, polished commercial hero shot through 00:12.000.

overall_soundscape:
Subtle ambient studio room tone with gentle, delicate water ripple accents and a faint soft swoosh accompanying the camera movement.

non_diegetic_music:
An airy, sophisticated ambient electronic track with soft synth pads, delicate chime melodies, and a gentle rhythmic pulse that builds subtly before fading softly at the end.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000012

ref2va-sample-2026083032-000007

Change the young woman's red jacket to a navy blue jacket, keeping everything else in the press conference video exactly the same.
source_preserving_editingdirect ablation8sgeneratedpending
<Video 1>source_edit
video · 8.0s
Exact H3 prompt
In <Video 1>, change the red athletic jacket worn by the speaking woman to a navy blue jacket while fully preserving her identity, braided hairstyle, speaking movements, facial expressions, microphone setup, blue backdrop, and camera framing across the full 8 seconds.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000013

ref2va-sample-2026083032-000007

Change the young woman's red jacket to a navy blue jacket, keeping everything else in the press conference video exactly the same.
source_preserving_editingContext-IR approximation8sgeneratedpending
<Video 1>source_edit
video · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the young blonde woman speaking at the press conference in <Video 1>.
<Video 1> is the source video for the target video edit.

summary:
[video editing] The target video is an edited version of <Video 1>, changing the red jacket worn by <Subject 1> to navy blue while keeping all other visual elements identical.

retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - her identity, braided hairstyle, facial expressions, and speaking actions are preserved, while her red jacket is changed to navy blue.
<Video 1> (timing, framing, and background): fully_preserved - the static medium close-up framing, branded backdrop, microphone, and lower-third graphics are retained.

detailed_description:
The target video maintains the broadcast press conference style, lighting, and static composition of <Video 1>.
[Shot 1] The static medium close-up from <Video 1> shows <Subject 1>, the young blonde woman with braided pigtails speaking into the microphone in front of the blue branded backdrop with lower-third text overlay. The red athletic jacket worn by <Subject 1> in <Video 1> is replaced with a navy blue athletic jacket of the same cut, retaining all original embroidery details and contours. Her facial expressions, head movements, and speech delivery precisely follow <Video 1> throughout the 8-second shot.

overall_soundscape:
Press room room tone with light background ambience.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000014

ref2va-sample-2026083032-000008

Animate the man in <Picture 1> so that his speaking performance, head movements, and facial expressions match the woman speaking in <Video 1>, keeping his appearance, clothes, and studio setting from the image.
motion_performance_transferdirect ablation5sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 5.0s
Exact H3 prompt
Generate a 5-second video animating the man from <Picture 1> in his studio setup, replicating the speaking motion, head nods, facial expressions, and timing from <Video 1> while retaining the identity, clothing, microphone, and background environment of <Picture 1>.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000015

ref2va-sample-2026083032-000008

Animate the man in <Picture 1> so that his speaking performance, head movements, and facial expressions match the woman speaking in <Video 1>, keeping his appearance, clothes, and studio setting from the image.
motion_performance_transferContext-IR approximation5sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 5.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the man from <Picture 1> seated behind his broadcast microphone in a memorabilia-filled studio, wearing headphones and a grey polo shirt.
<Subject 2> is the speaking performance and head motion from <Video 1>, featuring subtle downward glances, measured head nods, and articulate speech delivery.

summary:
[reference generation] The target video animates <Subject 1> in his studio environment, performing the radio news reading and speaking motion of <Subject 2> over 5 seconds.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the man's identity, grey polo shirt, headphones, desk microphone, and memorabilia-filled studio background are preserved from <Picture 1>.
<Subject 2> (motion performance): attribute_transfer - the mouth movements, head articulation, downward glances, and cadence from <Video 1> are transferred to <Subject 1>.

detailed_description:
The video has a realistic broadcast studio aesthetic with warm interior lighting.
[Shot 1] A stationary medium close-up shot frames <Subject 1>, the man from <Picture 1> wearing silver-and-white over-ear headphones and a grey collared polo shirt, seated behind a black broadcast microphone with his eclectic background of figurines, framed art, and brick walls. Throughout the 5 seconds, <Subject 1> articulates speech into the microphone, transferring the performance of <Subject 2> from <Video 1>. He glances down slightly as if reading copy, tilts and nods his head with measured cadence, and delivers his lines with natural lip and jaw movements corresponding to the pacing of <Video 1>. The camera remains static with stable depth of field focused on the presenter.

overall_soundscape:
Ambient broadcast studio room tone with subtle quiet presence.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000016

ref2va-sample-2026083032-000009

Create a sleek 5-second product commercial showcasing the cream tube from <Picture 1> resting on an elegant marble vanity with soft studio lighting and smooth camera motion.
product_objectdirect ablation5sgeneratedpending
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
A sleek 5-second commercial product video featuring the cosmetic tube from <Picture 1> resting on a polished marble vanity bathed in soft pastel studio lighting. The camera performs a slow, smooth arc tracking around the product, highlighting its clean packaging and glossy surface reflections.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000017

ref2va-sample-2026083032-000009

Create a sleek 5-second product commercial showcasing the cream tube from <Picture 1> resting on an elegant marble vanity with soft studio lighting and smooth camera motion.
product_objectContext-IR approximation5sgeneratedpending
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the cosmetic cream tube in <Picture 1>, featuring a white tube body, a white cap, and a pink-and-white illustrated label.

summary:
[reference generation] A 5-second commercial product showcase featuring <Subject 1> resting on an elegant marble surface with soft studio lighting and smooth camera movement.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the tube packaging, white cap, and pink illustrated label design are preserved.

detailed_description:
The target video is rendered in a premium commercial product advertising style with soft, diffused cosmetic lighting and shallow depth of field.
[Shot 1] The scene opens on a polished white marble tabletop under soft, warm studio light with subtle pastel reflections. In the center, <Subject 1>, the cosmetic cream tube from <Picture 1> with its distinct white cap and pink-and-white label, rests at a gentle diagonal angle atop a minimalist stone riser. The camera executes a slow, cinematic arcing dolly shot from left to right around <Subject 1>, capturing subtle specular highlights gliding across the smooth plastic tube. In the background, soft out-of-focus pastel tones and diffused light create a clean, modern aesthetic.

overall_soundscape:
A clean, quiet studio ambience with a gentle, soft room tone.

non_diegetic_music:
A modern, airy ambient electronic track with warm acoustic synth pads and a gentle, uplifting beat.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000018

ref2va-sample-2026083032-000010

Continue this cooking clip by wrapping the bacon weave tightly around the stuffed meatloaf.
continuation_next_shotdirect ablation10sgeneratedpending
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
Continue directly from the final frame of <Video 1>, showing the person using their hands and parchment paper to roll the woven bacon lattice tightly around the stuffed meatloaf until it is fully enclosed.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000019

ref2va-sample-2026083032-000010

Continue this cooking clip by wrapping the bacon weave tightly around the stuffed meatloaf.
continuation_next_shotContext-IR approximation10sgeneratedpending
<Video 1>continuation_context
video · 10.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the continuation source video showing a stuffed ground beef log placed on a woven bacon lattice.

summary:
[video continuation] Continue from the end of <Video 1>, showing the person rolling the bacon weave completely around the stuffed meatloaf log.

retention_analysis:
<Video 1> (continuation source): fully_preserved - the tabletop environment, checkered tablecloth, tattooed arms, and stuffed meatloaf on the bacon lattice continue directly from the final frame.

detailed_description:
The target continuation maintains a natural kitchen demonstration style with overhead medium-close framing and warm indoor lighting.
[Shot 1] Continuing seamlessly from the final frame of <Video 1>, the person grasps the edges of the parchment paper beneath the bacon weave and begins rolling the woven bacon tightly around the stuffed ground beef log. Working from the top down, the person uses both hands to press and mold the bacon lattice firmly over the meat, tucking in the ends to ensure the mac-and-cheese filling remains securely sealed inside the bacon cylinder. The camera remains fixed in the overhead angle looking down at the checkered tablecloth.

overall_soundscape:
Subtle kitchen ambient tone with natural Foley sounds of parchment paper crinkling and meat being pressed and rolled.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000020

ref2va-sample-2026083032-000011

Create an 8-second video of the woman in the photo walking and posing along a runway stage under studio lighting.
character_subjectdirect ablation8sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate an 8-second video featuring <Subject 1>, the woman from <Picture 1> with her dark hair updo, peach hair flower, white t-shirt, brocade peplum mini skirt, and black wedge heels. She walks confidently forward down a fashion runway stage under soft spotlights, pauses to pose with a hand on her hip, and pivots gracefully.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000021

ref2va-sample-2026083032-000011

Create an 8-second video of the woman in the photo walking and posing along a runway stage under studio lighting.
character_subjectContext-IR approximation8sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman in <Picture 1>, with dark hair styled in a high updo with a peach-orange flower hairpiece, wearing a white short-sleeve V-neck t-shirt, a metallic brocade peplum mini skirt, and black platform wedge ankle-strap heels.

summary:
[reference generation] The target video shows <Subject 1> walking forward down a dark fashion runway, pausing to pose, and turning gracefully under soft stage lighting.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the woman's facial features, hairstyle with flower accessory, white top, brocade peplum skirt, and black wedge heels are faithfully preserved.

detailed_description:
The target video has a realistic fashion runway presentation style with clean stage lighting and soft spotlighting.
[Shot 1] A full-body shot captures <Subject 1>, the woman from <Picture 1> with her dark hair swept into an elegant high updo accented with an orange-peach flower, wearing her white t-shirt, brocade peplum mini skirt, and black wedge platform heels. She stands on a dark raised runway stage against a dark backdrop. She smiles with poised confidence, resting one hand on her hip before fluidly stepping forward down the runway toward the camera. Her steps are deliberate and rhythmic. As she reaches the end of the catwalk, she pauses, rests her opposite hand gently on her waist, and pivots smoothly, offering a confident glance back over her shoulder before turning to walk away into the soft stage illumination. The camera maintains a smooth medium-full tracking framing that slowly pans to keep her centered.

overall_soundscape:
Subtle ambient room resonance of an indoor auditorium and the muted rhythmic tap of heels on the wooden runway floor.

non_diegetic_music:
A steady, stylish mid-tempo electronic runway soundtrack with deep bass pulses and light synth chords.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000022

ref2va-sample-2026083032-000012

Create a smooth 3D product commercial showcasing the hand cream tube from <Picture 1> on an elegant pedestal with studio lighting.
product_objectdirect ablation10sgeneratedpending
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
Generate a 10-second 3D product commercial video featuring <Subject 1>, the hand cream tube from <Picture 1>. Place the hand cream on a smooth minimalist studio pedestal while the camera performs a smooth, elegant orbital sweep around it under soft luxury studio lighting.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000023

ref2va-sample-2026083032-000012

Create a smooth 3D product commercial showcasing the hand cream tube from <Picture 1> on an elegant pedestal with studio lighting.
product_objectContext-IR approximation10sgeneratedpending
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the tube of hand cream in <Picture 1>, featuring a pale yellow squeeze tube with a translucent cap and product labeling.

summary:
[reference generation] The 10-second target video is a 3D product showcase commercial displaying <Subject 1> on an elegant studio pedestal with soft ambient lighting.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the tube shape, pale yellow color, translucent cap, and packaging label details are retained.

detailed_description:
The target video is a polished 3D commercial product advertisement featuring clean studio lighting, shallow depth of field, and a minimalist luxury aesthetic.
[Shot 1] A medium close-up opens on <Subject 1>, the pale yellow hand cream tube from <Picture 1>, standing upright on a smooth circular pastel pedestal in a clean studio setting. Soft studio key lighting highlights the smooth matte surface of the tube and translucent cap. The camera performs a slow, elegant orbital sweep around <Subject 1> as gentle warm light glints subtly across the packaging. The background features a soft-focus clean studio gradient with warm pastel tones. The tube remains pristine and stable at the center of the frame throughout the smooth revolving camera movement until the shot gently concludes.

overall_soundscape:
A subtle, quiet indoor room tone with gentle studio air presence.

non_diegetic_music:
An elegant, airy ambient electronic background track with gentle synth pads and a soft, calm tempo.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000024

ref2va-sample-2026083032-000013

Create a 12-second realistic video where the woman in the terracotta t-shirt and denim shorts from Picture 1 meets the woman in the blue sequin fringe dress from Picture 2 at the reception desk shown in Picture 3, interacting warmly through gestures.
multi_subject_interactiondirect ablation12sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
A 12-second realistic video in a reception lobby (<Picture 3>). The woman from <Picture 1> in a terracotta t-shirt and denim shorts walks up to the counter where the woman from <Picture 2> in a royal blue sequin fringe dress is standing. In [Shot 1], the woman from <Picture 1> approaches, smiles, and nods politely; the woman from <Picture 2> turns and gestures warmly toward the counter with her gloved hand. In [Shot 2] at 00:06.500, a medium two-shot shows them looking at a brochure together on the counter, exchanging friendly nods and pointing gestures in a natural non-verbal interaction.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000025

ref2va-sample-2026083032-000013

Create a 12-second realistic video where the woman in the terracotta t-shirt and denim shorts from Picture 1 meets the woman in the blue sequin fringe dress from Picture 2 at the reception desk shown in Picture 3, interacting warmly through gestures.
multi_subject_interactionContext-IR approximation12sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the young woman in <Picture 1>, with shoulder-length wavy brown hair, wearing a terracotta V-neck t-shirt, distressed denim shorts, and white sneakers.
<Subject 2> is the woman in <Picture 2>, with styled dark hair and a hair accessory, wearing a royal blue sequined flapper dress with fringe, matching blue opera gloves, fishnet tights, and heels.
<Subject 3> is the reception lobby environment in <Picture 3>, featuring a front service counter, overhead warm pendant lighting, wood-paneled walls, and speckled tiled flooring.

summary:
[reference generation] The 12-second target video portrays <Subject 1> and <Subject 2> interacting non-verbally near the front desk counter inside <Subject 3> across two cinematic shots.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - facial features, hair, terracotta V-neck shirt, and distressed denim shorts are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - facial features, hairstyle, blue sequined fringe dress, blue gloves, and fishnets are retained.
<Subject 3> (appears in [Shot 1], [Shot 2]): fully_preserved - reception counter layout, paneled walls, and ambient indoor lighting are retained.

detailed_description:
The target video is filmed in a naturalistic cinematic style with warm indoor hotel lighting.
[Shot 1] A medium-wide shot establishes <Subject 3>, the reception lobby with its warm-toned service counter and wood-paneled walls. <Subject 2>, dressed in her royal blue sequined fringe dress and blue gloves, stands beside the counter looking at a travel booklet. <Subject 1>, wearing her terracotta t-shirt and denim shorts, enters from the left and walks toward the counter. <Subject 1> makes eye contact with <Subject 2> and offers a polite, friendly nod and smile. <Subject 2> looks up, reciprocates with a warm smile, and makes a welcoming open-palm gesture toward the desk. The camera gently tracks right with smooth, steady framing.
[Shot 2] At 00:06.500, the shot cuts to a medium two-shot at eye level. <Subject 1> and <Subject 2> stand side by side by the counter in <Subject 3>. <Subject 1> points toward a paper map on the counter with an inquiring expression. <Subject 2> leans in slightly, nodding affirmatively, and points with her gloved index finger to a specific spot on the map, gesturing clear directions. <Subject 1> nods with understanding and smiles gratefully, raising a hand in a brief gesture of thanks as the scene closes.

overall_soundscape:
Gentle indoor ambient room tone of a lobby, subtle rustling of paper, soft shoe scuffs on tiled floor, and quiet ventilation hum.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000026

ref2va-sample-2026083032-000014

In this video, change the woman's pink shirt to a royal blue shirt while keeping her identity, writing motion, desk setting, and timing exactly the same.
source_preserving_editingdirect ablation5sgeneratedpending
<Video 1>source_edit
video · 5.0s
Exact H3 prompt
Edit <Video 1> to change the pink collared shirt worn by the woman writing notes (<Subject 1>) to a royal blue shirt, while preserving her facial identity, writing action, desk environment, background, camera framing, cuts, and overall timing.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000027

ref2va-sample-2026083032-000014

In this video, change the woman's pink shirt to a royal blue shirt while keeping her identity, writing motion, desk setting, and timing exactly the same.
source_preserving_editingContext-IR approximation5sgeneratedpending
<Video 1>source_edit
video · 5.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the source video for the target video edit.
<Subject 1> is the young woman in <Video 1> writing notes at the desk.

summary:
[video editing] The target video is an edited version of <Video 1>, altering the color of <Subject 1>'s collared shirt from pink to royal blue while retaining all original motion, framing, lighting, and timeline pacing.

retention_analysis:
<Video 1> (shot structure and timing): fully_preserved - the camera framing, cuts, background office setting, and timing are fully retained.
<Subject 1> (appears in [Shot 1], [Shot 2]): partially_preserved - the woman's facial identity, tied-back hairstyle, downward gaze, posture, and note-writing hand movements are fully preserved, with only her shirt color modified from pink to royal blue.

detailed_description:
The target video is an edited version of <Video 1> in a clean broadcast documentary style.
[Shot 1] The opening follows <Video 1> showing colleagues seated at an office desk.
[Shot 2] At 00:02.000, following the cut in <Video 1>, the camera focuses in medium close-up on <Subject 1>, the young woman with tied-back dark hair seated at the desk writing notes on paper with a pen. Her button-up collared shirt is now a solid royal blue instead of the original pink. Her calm facial expression, focused downward gaze, natural hand movement while writing, and the background office desks with laptops and seated colleagues remain identical to <Video 1>.

overall_soundscape:
Subtle indoor office room tone and quiet paper-handling ambience continue throughout the scene.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000028

ref2va-sample-2026083032-000015

Create a realistic video where the woman from Picture 1 and the man in the embellished suit from Picture 2 meet and interact inside the reception lounge shown in Picture 3.
multi_subject_interactiondirect ablation12sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
Generate a realistic 12-second video featuring <Subject 1> from <Picture 1> and <Subject 2> from <Picture 2> interacting non-verbally inside the modern reception lounge <Subject 3> from <Picture 3>. In an opening medium-wide shot, <Subject 1> stands by the white reception counter reviewing a tablet when <Subject 2> enters the room and approaches; both exchange polite nods. At 00:06.500, cut to a medium two-shot where <Subject 2> gestures toward the reception desk and <Subject 1> responds with an accommodating gesture and warm smile.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000029

ref2va-sample-2026083032-000015

Create a realistic video where the woman from Picture 1 and the man in the embellished suit from Picture 2 meet and interact inside the reception lounge shown in Picture 3.
multi_subject_interactionContext-IR approximation12sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman in <Picture 1>, wearing a black-and-white striped crop top, open black knit cardigan, blue jeans, and dark flat shoes.
<Subject 2> is the man in <Picture 2>, wearing a black tailored suit with colorful jewel embellishments on the lapels, pockets, and shoes.
<Subject 3> is the modern reception lounge environment in <Picture 3>, featuring a white front service counter, vibrant blue accent wall, ceiling track lighting, and grey flooring.

summary:
[reference generation] The target video portrays a non-verbal interaction between <Subject 1> and <Subject 2> within the modern reception lounge of <Subject 3> across two sequential shots.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the woman's facial features, hairstyle, striped crop top, cardigan, and jeans are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the man's facial features, haircut, and embellished black suit are retained.
<Subject 3> (appears in [Shot 1], [Shot 2]): fully_preserved - the white counter, blue accent walls, ceiling lighting fixtures, and interior layout are retained.

detailed_description:
The target video is shot in a realistic, cinematic style with clean, bright interior illumination.
[Shot 1] A medium-wide shot opens inside <Subject 3>, showcasing the white service counter and vibrant blue accent wall. <Subject 1>, wearing her striped crop top, dark cardigan, and blue jeans, stands near the counter reviewing a tablet device. From the hallway on the right, <Subject 2>, dressed in his black suit with jewel-studded lapels, walks into the room with steady, measured strides. As he approaches the counter area, <Subject 1> looks up, notices his presence, and offers a polite, welcoming nod and warm smile. <Subject 2> slows his pace, returning a gentle nod of acknowledgement as he pauses near the counter.
[Shot 2] At 00:06.500, the shot cuts to a medium two-shot tracking slightly toward the subjects. <Subject 2> gestures toward the display area behind the counter with an open hand, maintaining a calm and attentive expression. <Subject 1> turns slightly in the direction of his gesture, smiling receptively and gesturing back toward the white counter top with an agreeable motion. Both subjects share a brief, respectful visual exchange as the camera gently settles into a stable two-shot composition before the scene concludes.

overall_soundscape:
Subtle modern office ambience with faint air conditioning hum, quiet indoor room resonance, and soft footsteps on the floor.

non_diegetic_music:
A gentle, contemporary ambient electronic track with soft synthesizer pads and light melodic pulses plays at a steady, low volume throughout.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000030

ref2va-sample-2026083032-000016

Animate the stylized character from the image aiming and firing his mechanical crossbow in a misty bamboo forest.
character_subjectdirect ablation12sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
Generate a 12-second stylized cinematic video of the warrior <Subject 1> from <Picture 1> in a misty bamboo forest at twilight. <Subject 1> raises his mechanical crossbow, aims steadily as gears whir into position, and fires a glowing energy bolt into the distance while the wind billows his robes and hair.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000031

ref2va-sample-2026083032-000016

Animate the stylized character from the image aiming and firing his mechanical crossbow in a misty bamboo forest.
character_subjectContext-IR approximation12sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the stylized young warrior in <Picture 1>, featuring long wavy ash-brown hair, layered black and white martial robes with golden embroidery and flowing sashes, golden mechanical prosthetic legs, and a handheld mechanical crossbow.

summary:
[reference generation] The target video depicts <Subject 1> in a misty bamboo forest at twilight, raising his mechanical crossbow, locking onto a distant target, and firing a high-speed bolt as wind swirls around him.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the character identity, costume details, golden mechanical legs, and crossbow from <Picture 1> are retained.

detailed_description:
The target video is rendered in a high-end stylized anime/fantasy illustration aesthetic with atmospheric lighting, drifting mist, and subtle particle effects.
[Shot 1] A medium shot reveals <Subject 1> standing firmly on the moss-covered ground of a tranquil bamboo forest at twilight. A cool breeze causes the tall bamboo stalks to sway gently, sending pale green leaves drifting through the air. The long ribbons and dark sashes of <Subject 1>'s layered robe flutter in the wind, while his golden mechanical prosthetic legs catch subtle moonlight filtering through the canopy. With focused concentration, he raises the mechanical crossbow with his right arm, aligning his sight toward the off-screen distance as the internal gears and tension cords of the weapon softly click into position.
[Shot 2] At 00:06.000, the shot cuts to a dynamic low-angle tracking shot orbiting smoothly around <Subject 1>. A faint golden energy gathers at the tip of the crossbow bolt. His ash-brown hair billows backward in the strengthening wind. He smoothly squeezes the trigger mechanism; the crossbow releases a high-speed projectile with a sharp mechanical snap, unleashing a bright golden streak that cuts through the forest fog. <Subject 1> lowers his weapon with steady discipline, his gaze tracking the trajectory into the distance as ambient mist swirls past him.

overall_soundscape:
A gentle breeze rustling through bamboo groves, soft footfalls on damp moss, the intricate mechanical clicks and whirs of the crossbow mechanism, followed by a resonant mechanical snap and whoosh upon firing.

non_diegetic_music:
An atmospheric fantasy soundtrack blending soft guzheng notes with subtle orchestral string pads, swelling gently during the weapon release and fading into a serene ambiance.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000032

ref2va-sample-2026083032-000017

Generate an 8-second realistic video featuring the woman in the hat from Picture 1 and the woman in the patterned dress from Picture 2 meeting at the wooden reception desk area shown in Picture 3, interacting through gestures and expressions without spoken dialogue.
multi_subject_interactiondirect ablation8sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
Generate an 8-second realistic video in the reception area from <Picture 3>. <Subject 1>, the woman from <Picture 1> in the wide-brim hat, grey top, and rolled jeans, stands near the wooden desk as <Subject 2>, the woman from <Picture 2> in the sleeveless patterned dress and ponytail, walks up toward the counter. They exchange a polite smile and an acknowledging nod, interacting non-verbally with natural posture and gestures.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000033

ref2va-sample-2026083032-000017

Generate an 8-second realistic video featuring the woman in the hat from Picture 1 and the woman in the patterned dress from Picture 2 meeting at the wooden reception desk area shown in Picture 3, interacting through gestures and expressions without spoken dialogue.
multi_subject_interactionContext-IR approximation8sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman from <Picture 1>, wearing a wide-brim hat, long-sleeve light grey top, light blue cuffed jeans, black ankle boots, and layered necklaces.
<Subject 2> is the woman from <Picture 2>, with hair in a sleek ponytail, wearing a sleeveless high-neck patterned dress with a dark waist belt and strappy heels.
<Subject 3> is the indoor reception environment from <Picture 3>, featuring a tall wooden front desk with a table lamp, framed pictures on a beige wall, and carpeted flooring.

summary:
[reference generation] The target video shows <Subject 1> standing near the reception desk in <Subject 3> as <Subject 2> approaches the counter, where they share a polite non-verbal interaction with nods and friendly expressions.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the woman's facial identity, wide-brim hat, long-sleeve grey top, cuffed jeans, boots, and necklace styling are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the woman's facial identity, ponytail hairstyle, high-neck patterned dress, belt, and heel footwear are retained.
<Subject 3> (appears in [Shot 1]): fully_preserved - the wooden reception counter, table lamp, framed artwork on the beige wall, and carpeted interior setting are retained.

detailed_description:
The target video has a clean, naturalistic cinematic look with warm, soft interior lighting.
[Shot 1] A steady medium-wide shot captures <Subject 3>, the office reception area with its polished wooden counter, table lamp, framed wall art, and neutral carpet. <Subject 1>, the woman in the wide-brim hat, grey long-sleeve top, and cuffed blue jeans, stands behind the wooden desk leaning slightly forward in a relaxed, welcoming posture. From the right side of the frame, <Subject 2>, wearing the sleeveless patterned dress and sleek ponytail, walks gracefully toward the front desk. As <Subject 2> reaches the counter, <Subject 1> smiles warmly and offers a polite nod of greeting. <Subject 2> pauses, returns the friendly smile with an acknowledging tilt of her head, and places one hand gently on the desk edge. The camera maintains a smooth, subtle slow push-in as both women hold their pleasant non-verbal rapport until the shot ends.

overall_soundscape:
Faint indoor office room tone with soft footsteps on carpet and the subtle rustle of clothing fabric.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000034

ref2va-sample-2026083032-000018

Animate the man in Picture 1 to speak to the camera with the same head movements and facial expressions as the person in Video 1.
motion_performance_transferdirect ablation8sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 8.0s
Exact H3 prompt
Generate an 8-second video animating the man from <Picture 1> in a static close-up shot, applying the head movements, facial expressions, and speaking cadence from <Video 1> while retaining the appearance, clothing, and dark background of <Picture 1>.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000035

ref2va-sample-2026083032-000018

Animate the man in Picture 1 to speak to the camera with the same head movements and facial expressions as the person in Video 1.
motion_performance_transferContext-IR approximation8sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the man whose facial identity, long dark wavy hair, beard, black t-shirt, and dark room background are established in <Picture 1>, and whose conversational head motion, facial expressions, and speaking gestures are referenced from <Video 1>.

summary:
[reference generation] The target video animates <Subject 1> in a static close-up shot, applying the talking performance and head movements of <Video 1> to the identity and setting of <Picture 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the man's identity, hair, beard, black clothing, and dark indoor setting from <Picture 1> are preserved, animated by the performance from <Video 1>.

detailed_description:
The target video is a realistic conversational close-up with soft, low-key indoor lighting.
[Shot 1] In a static close-up framing, <Subject 1>, the man with long wavy dark hair, a beard, and a black t-shirt from <Picture 1>, faces the camera in a dimly lit room. Driven by the motion performance from <Video 1>, <Subject 1> speaks with animated head movements, subtle nodding, expressive eyebrow shifts, and continuous natural lip motion across the entire 8 seconds, maintaining direct eye contact with the lens in a steady, conversational delivery.

overall_soundscape:
Faint indoor room tone and subtle microphone ambient noise.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000036

ref2va-sample-2026083032-000019

Animate the woman from the photo in her kitchen performing the exact gestures, head turns, and talking motions from the video over 10 seconds.
motion_performance_transferdirect ablation10sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 10.0s
Exact H3 prompt
Generate a 10-second video animating <Subject 1>, the woman in the kitchen from <Picture 1>, to perform the exact talking gestures, head turns, and body movements of <Subject 2> from <Video 1>, retaining the appearance and kitchen background of <Picture 1> while adopting the motion timing of <Video 1>.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000037

ref2va-sample-2026083032-000019

Animate the woman from the photo in her kitchen performing the exact gestures, head turns, and talking motions from the video over 10 seconds.
motion_performance_transferContext-IR approximation10sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 10.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the curly-haired blonde woman in <Picture 1>, wearing a sleeveless patterned top, situated in the kitchen setting from <Picture 1>.
<Subject 2> is the physical performance and body motion from <Video 1>, including the conversational gesturing, turning toward the background, and raising a bottle while speaking.

summary:
[reference generation] The target video animates <Subject 1> performing the motion, posture, head turns, and hand gestures of <Subject 2> across a single 10-second shot.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the woman's identity, curly blonde hair, sleeveless top, and the kitchen background with cabinets, sink, and stove are retained.
<Subject 2> (motion and performance): fully_preserved - the body movements, head turns, hand gestures, and timing from <Video 1> are transferred directly to <Subject 1>.

detailed_description:
The target video is a realistic indoor vlog-style recording with natural indoor kitchen lighting.
[Shot 1] A medium shot frames <Subject 1>, the woman with curly blonde hair and a patterned sleeveless top standing in her kitchen from <Picture 1>. She faces the camera and begins speaking with expressive facial movements and hand gestures matching <Subject 2> from <Video 1>. At 00:02.000, following the motion of <Subject 2>, she turns around toward the kitchen cabinets behind her, gesturing toward them before turning back toward the camera at 00:04.000. She then raises a bottle in her right hand to show it to the camera while continuing to speak with lively facial expressions. At 00:07.000, she turns back toward the cabinet surface, reaching out with her hand before turning back to face the camera at 00:09.000, finishing her spoken delivery as the shot concludes.

overall_soundscape:
Natural indoor room tone with subtle kitchen ambience.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000038

ref2va-sample-2026083032-000006

Create a 10-second video where the man from Picture 1 and the woman from Picture 2 meet in the reception office from Picture 3, interacting non-verbally as she approaches the desk and they exchange a polite nod.
multi_subject_interactiondirect ablation10sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
Generate a 10-second realistic video set inside the reception area from <Picture 3>. The man from <Picture 1> stands behind the wooden reception counter, while the woman from <Picture 2> enters from the hallway holding her clutch. She approaches the desk, pauses, and makes eye contact with him as both exchange a polite, subtle nod under warm office lighting.
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA
ref2va-sample-2026083400-000039

ref2va-sample-2026083032-000006

Create a 10-second video where the man from Picture 1 and the woman from Picture 2 meet in the reception office from Picture 3, interacting non-verbally as she approaches the desk and they exchange a polite nod.
multi_subject_interactionContext-IR approximation10sgeneratedpending
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
<Picture 3>
<Picture 3>environment
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the man in <Picture 1>, with short dark hair, wearing a long black coat, an open white shirt, and metallic gold patterned trousers.
<Subject 2> is the woman in <Picture 2>, with shoulder-length brown hair, wearing a red blazer, a long white shirt, distressed light-wash jeans, black ankle boots, and holding a black clutch.
<Subject 3> is the office reception environment in <Picture 3>, featuring a tall dark-wood desk, a shaded table lamp, three framed wall prints, beige walls, and carpeted flooring.

summary:
[reference generation] In the reception area of <Subject 3>, <Subject 1> stands behind the counter as <Subject 2> walks into the room, approaches the desk, and exchanges a subtle polite nod with him.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the man's facial features, dark hair, black coat, white shirt, and gold trousers are preserved.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the woman's facial features, brown hair, red blazer, white shirt, ripped jeans, boots, and clutch are preserved.
<Subject 3> (appears in [Shot 1], [Shot 2]): fully_preserved - the wooden reception desk, lamp, wall art, and office layout are preserved.

detailed_description:
The target video has a realistic cinematic look with warm interior office lighting and subtle camera movement.
[Shot 1] The scene opens in a medium wide shot establishing <Subject 3>, the quiet reception area with its tall wooden counter, small lamp, framed wall pictures, and carpeted hallway. <Subject 1>, wearing his long black coat, white shirt, and gold trousers, stands behind the wooden reception desk reviewing documents on the counter. From the hallway on the right, <Subject 2>, dressed in her red blazer, long white shirt, ripped jeans, and black ankle boots while holding her clutch, walks smoothly toward the counter. Her footsteps land softly on the carpet as she approaches the desk with a calm, composed expression.
[Shot 2] At 00:05.500, the shot cuts to a medium two-shot framing both characters across the reception counter. <Subject 2> reaches the desk, pausing and resting her free hand gently on the wooden surface. She looks up and makes eye contact with <Subject 1>. <Subject 1> looks up from the counter, smiles gently, and gives a respectful, slight nod of greeting. <Subject 2> responds with a gentle nod and a brief, pleasant smile, maintaining eye contact as the shot slowly eases in until the final frame.

overall_soundscape:
Quiet indoor office room tone, soft footsteps muffled by carpeting, and the subtle rustle of clothing.

non_diegetic_music:
N/A
Final delivered output
H3 generation in progress
PENDING
QA pending
Gemini 3.7 independent output QA