H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-30 · no auto refresh

Baseline 20

Original accepted-20 run, with the current evaluator applied to the 17 prompt-valid cases.

Current strict QA: 12 / 17 accepted
ref2va-sample-2026082240-004277

ref2va-sample-2026082240-004277

Generate an 8-second realistic video of this woman walking through a sunlit modern city plaza. She should stroll forward gracefully, pause near a fountain, and look at the camera with a warm smile.
character_subjectContext-IR approximation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman whose identity and appearance come from <Picture 1>.

summary:
[reference generation] The target video follows <Subject 1> as she walks through a bright, modern urban plaza on a sunny afternoon, pauses beside a contemporary fountain, and looks toward the camera with a gentle, confident smile.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the woman's facial identity, hairstyle, outfit, and overall visual appearance from <Picture 1> are faithfully retained throughout.

detailed_description:
The target video is captured in a bright, elegant commercial cinematic style with warm afternoon sunlight and subtle handheld camera movement.
[Shot 1] The scene opens in a wide shot within an expansive modern urban plaza paved with sleek granite tiles and bordered by glass architecture. <Subject 1>, wearing her outfit from <Picture 1>, walks gracefully forward along the open walkway toward the camera. Warm sunlight casts soft shadows behind her as she takes poised, steady steps. A gentle breeze lightly moves her hair. The camera smoothly tracks backward while keeping her centered in a fluid medium-wide composition.
[Shot 2] At 00:04.500, the shot cuts to a medium close-up of <Subject 1> coming to a gentle halt near a sleek stone fountain. She turns slightly, letting her gaze sweep across the surrounding plaza before turning her eyes directly toward the lens. Her expression relaxes into a soft, genuine smile as specular highlights from the fountain water sparkle softly in the warm background bokeh. The camera slowly pushes in slightly on her face until the final frame.

overall_soundscape:
Gentle outdoor urban plaza ambience, soft flowing water sounds from the fountain, and rhythmic footsteps on stone pavement.

non_diegetic_music:
An airy, uplifting acoustic instrumental featuring soft guitar strumming and gentle piano accents, maintaining a warm and elegant atmosphere throughout.
Final delivered output
ACCEPTED
The generated video successfully preserves the identity and wardrobe from the source reference while executing the requested walking motion, camera cuts, and cheerful expression in a high-fidelity plaza environment.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 167.8s generation
ref2va-sample-2026082240-000983

ref2va-sample-2026082240-000983

Generate an 8-second anime clip of this character standing in a breezy school courtyard, looking shyly toward the camera with her hands clasped before offering a warm, gentle smile.
character_subjectContext-IR approximation8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the character depicted in <Picture 1>.

summary:
[reference generation] An 8-second anime sequence featuring <Subject 1> standing in a sunlit school courtyard surrounded by falling cherry blossom petals, shyly looking toward the camera and breaking into a gentle smile.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the character's identity, visual design, attire, and anime art style from <Picture 1> are fully preserved.

detailed_description:
The target video is animated in a vibrant, polished Japanese anime aesthetic with warm, soft afternoon sunlight and dynamic natural lighting.
[Shot 1] A medium full shot establishes <Subject 1> standing in the center of a sunlit school courtyard lined with stone pathways and blooming cherry blossom trees. She maintains a slightly timid posture with her hands clasped gently in front of her chest. A soft spring breeze rustles the nearby foliage, causing her hair, ears, tail, and skirt to sway naturally while delicate flower petals drift across the frame. The camera executes a slow, subtle push-in toward her. <Subject 1> looks directly toward the viewer, blinks softly with a bashful expression, and gradually eases into a sweet, shy smile, offering a slight tilt of her head.

overall_soundscape:
A gentle breeze rustles through tree canopies, accompanied by distant ambient courtyard room tone and faint bird chirps.

non_diegetic_music:
A delicate, uplifting acoustic guitar and soft piano melody playing at a relaxed tempo.
Final delivered output
ACCEPTED
The generated video accurately reproduces the reference character and follows the prompt's scene description, facial expression, and temporal progression with high consistency.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 150.7s generation
ref2va-sample-2026082240-004547

ref2va-sample-2026082240-004547

Animate the girl from this image acting as a live news reporter outdoors at night, mimicking the gestures, head movements, and speaking performance from the video.
character_subjectfree_text12sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 7.674s
Exact H3 prompt
subject_definitions:
<Subject 1> is the anime character from <Picture 1>, adopting the live television reporter motion and speaking gestures from <Video 1>.

summary:
[reference generation] The target video features <Subject 1> standing outdoors at night in front of a brick building, delivering a live news report straight to the camera while holding a microphone, retargeting the head tilts, expressive facial cadence, and body motion from <Video 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): attribute_transfer - the character design from <Picture 1> is preserved while adopting the speaking motion, gestures, and live-reporting presentation from <Video 1>.

detailed_description:
The target video is rendered in a vibrant 2D anime style with clear line art and outdoor nighttime lighting.
[Shot 1] A medium shot frames <Subject 1> (S1) standing in front of a dark brick building at night. Holding a news microphone in hand, <Subject 1> (S1) speaks directly into the lens with animated facial expressions, rhythmic head tilts, and hand gestures retargeting the performance from <Video 1>. A lower-third news banner is visible at the bottom of the screen. The camera remains in a steady tripod broadcast framing throughout the 12-second live report.

overall_soundscape:
Faint outdoor nighttime street ambience with distant traffic hum.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
Successful character motion retargeting preserving identity and recreating the news report presentation faithfully.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 current output QA · 456.4s generation
ref2va-sample-2026082240-002931

ref2va-sample-2026082240-002931

Transfer the speaking head motion and facial performance from the video onto the woman in this image, keeping her identity and clothing unchanged.
motion_performance_transferfree_text8sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 6.673s
Exact H3 prompt
subject_definitions:
<Subject 1> is the subject whose visual identity and attire come from <Picture 1> and whose speaking performance and head movement are transferred from <Video 1>.
<Video 1> is the motion performance reference video.

summary:
[reference generation] The target video animates <Subject 1>, applying the facial expressions, head tilts, and speech cadence from <Video 1> to the visual subject from <Picture 1> over an 8-second single shot.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the identity, hair, and wardrobe from <Picture 1> are retained while driven by performance motion.
<Video 1> (motion performance): attribute_transfer - the head gestures and facial articulation dynamics are transferred to <Subject 1>.

detailed_description:
The target video features a realistic, clean portrait aesthetic with soft ambient lighting.
[Shot 1] A medium shot frames <Subject 1> in front of the window setting. Driven by the performance in <Video 1>, <Subject 1> speaks directly toward the camera, shifting head angles with natural emphasis and expressive facial movements. Around 00:05.000, <Subject 1> tilts her head downward slightly while continuing to move her lips naturally until the shot concludes at 00:08.000.

overall_soundscape:
Quiet indoor room ambience.

non_diegetic_music:
N/A
Final delivered output
REJECTED
Target identity is retained well from the reference image, but speech motion tracking is inaccurate and the audio was replaced with unrelated speech.
failure: motion_transfer_failed, audio_missing_or_changed
primary_transformation: 3source_preservation: 3target_identity: 4source_leakage: 2motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 current output QA · 291.4s generation
ref2va-sample-2026082240-004938

ref2va-sample-2026082240-004938

Create a video where the character from this image balances and walks along a high slackline stretched over crashing ocean waves, following the balancing motion and camera perspective in the video.
motion_performance_transferfree_text8sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 6.48s
Exact H3 prompt
subject_definitions:
<Subject 1> is the character whose identity and visual design are provided by <Picture 1>, performing the slackline balancing and walking motion referenced from <Video 1>.

summary:
[reference generation] The target video portrays <Subject 1> balancing and walking along a taut highline suspended above turbulent ocean waves between dark coastal cliffs, matching the high-angle rear perspective and careful locomotive balancing mechanics seen in <Video 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): attribute_transfer - the visual identity is derived from <Picture 1>, while the slackline posture, arm balancing adjustments, and walking motion are transferred from <Video 1>.

detailed_description:
The target video is rendered in a vibrant anime visual style with dynamic outdoor lighting and energetic wave animation.
[Shot 1] From a high-angle rear-quarter perspective looking down past dark coastal cliffs, <Subject 1> is positioned mid-line on a taut slackline suspended high over churned, white-capped turquoise ocean water. Mimicking the balancing performance in <Video 1>, <Subject 1> extends both arms high to the sides, actively flexing wrists and shifting weight to keep center of gravity steady. The camera slowly tracks forward and zooms slightly in from behind and above as <Subject 1> cautiously steps one foot in front of the other along the narrow strap, with sea breeze rustling clothes as waves crash heavily on the cliffside below until the video ends.

overall_soundscape:
The continuous heavy roar of ocean swell crashing against rocky cliffs, mixed with the rushing sound of open coastal wind.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
Character, slackline action, setting, pose, and camera adhere.
manual audit; excluded from current 17-case re-QA · 248.5s generation
ref2va-sample-2026082240-009763

ref2va-sample-2026082240-009763

Animate the man from this image to recreate the direct-to-camera speaking performance and mug-holding gestures from the video.
motion_performance_transferContext-IR approximation12sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>motion_performance
video · 13.9s
Exact H3 prompt
subject_definitions:
<Subject 1> is the man whose visual identity is referenced from <Picture 1> and whose upper-body speaking motion and mug-holding performance are referenced from <Video 1>.
<Video 1> is the performance video providing the upper-body gestures, head movement, and conversational timing for <Subject 1>.

summary:
[reference generation] The target video portrays <Subject 1>, with visual identity derived from <Picture 1>, performing the direct-to-camera vlog-style speaking and mug-holding actions transferred from <Video 1> in a kitchen setting.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the facial identity, hairstyle, and physique from <Picture 1> are retained while executing the performance from <Video 1>.
<Video 1> (motion performance): fully_preserved - the natural head tilts, conversational lip sync motion, and two-handed mug-holding gestures are transferred throughout the single shot.

detailed_description:
The target video features a realistic, casual indoor vlog aesthetic with soft ambient kitchen lighting.
[Shot 1] A medium close-up establishes <Subject 1> positioned centrally against a kitchen backdrop, looking directly into the camera lens. Holding a white ceramic mug with both hands at chest level, <Subject 1> carries out the continuous conversational performance transferred from <Video 1>. He articulates smoothly with expressive head nods, subtle facial expressions, and natural blinking while maintaining a steady stance. As the video progresses, his hands lightly adjust their grip on the mug in sync with his conversational pacing across the continuous 12-second shot, with the camera remaining locked off in a fixed medium framing.

overall_soundscape:
Soft indoor room tone with subtle domestic kitchen ambience.

non_diegetic_music:
N/A
Final delivered output
REJECTED
Identity transfer failed completely. The subject from the source video remains instead of being replaced by the reference man.
failure: primary_transformation_missing, target_identity_lost, source_identity_leakage
primary_transformation: 2source_preservation: 3target_identity: 2source_leakage: 4motion_or_temporal: 3audio: 2artifact: 3
Gemini 3.7 current output QA · 659.2s generation
ref2va-sample-2026082240-006062

ref2va-sample-2026082240-006062

Create an 8-second video of two garage mechanics hosting a video show behind a workbench, transferring the conversational gestures, head rubbing, and finger-raising timing from the two men in the video.
motion_performance_transferContext-IR approximation8sgeneratedrejected
<Video 1>motion_performance
video · 7.04s
Exact H3 prompt
subject_definitions:
<Subject 1> is the motion and physical performance of the two men in <Video 1>, including their seating arrangement, head gestures, conversational laughter, left-hand temple rub, and the one-finger raised pointing emphasis at the conclusion.

summary:
[reference generation] The target video portrays two mechanics sitting side-by-side behind a cluttered garage workbench hosting an online workshop show, transferring the conversational performance, reactive gesturing, and comedic timing from <Subject 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): attribute_transfer - the seated interactive dynamics, casual nodding, head-rubbing gesture on the left, and emphatic one-finger gesture near the end are transferred onto the two garage mechanic characters.

detailed_description:
The target video uses a bright, realistic multi-camera vlog style with natural garage workshop lighting and warm workbench textures.
[Shot 1] A static medium shot captures two male mechanics sitting side-by-side behind a wooden workbench filled with vintage motorcycle parts, wrenches, and an iron vise. The mechanic on the right, wearing a grey graphic T-shirt, turns toward his co-host with a wide grin, speaking lively and gesturing casually with his hands following the timing in <Subject 1>. The mechanic on the left, wearing a dark navy work shirt and a baseball cap, rests his chin and right arm against the table, listening with a half-smile before lifting his left hand to rub his temple and forehead in mild exasperation. When the mechanic on the right finishes his prompt with a light chuckle, the mechanic on the left looks directly into the lens and raises his index finger firmly in the air with comedic conviction, holding the single-finger gesture until the end of the shot.

overall_soundscape:
Warm indoor garage ambience, faint distant street hum, subtle shuffling of forearms on the wooden tabletop, and quiet mechanical room tone.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The performance transfer successfully captures the timing, gestures, and expressions of the source reference with stable character motion.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 3motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 current output QA · 234.9s generation
ref2va-sample-2026082240-007682

ref2va-sample-2026082240-007682

Create a clean, 12-second commercial product video featuring the tube of cream shown in this image. Show it being picked up, opened, and applied delicately on clean skin in a bright nursery or skincare studio setup.
product_objectfree_text12sgeneratedaccepted
<Picture 1>
<Picture 1>object_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the product tube from <Picture 1>.

summary:
[reference generation] A 12-second commercial product presentation highlighting <Subject 1> on a clean, sunlit pastel tabletop, showing hands gently unscrewing its cap and dispensing a small drop of soothing cream.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the visual identity, form, labeling, and dimensions of the product tube from <Picture 1> are faithfully maintained across all shots.

detailed_description:
The video is styled as a soft, bright, high-end skincare commercial with clean studio lighting, pastel tones, and smooth cinematic focus pulls.
[Shot 1] A close-up establishes <Subject 1> resting horizontally across a soft beige textured mat surrounded by subtle botanical accents and soft daylight. A woman's manicured hand smoothly enters the frame from the right, picking up <Subject 1> and gently tilting it upward toward the lens to showcase the front packaging clearly as the camera executes a slight slow-push forward.
[Shot 2] At 00:04.500, the shot cuts to an extreme close-up of the user's hands twisting open the white screw cap of <Subject 1>. Soft physical clicks and friction sounds accompany the unscrewing action. Once the cap is removed, the hands gently squeeze the flexible body of <Subject 1>, dispensing a small, smooth pearl of white soothing cream onto an index fingertip.
[Shot 3] At 00:08.500, the shot cuts to a medium close-up from an over-the-shoulder angle. The fingertip dabs and softly blends the white cream onto the back of a hand in smooth, circular moisturizing motions, showing its light texture absorbing cleanly into the skin. <Subject 1> stands upright in soft focus in the background on the sunlit table as the shot glides gently to a halt.

overall_soundscape:
Subtle studio room tone, soft plastic cap unscrewing clicks, delicate cream dispensing texture sounds, and gentle skin-rubbing foley.

non_diegetic_music:
A gentle, warm, and uplifting acoustic guitar and ambient chime commercial track playing softly throughout the 12 seconds.
Final delivered output
ACCEPTED
The generated video accurately reflects the reference product identity while executing the commercial prompt smoothly across multiple sequential shots.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 209.7s generation
ref2va-sample-2026082240-000961

ref2va-sample-2026082240-000961

Create a 5-second animated scene of an opulent palace interior featuring a sweeping grand staircase, glowing stained glass, and shimmering light particles, matching the vibrant anime visual style of this image.
environment_styleContext-IR approximation5sgeneratedaccepted
<Picture 1>
<Picture 1>visual_style
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the 2D anime illustration visual style from <Picture 1>, characterized by rich saturated warm tones, ornate gilded architectural elements, soft pink-and-gold ambient highlights, and delicate glowing light sparkles.

summary:
[reference generation] The target video is a 5-second cinematic environment showcase of a luxurious royal ballroom staircase, rendered in the distinct aesthetic and lighting treatment of <Subject 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the vibrant 2D anime art style, ornate golden architecture, warm color grading, and sparkling particle aesthetic from <Picture 1> are retained throughout the shot.

detailed_description:
The target video is rendered in a vibrant, polished 2D anime aesthetic with opulent architectural detail and ethereal lighting.
[Shot 1] The scene opens on a wide, sweeping view of an opulent royal hall rendered in the visual style of <Subject 1>. A grand staircase adorned with plush crimson carpeting and intricately carved golden balustrades curves downward through the center of the hall. Tall arched stained-glass windows along the right wall cast patterned beams of colorful morning light across the polished marble floor and steps. Translucent pink flower petals and glistening magical light sparkles drift gently through the sunbeams in slow, floating motions. The camera executes a slow, graceful forward tracking shot while slightly panning upward toward the towering vaulted ceiling and gilded chandeliers, emphasizing the grandeur and warmth of the hall as the shimmering light particles continuously glint in the air.

overall_soundscape:
A serene indoor palace ambience with subtle, gentle acoustic reverb and soft, high-register magical crystalline chimes as light sparkles drift through the air.

non_diegetic_music:
A graceful, slow-tempo waltz melody played on solo harp and delicate string accompaniment, evoking a sense of royal elegance and wonder.
Final delivered output
ACCEPTED
The generated video successfully adopts the visual style and environment motifs of the reference image, producing a smooth and coherent animation.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 current output QA · 111.4s generation
ref2va-sample-2026082240-002728

ref2va-sample-2026082240-002728

Create a 10-second animated video featuring a squad of anime fantasy characters standing ready in a ruined, smokey wasteland, using the visual art style from this image and the supplied music as the background track.
environment_styleContext-IR approximation10sgeneratedpartial
<Picture 1>
<Picture 1>visual_style
image
<Audio 1>background_music
audio · 15.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the anime concept art visual style derived from <Picture 1>, featuring dramatic cinematic lighting, dynamic brushwork, painterly textures, and high-contrast atmospheric smoke and debris.
<Audio 1> is the background electronic music track.

summary:
[reference generation + audio reuse] A 10-second stylized anime cinematic sequence rendered in the visual style of <Subject 1>, featuring a group of heroic adventurers standing resiliently in a ruined landscape accompanied by <Audio 1> as background music.

retention_analysis:
<Subject 1> (visual style): fully_preserved - the painterly rendering, warm rim lighting, dramatic contrast, and anime aesthetic of <Picture 1> are maintained throughout.
<Audio 1>: fully_copy - <Audio 1> is reused directly as the background music track.

detailed_description:
The video adopts the dramatic anime concept art aesthetic of <Subject 1> with atmospheric lighting and swirling embers.
[Shot 1] The scene opens on a low-angle wide shot of a ruined fantasy battlefield beneath stormy skies. In the foreground, stylized adventurers stand poised in ready combat stances with cloaks and hair billowing in the wind. Sparks and smoke drift across the frame. The camera slowly pushes forward and arcs upward, emphasizing the dramatic scale of the characters against the smoldering backdrop.
[Shot 2] At 00:05.000, the shot cuts to a medium profile shot panning across the frontliner, a red-haired warrior gazing determinedly toward an unseen threat on the horizon. The rim lighting catches floating embers as the camera tracks past her staff and mechanical accents, holding the tense, heroic tableau to the end.

overall_soundscape:
Subtle howling wind and distant crackling fire ambience are faintly audible beneath the music.

non_diegetic_music:
<Audio 1> plays continuously as the non-diegetic background score throughout the video.
Final delivered output
PARTIAL
Broad style works; composition and glyph artifacts drift.
manual audit; excluded from current 17-case re-QA · 201.3s generation
ref2va-sample-2026082240-001452

ref2va-sample-2026082240-001452

Convert this video of assembling a wooden coop into high-contrast black and white, while keeping the exact action, timing, and camera framing intact.
source_preserving_editingContext-IR approximation10sgeneratedrejected
<Video 1>source_edit
video · 10.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman in the jacket and jeans, the man in glasses and t-shirt, and the wooden coop structure in the backyard setting from <Video 1>.
<Video 1> is the source video for the target video edit.

summary:
[video editing] The target video is an edited version of <Video 1>, applying a high-contrast black-and-white monochrome grade across the entire 10-second duration while retaining the original actions, camera positioning, and temporal flow of <Subject 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - all physical actions, subjects, and backyard environment are preserved while the color is transformed into high-contrast black and white.
<Video 1> (source video): fully_preserved - the continuous 10-second shot structure, framing, and pacing are preserved 1:1.

detailed_description:
The target video is rendered in a high-contrast black-and-white aesthetic with sharp midtones, deep black shadows, and crisp highlights.
[Shot 1] A continuous stationary medium shot captures <Subject 1> in the backyard from <Video 1>. A woman and a man lift and carry a wooden two-tiered animal coop together. They step forward and carefully lower the structure onto a rectangular wooden ground base positioned against a vertical wooden fence. As they settle the coop into place and adjust their footing, the stylized text overlay reading "PUTTING IT ALL TOGETHER" appears at the bottom center of the frame. The motion, timing, and camera angle follow <Video 1> across the entire 10-second duration in black and white.

overall_soundscape:
N/A

non_diegetic_music:
N/A
Final delivered output
REJECTED
The requested high-contrast black and white transformation was not applied, leaving the source footage in full color.
failure: primary_transformation_missing
primary_transformation: 2source_preservation: 4target_identity: 3source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 current output QA · 380.2s generation
ref2va-sample-2026082240-001601

ref2va-sample-2026082240-001601

Please apply a crisp, balanced color grade with enhanced contrast to this video showing the vintage Apple mouse, while keeping all the original camera movement, hand actions, and spoken dialogue perfectly intact.
source_preserving_editingContext-IR approximation5sgeneratedaccepted
<Video 1>source_edit
video · 5.0s
Exact H3 prompt
subject_definitions:
<Video 1> is the source video for the target video edit.
<Subject 1> is the desk setup, keyboard peripherals, and transparent black Apple Pro Mouse from <Video 1>.

summary:
[video editing] The target video is an edited version of <Video 1>. The edit modifies <Video 1> by replacing the original ambient lighting on <Subject 1> with high-contrast studio illumination, deepening the black tones to crushed blacks, and adding crisp specular highlights across all glossy surfaces, while preserving the camera framing, hand motion, and dialogue from <Video 1>.

retention_analysis:
<Video 1> (source video): partially_preserved - the temporal pacing, camera framing, hand action, and synchronized speech from <Video 1> are retained, while the original lighting and color grade are replaced with high-contrast studio illumination and deepened black levels.
<Subject 1> (appears in [Shot 1]): partially_preserved - the physical geometry, position, and handling of the Apple Pro Mouse, keyboard, and desk setup from <Video 1> are retained, while their surface appearance is modified by changing the illumination to high-contrast studio lighting with deeper black tones and intensified specular highlights on the transparent casing.

detailed_description:
The target video is an edited version of <Video 1> featuring high-contrast studio lighting with deep crushed blacks, clean directional key light, and crisp specular highlights across all glossy surfaces.
[Shot 1] From <Video 1>, the scene opens on a high-angle close-up of a dark desk surface next to a black keyboard, edited by replacing the original lighting with high-contrast studio illumination and deep rich blacks. A hand reaches in from the upper right, takes hold of the transparent black Apple Pro Mouse (<Subject 1>), and picks it up from the desk. An off-screen speaker (S1) speaks in a conversational, casual tone, <d>[English] This is not the original hockey puck mouse. It it's the more expenser Apple Pro Mouse.</d> While speaking, the hand tilts and rotates the mouse, turning it upside down to display the glowing red optical sensor and the underside casing labeled "Apple Pro Mouse" directly toward the camera until the end of the shot.

overall_soundscape:
Quiet indoor room tone accompanied by subtle rustling and handling sounds as the mouse is lifted and turned over.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The requested high-contrast studio relighting is successfully applied while accurately preserving the original motion, framing, and synchronized audio.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 119.6s generation
ref2va-sample-2026082240-001528

ref2va-sample-2026082240-001528

Replace the woman speaking in the video with the woman from this image, keeping the same dialogue, gestures, framing, and room background intact.
source_preserving_editingContext-IR approximation8sgeneratedrejected
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the young woman in <Picture 1>.
<Video 1> is the source video for the target video edit.

summary:
[video editing + reference generation] The target video is an edited version of <Video 1>, replacing the original speaker with <Subject 1> while preserving the indoor background, framing, gestures, and spoken dialogue.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the visual identity from <Picture 1> is applied to the speaking subject.
<Video 1> (timing, gestures, background, and dialogue): partially_preserved - the original speaker's facial identity and outfit are replaced by <Subject 1>, while the room background, physical gestures, framing, and vocal delivery are retained.

detailed_description:
The target video maintains a realistic documentary interview style with soft indoor lighting.
[Shot 1] In a static medium close-up, <Subject 1> (S1) stands indoors in front of a neutral beige wall decorated with small framed animal illustrations from <Video 1>. Facing slightly to the right of the camera, <Subject 1> (S1) speaks with expressive hand gestures, raising her hands to emphasize her words: <d>[English] Ah-da-ba da-da-da. Then I'm going to say something back to you that's kind of like that, and I might even like use a gesture like, ah-da-ba.</d> During the explanation, she gestures with her palms open and shrugs slightly with both hands raised outward toward the end of her sentence.

overall_soundscape:
Clear indoor acoustics with subtle room tone.

non_diegetic_music:
N/A
Final delivered output
REJECTED
The requested subject replacement was not performed, leaving the original video performer unchanged.
failure: primary_transformation_missing, target_identity_lost
primary_transformation: 2source_preservation: 4target_identity: 2source_leakage: 4motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 335.0s generation
ref2va-sample-2026082240-009768

ref2va-sample-2026082240-009768

Replace the man in the guitar video with the woman from this image, keeping all his movements, the guitar, and the studio setup the same.
source_preserving_editingContext-IR approximation8sgeneratedpartial
<Picture 1>
<Picture 1>subject_identity
image
<Video 1>source_edit
video · 8.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the woman from <Picture 1>, providing the facial identity that replaces the seated guitarist in <Video 1>.
<Video 1> is the source video being edited, supplying the studio setting, guitar, seated pose, gestures, and speech timing.

summary:
[video editing + reference generation] The target video is an edited version of <Video 1> where the seated guitarist's face and identity are replaced with <Subject 1> from <Picture 1>, retaining the original movements, clothing, guitar, amplifier, and setting.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - facial appearance and identity from <Picture 1> are mapped onto the performer.
<Video 1> (visual and temporal structure): fully_preserved - the static camera, grey studio background, plaid shirt, hat, seafoam green guitar, amplifier, and synchronized physical gestures are preserved.

detailed_description:
The target video retains the realistic studio tutorial aesthetic from <Video 1> with clean neutral lighting.
[Shot 1] In a static medium shot, <Subject 1> (S1), wearing the fedora, plaid button-down shirt over a t-shirt, and dark trousers from <Video 1>, sits holding the light-green electric guitar in front of a two-tone amplifier against a plain grey backdrop. Looking directly into the camera, she speaks and performs the exact hand gestures from <Video 1>, pointing downward, touching her nose, gesturing outward to emphasize her words, and holding the neck and body of the guitar throughout the continuous shot.

overall_soundscape:
Clear indoor studio voice audio and subtle room tone preserved from <Video 1>.

non_diegetic_music:
N/A
Final delivered output
PARTIAL
Replacement is broad; exact face identity is moderate.
manual audit; excluded from current 17-case re-QA · 334.3s generation
ref2va-sample-2026082240-002105

ref2va-sample-2026082240-002105

Edit this news report video while keeping the reporter's speech, dialogue timing, and background action fully intact.
source_preserving_editingfree_text10sgeneratedaccepted
<Video 1>source_edit
video · 10.0s
<Audio 1>source_soundtrack
audio · 10.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the male reporter in <Video 1>, wearing a red polo shirt with short dark hair and a beard.
<Video 1> is the source video for the target video edit.
<Audio 1> is the synchronized audio track of <Video 1> and is reused in the target video.

summary:
[video editing + audio reuse] The target video is an edited version of <Video 1>. It replaces <Subject 1>'s red polo shirt with a navy blue polo shirt while preserving his facial features, reporting gestures, the background parking lot outside the stadium, and the synchronized audio from <Audio 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1]): partially_preserved - the reporter's facial appearance, hairstyle, gestures, and reporting delivery are preserved, while his red polo shirt is replaced with a navy blue polo shirt.
<Video 1> (appears in [Shot 1]): partially_preserved - the camera framing, background buses, and overall timeline are preserved, while an edit replaces <Subject 1>'s red polo shirt with a navy blue polo shirt.
<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the complete final audio track.

detailed_description:
The target video is in a realistic news broadcast style with natural outdoor daylight.
[Shot 1] Directly modifying <Video 1>, the video replaces <Subject 1>'s original red polo shirt with a navy blue polo shirt throughout the shot. In a medium shot, <Subject 1> (S1), the male reporter with short dark hair and a beard now dressed in the navy blue polo shirt, stands in an open parking lot in front of dark buses and security personnel. <Subject 1> (S1) gestures with his hands while reporting directly to the camera, synchronized with <Audio 1>: <d>[Spanish] ...desde el Estadio Azteca por donde entrará la porra de los visitantes. Se calcula que habrá 3,000 efectivos los que se encargarán de garantizar la seguridad de los más de 80,000 asistentes esta noche al Estadio Azteca.</d> The camera stays fixed on <Subject 1> as he completes his report.

overall_soundscape:
The copied outdoor ambience, faint background activity, and wind noise from <Audio 1> continue throughout.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The edit accurately changes the polo shirt color to navy blue while preserving all other source details, motion, and audio.
failure: none
primary_transformation: 5source_preservation: 5target_identity: 5source_leakage: 1motion_or_temporal: 5audio: 5artifact: 5
Gemini 3.7 current output QA · 383.4s generation
ref2va-sample-2026082240-000145

ref2va-sample-2026082240-000145

Continue the crafting tutorial from the video, showing how to attach the remaining bead strands along the gold stand arms to finish the centerpiece.
continuation_next_shotfree_text12sgeneratedaccepted
<Video 1>continuation_context
video · 5.0s
Exact H3 prompt
subject_definitions:
<Subject 1> is the gold decorative centerpiece stand and matching bead garlands from <Video 1>.
<Subject 2> is the crafter's hands with purple glitter nail polish from <Video 1>.
<Video 1> is the source video providing continuation context.

summary:
[video continuation] The video continues directly from the end of <Video 1>, showing <Subject 2> securing bead strands along the gold arms of <Subject 1> using glue and a wooden craft stick.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the gold stand and beaded garlands are maintained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the crafter's hands and manicure style continue consistently.
<Video 1> (continuation source): fully_preserved - the action resumes seamlessly from the final frame of <Video 1>.

detailed_description:
The target video maintains a bright DIY tutorial aesthetic with clean tabletop lighting.
[Shot 1] Continuing directly from the final frame of <Video 1>, the camera remains in a steady medium close-up on <Subject 1> atop the white table. <Subject 2> presses the golden bead strand securely into the top groove using a wooden craft stick. <Subject 2> then slightly rotates <Subject 1>, dispenses a thin bead of hot glue along the adjacent arm, and presses the next section of gold beads into place. The hands carefully clear away stray adhesive strings, ensuring each strand drapes symmetrically across all four arms of the centerpiece.

overall_soundscape:
Quiet craft room ambience with subtle clicks of acrylic beads, glue gun handling sounds, and soft wooden stick taps.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The generated video effectively continues the crafting sequence from Video 1 with strong identity preservation, plausible motion, and consistent environment.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 current output QA · 276.2s generation
ref2va-sample-2026082240-007247

ref2va-sample-2026082240-007247

Create a 12-second anime video featuring the character from the first image and the character in the pink gown from the second image meeting inside the elegant grand ballroom with red stairs.
multi_subject_interactionContext-IR approximation12sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the anime character in <Picture 1>, with short brownish hair, purple eyes, a floral head wreath, and a layered white and pale green frilly dress.
<Subject 2> is the anime character in <Picture 2>, with orange hair in a side ponytail, blue eyes, a white hair flower, and an ornate pink Victorian-style gown.

summary:
[reference generation] The 12-second target video portrays <Subject 2> descending a red carpeted staircase in an opulent palace hall to greet <Subject 1>, who approaches holding a four-leaf clover.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the character appearance, green-white dress, and floral crown are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the character appearance, orange side-ponytail, and pink gown are retained.

detailed_description:
The target video is rendered in a vibrant, polished anime art style with soft indoor lighting and drifting flower petals.
[Shot 1] A medium-wide shot shows <Subject 2>, wearing her pink ornate gown, stepping gracefully down the grand red-carpeted staircase with golden railings. At the base of the stairs, <Subject 1> kneels and then gently stands up, dressed in her layered pale green and white ruffled dress with a flower crown atop her hair. <Subject 1> looks up with a warm smile, presenting a green four-leaf clover in her hands.
[Shot 2] At 00:06.500, the shot cuts to a medium two-shot at the bottom of the staircase. <Subject 2> approaches with an energetic, cheerful smile, extending her hands toward <Subject 1>. The two exchange joyful nods as pink petals drift gently through the golden hall atmosphere until the scene fades out.

overall_soundscape:
Soft rustling of dresses, gentle footsteps on carpeted stairs, and faint indoor palace room acoustics.

non_diegetic_music:
A light, uplifting orchestral anime instrumental piece featuring gentle strings, flute, and piano playing at a moderate tempo.
Final delivered output
ACCEPTED
The generated video accurately translates the reference inputs into the requested multi-subject scene with consistent character models, smooth transitions, and fitting orchestral audio.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 248.3s generation
ref2va-sample-2026082240-008492

ref2va-sample-2026082240-008492

Create a 12-second anime video where the young man playing the white piano from the second picture is observed warmly by the blonde girl from the first picture in a concert stage setting.
multi_subject_interactionContext-IR approximation12sgeneratedaccepted
<Picture 1>
<Picture 1>subject_identity
image
<Picture 2>
<Picture 2>subject_identity
image
Exact H3 prompt
subject_definitions:
<Subject 1> is the anime-styled girl in <Picture 1>, with short blonde hair, a hair ornament, and a white dress with gold accents.
<Subject 2> is the anime-styled young man in <Picture 2>, with light-blue hair, a winged halo headpiece, dark gloves, and white-and-blue clothing, playing a white grand piano.

summary:
[reference generation] The target video portrays <Subject 2> playing a grand piano on a spotlighted stage while <Subject 1> stands beside the instrument listening and smiling softly.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the blonde hairstyle, floral hair ornament, and white dress with gold trims are retained.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the light-blue hair, winged headpiece, ornate outfit, and piano performance are retained.

detailed_description:
The target video features an anime aesthetic with luminous stage lighting and drifting gold dust motes.
[Shot 1] A medium shot depicts <Subject 2>, the young man with light-blue hair and a winged halo headpiece, playing the keys of a white grand piano under warm spotlights. The camera pans smoothly to reveal <Subject 1>, the blonde girl in a white and gold dress, standing near the piano's rim, listening attentively.
[Shot 2] At 00:06.000, the shot cuts to a medium close-up of <Subject 1> looking toward <Subject 2>. Her expression softens into a warm smile as the piano melody continues to resonate throughout the hall.

overall_soundscape:
A resonant, lyrical acoustic grand piano performance echoes clearly in a spacious hall.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The generated video accurately composes both input subjects into a multi-shot scene where the pianist plays and the girl listens attentively, matching the prompt.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 285.8s generation
ref2va-sample-2026082240-003441

ref2va-sample-2026082240-003441

Create an 8-second video showing the orange sports car from the reference video maneuvering through turns on the race track.
nonhuman_dynamic_eventsContext-IR approximation8sgeneratedaccepted
<Video 1>dynamic_event
video · 4.004s
Exact H3 prompt
subject_definitions:
<Subject 1> is the orange sports car with a black rear wing and aerodynamic bodywork seen in <Video 1>.
<Subject 2> is the asphalt racetrack environment with yellow-and-black kerbs and grassy infields from <Video 1>.

summary:
[reference generation] The target video shows <Subject 1> dynamically navigating complex bends on the circuit of <Subject 2> in a continuous tracking shot.

retention_analysis:
<Subject 1> (appears in [Shot 1]): fully_preserved - the orange sports car's visual style, color, and racing spoiler are retained.
<Subject 2> (appears in [Shot 1]): fully_preserved - the track asphalt, yellow-and-black painted kerbs, and surrounding grass layout are retained.

detailed_description:
The target video features realistic motorsport broadcast footage captured in bright daylight with smooth tracking.
[Shot 1] An elevated medium-wide tracking shot follows <Subject 1>, the orange sports car with a rear wing, as it carves through a tight S-curve on <Subject 2>. The car initiates a sharp right turn, clipping the yellow-and-black inner curbing, before smoothly transferring weight into a broad left-hand sweep. The camera pans steadily across the green infield to keep the vehicle centered as it straightens out and accelerates powerfully down the ensuing straightaway.

overall_soundscape:
Aggressive high-revving sports car engine roar, transmission whine, and tire friction noise over asphalt.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The generated video effectively captures the orange sports car navigating curves with coherent motion and environment retention.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 3artifact: 4
Gemini 3.7 current output QA · 156.4s generation
ref2va-sample-2026082240-008990

ref2va-sample-2026082240-008990

Create a realistic 10-second video of a passenger train slowing down and pulling into a city railway station, following the motion and setting of the train in the reference video.
nonhuman_dynamic_eventsContext-IR approximation10sgeneratedaccepted
<Video 1>dynamic_event
video · 4.304s
Exact H3 prompt
subject_definitions:
<Subject 1> is the passenger commuter train in <Video 1>, featuring dark bodywork and a yellow front face.
<Subject 2> is the railway terminal station environment in <Video 1>, featuring parallel tracks, platform edges, and overhead catenary gantries.
<Video 1> is the motion reference for the train's deceleration and arrival sequence.

summary:
[reference generation] The target video shows a commuter train (<Subject 1>) approaching and coming to a smooth stop along the platform in a sunlit urban railway station (<Subject 2>), capturing the arrival motion referenced from <Video 1>.

retention_analysis:
<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - the commuter train's appearance, livery, and slow forward motion are preserved.
<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the railway station environment, overhead wire structures, and concrete platforms are preserved.
<Video 1> (arrival motion and deceleration dynamic): fully_preserved - the steady approach and pacing of the train pulling into the platform are maintained.

detailed_description:
The target video is rendered in a realistic, clear daylight style documenting railway transit.
[Shot 1] From a trackside perspective looking down the line, <Subject 1>, the commuter train, glides along the steel rails toward the platform of <Subject 2>. The train approaches under the overhead metal catenary gantries, slowing down gradually as it nears the station roof canopy.
[Shot 2] At 00:05.000, the shot cuts to a profile angle along the platform edge. <Subject 1> pulls alongside the passenger platform, rolling steadily until it gently stops, with light reflecting off its passenger windows and exterior body.

overall_soundscape:
The soundscape comprises the rhythmic rolling of steel wheels on rails, a low electric traction hum, soft pneumatic brake releases, and distant echoing station ambience.

non_diegetic_music:
N/A
Final delivered output
ACCEPTED
The generated video accurately reflects the commuter train approach and smooth station stop across both planned camera viewpoints with consistent visual fidelity.
failure: none
primary_transformation: 4source_preservation: 4target_identity: 4source_leakage: 2motion_or_temporal: 4audio: 4artifact: 4
Gemini 3.7 current output QA · 206.8s generation