H3 Omni Ref Iteration Review

Gemini Omni dropped · Original H3 · 4×GB200 optimized FastVideo path

Updated 2026-08-31 03:12 PDT · no auto refresh

Baseline 20

Original accepted-20 run, with the current evaluator applied to the 17 prompt-valid cases.

Current strict QA: 12 / 17 accepted
20 cases

Round 1 · Prompt A/B

Direct vs minimal six-section, matched seeds, with final soundtrack delivery where applicable.

2 accepted / 10 after final delivery
10 cases

Round 2 · Matched Seeds

Compact contract vs minimal six-section; deterministic silence/source soundtrack.

Core 2 / 8 · expansion gate failed
10 cases

Round 3 · Official Format A/B

Five baseline-success combinations with matched seeds: direct requests vs official six-section prompts. Gemini 3.7 prompt QA approved all five pairs.

Core 8 / 8 · pair coverage 4 / 4
10 cases

Promotion 20 · Matched A/B

Twenty semantically accepted real-workload combinations. Direct and Context-IR-style prompts use identical references, duration, seed, and optimized 4xGB200 runtime.

Generated 40 / 40 · QA 11/40 accepted · both accepted 2, direct only accepted 3, neither accepted 11, structured only accepted 4 · task gates: A1 0% (3 pairs, needs prompt iteration), A3 50% (4 pairs, needs prompt iteration), A4 33% (3 pairs, needs prompt iteration), B3 50% (4 pairs, capability canary), D2 33% (3 pairs, capability canary), D4 100% (3 pairs, needs prompt iteration)
40 cases

Repair 15 · Base-Compatible

Fifteen coherent real-workload combinations, three each for A1, A3, A4, B3, and D4. Direct prompts use only FastVideo-bound media labels; unsupported Subject labels are removed.

Generated 15 / 15 · QA 7/15 accepted
15 cases

Audio Repair 5 · No Invented Speech

Five A3/B3 retries with the same reference workloads but an explicit non-verbal contract: mouths at rest, no spoken dialogue, and ambient room tone only.

Generating 1 / 5 · QA 0/1 accepted · unintelligible speech 1
5 cases

Audio Repair · Silent Control

The first failed B3 output with its generated audio removed. This isolates speech corruption from visual motion-transfer and source-leakage failures.

Generated 1 / 1 · QA 0/1 accepted · unintelligible speech 0
1 cases

Evidence 22 · Fresh Sources, Two Seeds

Eleven fresh, non-overlapping source combinations: five A1, five A4, and one speech-safe A3. Each direct base-native prompt runs at two seeds to measure stability.

Generated 22 / 22 · QA 16/22 accepted · unintelligible speech 0 · incomplete 11 · stable sources 8/11 terminal
22 cases

10k Distribution v1.4

Deterministic task allocation revised from Pilot-80 evidence. Actual-media semantic acceptance remains a separate stage.

10,000 unique tasks · 750 exact dialogue · audit passed
10000 cases

Pilot 80 · Core Distribution

Eighty semantically accepted source combinations using direct, base-native prompts. Raw H3 and postprocessed delivery are reported separately.

Generated 80/80 · raw core 50/76 · delivery 69/80
80 cases

Exact Dialogue Stability 8

Four semantically accepted A1/A3 workloads, each repeated at two new seeds. This tests whether explicit short dialogue is more stable than requesting silence.

Generated 8/8 · raw QA 6/8
8 cases