Baseline 20
Original accepted-20 run, with the current evaluator applied to the 17 prompt-valid cases.
Current strict QA: 12 / 17 accepted
20 cases
Round 1 · Prompt A/B
Direct vs minimal six-section, matched seeds, with final soundtrack delivery where applicable.
2 accepted / 10 after final delivery
10 cases
Round 2 · Matched Seeds
Compact contract vs minimal six-section; deterministic silence/source soundtrack.
Core 2 / 8 · expansion gate failed
10 cases
Round 3 · Official Format A/B
Five baseline-success combinations with matched seeds: direct requests vs official six-section prompts. Gemini 3.7 prompt QA approved all five pairs.
Core 8 / 8 · pair coverage 4 / 4
10 cases
Promotion 20 · Matched A/B
Twenty semantically accepted real-workload combinations. Direct and Context-IR-style prompts use identical references, duration, seed, and optimized 4xGB200 runtime.
Generating 11 / 40 · QA 7/11 accepted · both accepted 2, direct only accepted 1, incomplete 1, neither accepted 1, structured only accepted 1 · task gates: A1 100% (1 pairs, insufficient evidence), B3 100% (1 pairs, insufficient evidence), D2 0% (1 pairs, insufficient evidence), D4 100% (2 pairs, insufficient evidence)
40 cases