Aug 3, 2026
Prologue Intent Independent Video Agent Evaluation
A Prologue-commissioned blind study comparing Prologue Intent with six leading AI video agents. Thirty cinematic prompts from open T2F-Bench. Fifteen metrics across three pillars. Outside reviewers score without knowing which product made which clip.
Overall quality
All 15 metrics weighted equally (0–100, higher is better).
Scores coming soon
Blind reviewer scoring is in progress. This chart will update when the pilot publishes final results.
prologue
Methodology
- Prompts: 30 screenplays sampled from T2F-Bench (open, MIT-licensed).
- Agents: 7 — Prologue Intent plus Runway Agent, Utopai PAI 2.0, MiniMax Hub, Luma Creative Agents, TapNow Agent, and Kling Canvas Agent.
- Metrics: 15 scores (0–100, equal weight) in Narrative Coherence, Cinematic Language, Production Quality.
- Review: blind, independent human reviewers. Two reviewers per clip; third adjudicator on large disagreements.
Three pillars
Narrative Coherence — Does the story hold together across scenes? Beats, causality, identity, and spatial continuity.
Cinematic Language — Does it feel directed? Shot grammar, pacing, composition, emotional intent, and style consistency.
Production Quality — Is it technically usable? Stability, artifacts, audio fit, cut quality, and overall polish.
What this is not
This pilot is inspired by frameworks like Physion Arc 1.0 but is not an official Physion result. We are using it to stress-test Prologue Intent before seeking inclusion in Physion's benchmark. We will not publish invented leaderboard numbers.
