Maor Bril argues that frame-level metrics do not establish whether a video tells the intended story. His Character.ai workflow evaluates consistency, physics, pacing and audio alignment, starting with repeatable metrics and human-calibrated model judges.
Maor Bril distills those judgments into a smaller vision-language model to make evaluation fast enough for the generation loop. Pairwise comparisons replace arbitrary absolute scores, but an initial dataset teaches shortcuts: the model rewards visual coherence rather than the intended quality dimensions.
Maor Bril describes repairing the training data with consistent encoding and annotation, avoiding a simple real-versus-AI detector. Agents then use the evaluator to detect and repair drift earlier. In the Q&A, he discusses human taste, serving economics and audio timing, while acknowledging that lip-sync evaluation remains unresolved.
Watch on YouTube




