Close article
Insights

Build an AI Creative-Direction Regression Set.

Test more than technical validity: build a compact evidence set that shows whether an AI workflow still serves its intended creative direction.

4 min read

An AI workflow can pass every technical check and still become creatively worse. Its output may be fluent, correctly formatted and safe, yet gradually lose the voice, visual logic or product purpose that made it useful. Viotus publicly describes AI as a capability directed by people: judgement, authorship and final responsibility remain human. Turning that principle into practice requires a regression set that tests direction, not merely operation.

Define direction as observable intent

Begin with four to eight intentions that a reviewer can recognise in an output. Avoid labels such as ‘better’, ‘cinematic’ or ‘on brand’ unless the team can describe their evidence. An intention should connect a creative choice to its purpose: preserve a restrained voice so guidance feels trustworthy, keep environmental variation subordinate to a location’s authored identity, or favour responsive behaviour only when it strengthens immersion. Each statement gives reviewers something concrete to inspect.

For every intention, record one representative input, the expected qualities and the reason those qualities matter. Do not freeze a single ideal answer; that would reward imitation rather than direction. Preserve a range of acceptable outcomes instead. Viotus’s public description of Dreams Of Lumina offers the useful boundary: intelligence can support richer behaviour and faster iteration without replacing authored direction. The test therefore asks whether variation serves the authored purpose, not whether every run is identical.

Pair acceptable range with forbidden drift

A regression set becomes sharper when it contains rejection evidence. For each intention, add two or three drift patterns that can look polished while violating the purpose. Examples include a confident voice that erases uncertainty, extra detail that weakens a focal point, or responsive behaviour that attracts attention but breaks the atmosphere. Describe the observable fault and its consequence. This keeps reviewers from rejecting work merely because it differs from a favourite example.

Separate a hard boundary from a preference. A hard boundary contradicts authorship, safety or the product’s declared direction and should fail the case. A preference can be discussed without blocking the change. Record borderline samples as well as obvious failures: subtle regressions are precisely what ordinary format and validity checks miss. The set should make disagreement inspectable by naming which intention is affected, rather than collapsing every concern into a vague quality score.

Run a blinded comparison with rationale

Before changing a model, prompt, retrieval source or workflow step, run the same set through the current and candidate systems. Remove labels that reveal which output is new. Ask at least one accountable human reviewer to mark each intention as preserved, weakened or contradicted, then cite the observed passage or feature. The rationale matters more than a winner count because it reveals whether reviewers are applying the same creative standard.

Review cases individually before looking at the aggregate. Averages can hide a severe failure in a rare but defining intention. Where reviewers differ, compare their cited evidence and refine the case language; do not automatically turn a disputed preference into a rule. Keep the original output, decision and rationale together so the next change is compared against stable evidence. Human accountability means the team owns both the criteria and the decision, not that a person merely clicks approve.

Set replacement thresholds before the result

Decide acceptance rules before seeing the candidate outputs. A practical threshold might require no hard-boundary failures, no repeated weakening of the same core intention, and a written resolution for every material disagreement. The exact rule belongs to the product team; the important design move is to prevent enthusiasm for speed or novelty from rewriting the standard after the test. A technically superior candidate may still be unsuitable for the creative purpose.

Update the set when the authored direction changes, a new failure appears or an old case stops representing real work. Do not silently edit history: version the intention, explain the change and retain the earlier decision. This creates a compact memory of what the team means by quality. Read [Viotus’s human-direction principle](/news/human-direction-for-ai-powered-products/) and the public [Dreams Of Lumina direction](/products/worlds/dreams-of-lumina/) to see why capable technology still needs an explicit human standard.