How to test changes to an AI design workflow
Changes to prompts, models, and components can bring back old mistakes. Build a practical test collection for an AI design workflow and keep it useful after handoff.
The team changes a prompt to make its AI-generated prototypes more consistent. The next demo looks better. A week later, someone notices that the new workflow has stopped producing the permission-denied state.
The example is invented, but the problem is easy to miss: improving the output you are looking at does not establish what happened to the rest of the work. A design team needs a repeatable way to check changes to its prompts, models, component libraries, and source material.
This is where an AI workflow agency should be able to help after the pilot. Ask for a small set of checks tied to actual design failures, plus a way for the team to maintain it. A successful demonstration is only the beginning of that arrangement.
What needs to stay true when the workflow changes?
Choose one workflow with a defined output. For this example, the input is an approved feature brief and a component library; the output is a prototype ready for a designer’s review. Publication and implementation remain separate decisions.
Write down what makes that output usable. It must use the supplied roles accurately, represent required states, and keep unsupported assumptions visible. A reviewer needs to distinguish an approved fact from something the model filled in.
In Building eval systems that improve your AI product, Hamel Husain and Shreya Shankar explain how examining real failures can inform evaluation. Their guidance distinguishes checks for known problems from the ongoing work of finding new ones. For a design team, that means preserving examples of mistakes people actually had to correct, then using them when the workflow changes.
An evaluation is simply a test of whether the output meets an explicit requirement. The useful question is specific: does this prototype give a viewer access to an editor-only action?
What belongs in a small test collection?
Start with sanitized examples the team understands. Keep the original input, the requirement, the observed failure, and the decision a reviewer should make. Remove client information that is not approved for the tool or the test collection.
The following matrix is an illustrative starting point for the prototype workflow. Replace its examples with your own.
| Test input | Failure to catch | How to check it |
|---|---|---|
| Brief with viewer and editor roles | Viewer receives an editing action | Inspect both role paths against the supplied rules |
| Existing form with a rejected submission | Error state disappears | Check that the state remains present and understandable |
| Brief with an unanswered policy question | Model invents a definitive rule | Verify that the uncertainty stays marked for review |
| Updated component library | Prototype uses a retired component | Compare component references with the approved library |
| Content containing a long customer name | Layout hides the main action | Inspect the rendered output at relevant widths |
| Brief revised after initial generation | Output silently uses an older requirement | Trace the affected decision to the current input |
This collection covers different kinds of failure. Some checks can run automatically, such as finding a retired component reference. Others need someone to inspect behavior and meaning. A machine can confirm that an error message exists without establishing that a user knows what to do next.
How do you compare the old and new versions?
Keep the inputs fixed while you compare the proposed change with the current workflow. Record the configuration each output used, including the relevant prompt and component-library versions. Otherwise, a better result could reflect a different brief rather than the change being tested.
Agree on the acceptance rules before looking at the outputs. For example, missing a required permission boundary can block adoption even if the new version produces cleaner layouts. Other differences may be acceptable with a named reviewer and a limited scope.
When generation varies, repeat the comparison enough to understand the instability that matters for this workflow. Preserve failures as well as attractive outputs. There is no universal run count that makes every design task reliable; choose the effort in proportion to the consequences and document the remaining uncertainty.
If the team uses an AI model to judge the output, compare that judgment with human reviews on examples it was not tuned against. Do not let an agreeable second model turn a missing state into a passing result.
What happens when a change fails?
Suppose the revised prompt improves layout consistency but drops the permission-denied state in the test example. Keep the current workflow in use while the team investigates. Record the failing input and the behavior that must change.
After a fix, rerun the relevant checks and check the other cases the fix could affect. A patch that adds a permission state everywhere may create a different problem in a flow where that state has no meaning.
Preserve a way to return to the previous usable setup where the tools allow it. If a provider removes a model or changes behavior outside the team’s control, use a documented fallback, such as manual completion of the affected step. A rollback plan should describe an available action.
Who keeps the checks useful after handoff?
Assign an internal owner before the agency leaves. That person needs access to the approved inputs, the test collection, the current configuration, and the record of unresolved failures. Test the handoff by asking a colleague to evaluate a small change without help from the original builders.
Keep reviewing ordinary work after the checks pass. When designers find a new recurring failure, decide whether to add an example or revise an existing requirement. A test collection that never changes gradually becomes a record of what the team used to care about.
The AI workflow pilot guide helps define the first adoption experiment; the handoff contract covers what the receiving team needs. This next step makes change review repeatable. The team should finish with a test collection it can run, an owner who can interpret it, and a clear reason to accept or reject the next change.