Skip to main content
  • Other

How to test changes to an AI design workflow

Changes to prompts, models, and components can bring back old mistakes. Build a practical test collection for an AI design workflow and keep it useful after handoff.

ShareLinkedInXEmail

The team changes a prompt to make its AI-generated prototypes more consistent. The next demo looks better. A week later, someone notices that the new workflow has stopped producing the permission-denied state.

The example is invented, but the problem is easy to miss: improving the output you are looking at does not establish what happened to the rest of the work. A design team needs a repeatable way to check changes to its prompts, models, component libraries, and source material.

This is where an AI workflow agency should be able to help after the pilot. Ask for a small set of checks tied to actual design failures, plus a way for the team to maintain it. A successful demonstration is only the beginning of that arrangement.

What needs to stay true when the workflow changes?

Choose one workflow with a defined output. For this example, the input is an approved feature brief and a component library; the output is a prototype ready for a designer’s review. Publication and implementation remain separate decisions.

Write down what makes that output usable. It must use the supplied roles accurately, represent required states, and keep unsupported assumptions visible. A reviewer needs to distinguish an approved fact from something the model filled in.

In Building eval systems that improve your AI product, Hamel Husain and Shreya Shankar explain how examining real failures can inform evaluation. Their guidance distinguishes checks for known problems from the ongoing work of finding new ones. For a design team, that means preserving examples of mistakes people actually had to correct, then using them when the workflow changes.

An evaluation is simply a test of whether the output meets an explicit requirement. The useful question is specific: does this prototype give a viewer access to an editor-only action?

What belongs in a small test collection?

Start with sanitized examples the team understands. Keep the original input, the requirement, the observed failure, and the decision a reviewer should make. Remove client information that is not approved for the tool or the test collection.

The following matrix is an illustrative starting point for the prototype workflow. Replace its examples with your own.

Illustrative test collection for an AI prototype workflow
Test inputFailure to catchHow to check it
Brief with viewer and editor rolesViewer receives an editing actionInspect both role paths against the supplied rules
Existing form with a rejected submissionError state disappearsCheck that the state remains present and understandable
Brief with an unanswered policy questionModel invents a definitive ruleVerify that the uncertainty stays marked for review
Updated component libraryPrototype uses a retired componentCompare component references with the approved library
Content containing a long customer nameLayout hides the main actionInspect the rendered output at relevant widths
Brief revised after initial generationOutput silently uses an older requirementTrace the affected decision to the current input

This collection covers different kinds of failure. Some checks can run automatically, such as finding a retired component reference. Others need someone to inspect behavior and meaning. A machine can confirm that an error message exists without establishing that a user knows what to do next.

How do you compare the old and new versions?

Keep the inputs fixed while you compare the proposed change with the current workflow. Record the configuration each output used, including the relevant prompt and component-library versions. Otherwise, a better result could reflect a different brief rather than the change being tested.

Agree on the acceptance rules before looking at the outputs. For example, missing a required permission boundary can block adoption even if the new version produces cleaner layouts. Other differences may be acceptable with a named reviewer and a limited scope.

When generation varies, repeat the comparison enough to understand the instability that matters for this workflow. Preserve failures as well as attractive outputs. There is no universal run count that makes every design task reliable; choose the effort in proportion to the consequences and document the remaining uncertainty.

If the team uses an AI model to judge the output, compare that judgment with human reviews on examples it was not tuned against. Do not let an agreeable second model turn a missing state into a passing result.

What happens when a change fails?

Suppose the revised prompt improves layout consistency but drops the permission-denied state in the test example. Keep the current workflow in use while the team investigates. Record the failing input and the behavior that must change.

After a fix, rerun the relevant checks and check the other cases the fix could affect. A patch that adds a permission state everywhere may create a different problem in a flow where that state has no meaning.

Preserve a way to return to the previous usable setup where the tools allow it. If a provider removes a model or changes behavior outside the team’s control, use a documented fallback, such as manual completion of the affected step. A rollback plan should describe an available action.

Who keeps the checks useful after handoff?

Assign an internal owner before the agency leaves. That person needs access to the approved inputs, the test collection, the current configuration, and the record of unresolved failures. Test the handoff by asking a colleague to evaluate a small change without help from the original builders.

Keep reviewing ordinary work after the checks pass. When designers find a new recurring failure, decide whether to add an example or revise an existing requirement. A test collection that never changes gradually becomes a record of what the team used to care about.

The AI workflow pilot guide helps define the first adoption experiment; the handoff contract covers what the receiving team needs. This next step makes change review repeatable. The team should finish with a test collection it can run, an owner who can interpret it, and a clear reason to accept or reject the next change.

Daniel Mercer

AI product and design workflows

Published by Humbleteam.

Back to top

Let's talk

Have questions? Ask AI
Opens a new chat with context about us pre-loaded — ask anything

We’ll reply within 24 hours with case studies, a timeline, and an estimate.

Prefer email? Write to hi@humbleteam.com
All set – our team’s on it. Expect a reply soon.
Send another one
Oops! Something went wrong while submitting the form.
Have questions? Ask AI
Opens a new chat with context about us pre-loaded — ask anything
Europe
Národní 135/14, Prague
Middle East
UAE, Dubai, Internet City Offices
We use cookies to enhance your browsing experience,
serve personalised ads or content, and analyse our traffic.
By clicking "Accept All", you consent to our use of cookies.
Privacy policy