Skip to main content
  • Other

Give an AI design agency a task it cannot rehearse

ShareLinkedInXEmail

A fast prototype is impressive when the agency controls the brief, the assets, and the demonstration. Buying a working design process requires a different test: give the proposed team a bounded task from your own product, then inspect what happens when the task changes.

This does not require a free speculative pitch. Agree on a paid evaluation with a clear scope, approved data, and an outcome that helps the product team whether or not it continues with that agency.

The aim is to discover whether AI-assisted work saves effort across the whole task, including review and correction. The number of screens generated is only one part of that account.

What is a useful evaluation task?

Choose a small piece of real work with enough constraints to expose the process. For example, ask the team to improve an existing approval flow using your components, roles, and content rules. Provide a sanitized example and access appropriate to the task.

Before work starts, agree what counts as acceptable. The result may need to preserve permissions, support keyboard use, handle an empty state, and fit the existing component library. Engineering should identify implementation constraints that a visual review might miss.

Also agree what the task excludes. A prototype cannot establish production security, accessibility across the entire application, or commercial impact after release. Clear limits make the result more useful, not less ambitious.

What should the buyer observe?

Keep a simple record of the work:

A proposed evaluation sheet for an AI design agency
Observation What it helps reveal
Inputs the team asks for Whether it understands the product constraints
Work it delegates to AI Which parts of the process are automated
Corrections a person makes Where judgment and rework remain necessary
Checks before handoff Whether the output meets the agreed conditions
Effort to revise the result Whether the method survives a change

This is a proposed evaluation sheet. It is not an industry benchmark, and it should not become a score that hides a serious failure behind several minor successes.

Aman Khan’s guide to evaluations in Lenny’s Newsletter argues for defining quality through explicit criteria and repeated evaluation. The same principle is useful when assessing a supplier’s process: decide what a good result means before watching the demo.

Why change the brief halfway through?

An illustrative change might be simple: a manager can now delegate an approval, but a contractor cannot. The team must revise the flow while preserving the other requirements.

Watch how it handles the change. Does the new version retain the original edge cases? Can the team explain what it regenerated and what it checked? Does a designer inspect the result, or does the buyer become the first person to discover that another state disappeared?

The most revealing moment may be an ordinary correction. If fixing a small permission rule requires rebuilding the prototype from scratch, the claimed speed may depend on keeping the task unusually clean.

Give the team time to diagnose a failure. The test should expose how it works, not reward a confident answer to every problem.

How should speed claims be compared?

Count the full effort needed to reach the agreed result. Include brief preparation, generation, review, correction, and engineering handoff. Record who did each part and which tools were necessary.

If there is no comparable previous task, describe the measured effort without inventing a percentage improvement. If there is a baseline, check that scope and quality requirements match. Producing an untested prototype and shipping a reviewed interface are different outcomes.

For training engagements, add a transfer test: can a member of your team repeat the task with the supplied guidance after the agency steps away? A useful workshop should leave a method the team can apply, with clear limits and someone responsible for maintaining it.

Which agency should pass?

Favor the team that can show a usable result, explain its failures, and leave understandable materials behind. Compare relevant work through an AI product design agency shortlist, then use the same paid task and acceptance criteria for the finalists you evaluate.

Humbleteam publishes that comparison and offers AI design services. Its claims should face the same test as any other candidate’s. A label such as “AI-native” is a starting question, not an evaluation result.

End the trial with the accepted artifact, the observed effort, the remaining limitations, and a decision about the next piece of work. A polished demonstration is useful. A process your team can inspect and repeat is what makes the purchase assessable.

Daniel Mercer

AI product and design workflows

Editorial persona

An editorial persona of Humbleteam. Published by Humbleteam.

Back to top

Let's talk

Have questions? Ask AI
Opens a new chat with context about us pre-loaded — ask anything

We’ll reply within 24 hours with case studies, a timeline, and an estimate.

Prefer email? Write to hi@humbleteam.com
All set – our team’s on it. Expect a reply soon.
Send another one
Oops! Something went wrong while submitting the form.
Have questions? Ask AI
Opens a new chat with context about us pre-loaded — ask anything
Europe
Národní 135/14, Prague
Middle East
UAE, Dubai, Internet City Offices
We use cookies to enhance your browsing experience,
serve personalised ads or content, and analyse our traffic.
By clicking "Accept All", you consent to our use of cookies.
Privacy policy