Skip to main content

How to audit an AI copilot that users don’t adopt

Sergey Krasotin
Design Director

Published

An AI copilot with low adoption needs a task-level diagnosis before a new interface. Check whether eligible users notice it, try it at a useful moment, get an acceptable result, and return when the task happens again. Each break calls for a different response.

The workflow below is a suggested audit structure for product teams. It is not a claim that every AI feature should become a habit: a quarterly reporting assistant has a different usage pattern from a daily support tool.

What counts as adoption for this particular task?

Start with a sentence that describes the job without mentioning AI. For a hypothetical account management product, that might be: “Prepare an accurate account update before a customer meeting.” Then define what makes the update acceptable and who can judge it.

Separate eligibility from activity. An account manager without a scheduled meeting may have no reason to use the feature this week. Counting that person as an adoption failure would mix lack of need with poor UX.

Write down the existing way to complete the task. Include the time spent checking and correcting the copilot’s output when comparing approaches. A fast draft that requires extensive repairs may leave the user with more work.

Where does the adoption funnel break?

Use product events to locate a break, then observe sessions to understand it. Events alone cannot explain why somebody abandoned an answer.

  • Eligible opportunity: the user has the task and access to the necessary data.
  • Discovery: the user encounters the entry point while doing related work.
  • Attempt: the user supplies enough context to request a result.
  • Useful completion: the user checks the result and completes the underlying task.
  • Return: the user chooses the feature again at the next relevant opportunity.

Define the events with engineering before reading the funnel. Opening a sidebar is not the same as attempting a task. Copying an answer is not evidence that the answer was correct. Keep those differences visible in the report.

How can the team tell a UX problem from an output problem?

Review representative interactions with permission and appropriate data handling. Ask a domain expert to describe each failure in plain language. “The summary attached the wrong renewal date” is more actionable than “The AI quality was poor.”

In their guide to building AI evaluation systems, Hamel Husain and Shreya Shankar describe using error analysis to identify concrete failures before choosing evaluation metrics. That supports a useful audit habit: inspect the actual work before deciding what to count.

For each failed task, record the input, the displayed output, the expected result, and what the user did next. Classify the likely cause without forcing certainty. Missing source data, confusing controls, and an incorrect answer may contribute to the same failure.

In the hypothetical account update, a wrong renewal date could come from an outdated document or from the model choosing the wrong source. Adding a clearer button would address neither cause. By contrast, a correct update hidden behind an unfamiliar icon creates a discovery question worth testing with design.

What does one useful error-analysis record look like?

Illustrative example, not a client result: an account assistant summarizes a renewal as October 31. The approved contract says September 30, while an older sales note contains the October date. The reviewer marks the answer unacceptable because it could change the account team’s next action.

  • Save the question, permitted source documents, returned answer, and expected date together.
  • Label the observed failure narrowly: the answer used an outdated source. Keep the underlying cause open until retrieval and answer generation are inspected.
  • Ask whether the user could find the supporting passage and correct the answer before using it.
  • After a change, repeat this example alongside correct answers and cases with genuinely conflicting records. A fix should not make the assistant conceal uncertainty.

Sources: Hamel Husain and Shreya Shankar on building AI evaluation systems

What should a task-based evaluation brief contain?

Use the following fields for every task selected for a redesign test. Complete them before drawing a replacement screen.

  • The user role, triggering situation, and current way of completing the task.
  • The source material the system may use, including missing or conflicting information.
  • The result a domain expert would accept and the errors that would make it unusable.
  • The steps the user must take to inspect, correct, and apply the result.
  • The observed problem, the proposed change, and the evidence that would disprove the proposal.
  • The owner of the design change and any separate data or engineering work.

Include an ordinary successful task as well as failure examples. Otherwise, the team may improve an awkward edge case while making the common task slower.

What should change first?

Choose a bounded test based on the diagnosed break. If users cannot anticipate the result, test a task-specific label and preview. If they cannot judge the answer, test source access and correction controls. If the answer is unsuitable, agree on the output or data repair before presenting a visual redesign as the solution.

Repeat the same task evaluation after the change. Record improvements alongside new failure modes, and compare repeat use only when users have had another relevant task opportunity.

For help defining the scope, see Humbleteam’s AI feature UX audit or compare AI product design agencies. If the copilot takes actions rather than drafting answers, use the separate guide to approval, undo, and human handoff to specify those controls.

Let's talk

Have questions? Ask AI
Opens a new chat with context about us pre-loaded — ask anything

We’ll reply within 24 hours with case studies, a timeline, and an estimate.

Prefer email? Write to hi@humbleteam.com
All set – our team’s on it. Expect a reply soon.
Send another one
Oops! Something went wrong while submitting the form.
Have questions? Ask AI
Opens a new chat with context about us pre-loaded — ask anything
Europe
Národní 135/14, Prague
Middle East
UAE, Dubai, Internet City Offices
We use cookies to enhance your browsing experience,
serve personalised ads or content, and analyse our traffic.
By clicking "Accept All", you consent to our use of cookies.
Privacy policy