Test Arabic and English AI search on the same customer task
A product team tests its new assistant in English. The answer is clear, the sources open correctly, and the next action works. The Arabic version uses the same layout and passes a translation review.
Then a customer asks an Arabic question containing an English product name and a reference number. The answer looks plausible, but the reference is awkward to read and the linked document belongs to a different product version.
This illustrative case shows why a UAE product needs more than translated assistant copy. Test retrieval, mixed-language input, presentation, and the next action together. The useful unit of design is the customer’s task across languages.
What should the test set contain?
Start with questions that customers already ask. Where suitable evidence is available, include support themes, search queries, and terms used by customer-facing staff. Remove personal information and use approved examples.
For each task, prepare equivalent Arabic and English cases. Equivalent does not always mean a literal translation. Include the way people actually refer to the product, while keeping the expected outcome consistent.
A proposed test matrix can stay compact:
| Variation | What the team should inspect |
|---|---|
| Arabic question, Arabic source | Answer quality and the source it uses |
| English question, English source | The same task and expected outcome |
| Arabic question with an English product name | Product matching and readable presentation |
| Either language with a reference number | Exact preservation of the identifier |
| A source available in only one language | A clear explanation of what is available |
| No reliable source for the request | A useful fallback without an invented answer |
This is a starting worksheet. It does not establish that a product supports every dialect, language combination, or customer group.
Which problems belong to the interface?
Some failures come from retrieval or source coverage. Others come from the way the answer is displayed. Keep those findings separate so the right team can address them.
The W3C guidance on bidirectional text explains why right-to-left text mixed with left-to-right phrases, numbers, and identifiers needs careful markup. A fully translated interface can still display a mixed-language answer poorly.
Test the actual rendered result. Inspect product codes, amounts where relevant, dates, links, and copied text. Use the interface with keyboard navigation and the assistive technologies relevant to the intended audience. A screenshot alone cannot establish that the answer is usable.
Avoid turning a technical issue into a cultural assumption. If a participant struggles with a reference number, observe the problem before deciding it reflects a preference shared by an entire market.
How should the assistant recover?
Suppose the Arabic question retrieves an English document that answers the task. The product should make the source language clear and explain any translation it provides. If the underlying document is outdated, translating it well does not make it reliable.
Let the user refine the request or open the source without restarting the entire task. For a support handoff, carry the question and relevant context through the approved channel so the customer does not have to reconstruct the problem.
Google’s PAIR guidebook codelab recommends explaining AI behavior in ways that help people act and providing a way forward after an error. Here, a useful explanation might identify the available document or route the customer to the appropriate support path. It should not claim certainty the system cannot support.
What should a UAE design partner prove?
Ask who will research and review the Arabic experience, how participants will be recruited, and how the team will test mixed-language states. Verify the people doing the work and their relevant experience; a regional sales address does not establish language capability.
Also ask for experience integrating AI into an existing product. A standalone chatbot demonstration says little about the permissions, documents, and workflows it will inherit.
Humbleteam’s Dubai product design guide and Gulf partner evaluation guide give buyers a broader starting point. Use the matrix above to make the AI and language requirements explicit in the brief, and confirm any specialist capability before contracting it.
The first milestone should be a tested set of customer tasks with separate results for each language condition. If one fails, retain that finding. An average score across languages should not hide the experience of the customers who cannot complete their work.