BUILD WITH YOUR CUSTOMERS IN MIND
Test AI agent experiences with synthetic customers
Dori helps builders examine customer-facing AI agents through a browser-accessible interface. Synthetic participants can explore whether a chat or agent experience understands their goal, explains its actions and helps them recover when an interaction goes wrong.
THE SHORT ANSWER
What does customer-facing AI agent testing evaluate?
It evaluates the experience a person has while using an agent: task clarity, questions, responses, visible actions, trust and recovery. Start with a realistic customer goal, observe the interaction and review the evidence behind each finding before deciding what to change.
When to use it
Use it for a web chat assistant, an agent embedded in your product or a browser-accessible agent workflow. Product teams and builder agents can use the observations to refine instructions, interaction design and how the interface communicates progress or uncertainty.
How to approach the test
Choose a customer goal
Use a concrete task such as finding an option that meets a budget, preparing a draft or understanding a recommendation. Define what a useful outcome would look like from the customer perspective.
Vary context and expectations
Explore participants with different familiarity, constraints and communication styles. Test ambiguity that a person could reasonably introduce without hiding the core task.
Review the full interaction
Inspect the prompts, replies and visible state changes. Check whether the agent asks useful questions, distinguishes proposed actions from completed actions and provides a clear path when it cannot proceed.
Turn evidence into a revision
Separate interface problems from response-quality problems. Adjust the flow, instructions or copy, then retest and use a broader evaluation programme for model reliability.
ILLUSTRATIVE EXAMPLE · NOT A LIVE TEST RESULT
From an observation to a product change
- task
- Ask an assistant to prepare a shortlist within a stated budget.
- participant
- A customer who expects to review options before taking an action.
- observation
- The assistant gives recommendations but does not explain which options satisfy the budget or what information is missing.
- evidence
- Illustrative transcript: budget constraint in the request → recommendations without a budget comparison → follow-up asking which options qualify. A real report should reference the actual exchange.
- interpretation
- The customer cannot confidently judge whether the response meets the constraint.
- change
- Make the budget comparison explicit and ask for missing information before claiming the goal is complete.
What to look for
- Goal understanding and clarification
- Response clarity and visible constraints
- Communication about actions and uncertainty
- Error recovery and next steps
What this test can and cannot tell you
This page describes testing an agent through its web interface. It is not a direct API benchmark, a comprehensive security assessment or proof of factual correctness. Synthetic participants can miss failures or share model biases; use human review, domain-specific evaluations and security testing where the decision requires them.
Common questions
Can I test an agent without a web interface?
This Dori workflow uses a browser-accessible product surface. Direct API evaluations need a separate test harness and suitable evaluation criteria.
What makes an agent finding actionable?
A finding should show the customer goal, the relevant exchange or visible behaviour, why it created difficulty and a specific revision to investigate.
Does a successful synthetic task prove reliability?
No. Agent behaviour can vary with wording, context and model updates. Test multiple scenarios and combine observations with task-specific evaluations and human review.
Build a better next version.
Explore Dori and create your account. Test execution depends on current availability and account access; check your account before planning a study.
Create your account