Pick the right tool for the job
The pattern is the one Inbox is built around: find a failure, trace it in AI Steps, fix it at the source, replay to confirm, then lock it in as a Test Suite scenario.
What to test
Cover each of these categories for every agent. Most teams start with the first three and add the rest within the first weeks.Structure your test sets
Keep each test set focused on one behavior and one judge. A failing run then tells you immediately what regressed. A small, well-chosen set beats a large random one.
When you write a CSV, the columns are
prompt and expected (with an optional tag). Write each expected outcome as the arc you want across the whole conversation, not exact wording. For example: “Agent verifies the account, explains the refund window, and escalates to billing only if the order is outside it.”
Run the regression check before every change
Make this the release gate for any change to a prompt, knowledge source, Rulebook rule, Action, or model.1
Test the change in isolation
For a rule change, run Suggested scenarios on the draft before publishing. For a knowledge or prompt change, open the conversations it should affect in Inbox and Replay them with Simulate rule execution.
2
Publish the change
A Test Suite run takes the agent state at the moment you click Run test, so publish first, or choose a saved Behavior version in the run dialog to test a specific version.
3
Run every affected test set
On each set’s detail page, click Run test and choose Simulate rule execution. It re-evaluates the latest rules while reusing saved results for matching external actions, and it is the recommended default because it avoids live side effects.
4
Read the run history, not just the latest number
A stable or rising pass rate is what you want. A drop without an intentional judge or scenario change means something regressed. Open the run matrix, expand the No pass rows, and read the judge’s reasoning.
5
Fix, then re-run
Fix at the source, click Run test again, and confirm the pass rate recovers before rolling out further.
6
Watch production after release
Check Analytics the following days for changes in AI resolution rate, escalation reasons, and CSAT. See Measure your resolution rate.
Habits that keep tests useful
- Add every fixed bug to a test set. Before you close the loop on a wrong reply, add the conversation with + Add from inbox and write the expected outcome. Your suite then grows with every failure you learn about.
- Change one variable at a time. If you edit a prompt and a rule together, a pass-rate change can’t tell you which one caused it.
- Edit the judge on purpose. Judges are versioned (judges v1, judges v2), so run history shows when a pass-rate change came from a stricter judge rather than a different agent.
- Write outcomes as arcs. The customer side is simulated and not perfectly deterministic. Expected outcomes that describe the arc you want are stable across runs; outcomes that demand exact wording flap between Pass and No Pass.
- Separate errors from failures. Simulations that hit the turn limit without resolving are marked as errors, not No Pass. A cluster of errors often points to an Action endpoint that is unreachable or returning errors, not a wrong answer.
- Compare bots with duplicated sets. To A/B test a prompt or knowledge structure, duplicate the set for a second bot, change one thing, and compare run histories. See Comparing bots.
Automate it with the API
Teams that ship often run their regression check from CI or a scheduled job.
Replays create separate replay conversations. They don’t overwrite the original conversation or apply any changes by themselves. See Replays for modes and model overrides, and API Keys for scopes.
Related
Test Suite
Judges, scenarios, run history, and the run matrix.
Inbox
AI Steps, replays, feedback, and adding conversations to test sets.
Refine with AI
Turn a wrong reply into a reviewed fix.
Intent Rules
Suggested scenarios and manual test runs for each rule.

