Test your Fini agent in layers: Suggested scenarios and ▶ Test Run for each Rulebook rule, Inbox conversations and Replays to verify individual fixes, and Test Suite sets as the regression check you run before every change ships. Each layer catches a different kind of failure, and together they let you change prompts, knowledge, and rules without customers finding the regressions first. This page is the playbook: which tool to use when, what to test, how to structure your test sets, and how to run them safely.

Pick the right tool for the job

The pattern is the one Inbox is built around: find a failure, trace it in AI Steps, fix it at the source, replay to confirm, then lock it in as a Test Suite scenario.

What to test

Cover each of these categories for every agent. Most teams start with the first three and add the rest within the first weeks.
Guardrail channel scope does not exclude dashboard tests, Test Suite, or replay evaluation. That means you can test a new guardrail policy against your scenarios before you enable it on a production channel.

Structure your test sets

Keep each test set focused on one behavior and one judge. A failing run then tells you immediately what regressed. A small, well-chosen set beats a large random one. When you write a CSV, the columns are prompt and expected (with an optional tag). Write each expected outcome as the arc you want across the whole conversation, not exact wording. For example: “Agent verifies the account, explains the refund window, and escalates to billing only if the order is outside it.”

Run the regression check before every change

Make this the release gate for any change to a prompt, knowledge source, Rulebook rule, Action, or model.
1

Test the change in isolation

For a rule change, run Suggested scenarios on the draft before publishing. For a knowledge or prompt change, open the conversations it should affect in Inbox and Replay them with Simulate rule execution.
2

Publish the change

A Test Suite run takes the agent state at the moment you click Run test, so publish first, or choose a saved Behavior version in the run dialog to test a specific version.
3

Run every affected test set

On each set’s detail page, click Run test and choose Simulate rule execution. It re-evaluates the latest rules while reusing saved results for matching external actions, and it is the recommended default because it avoids live side effects.
4

Read the run history, not just the latest number

A stable or rising pass rate is what you want. A drop without an intentional judge or scenario change means something regressed. Open the run matrix, expand the No pass rows, and read the judge’s reasoning.
5

Fix, then re-run

Fix at the source, click Run test again, and confirm the pass rate recovers before rolling out further.
6

Watch production after release

Check Analytics the following days for changes in AI resolution rate, escalation reasons, and CSAT. See Measure your resolution rate.
Choosing Execute live actions in a Test Suite run, or Run live actions in a replay, calls your real systems. If a scenario cancels a subscription or updates an address, it actually happens. Use Simulate rule execution by default, and point Actions at sandbox URLs for any bot you use for live-action testing.

Habits that keep tests useful

  • Add every fixed bug to a test set. Before you close the loop on a wrong reply, add the conversation with + Add from inbox and write the expected outcome. Your suite then grows with every failure you learn about.
  • Change one variable at a time. If you edit a prompt and a rule together, a pass-rate change can’t tell you which one caused it.
  • Edit the judge on purpose. Judges are versioned (judges v1, judges v2), so run history shows when a pass-rate change came from a stricter judge rather than a different agent.
  • Write outcomes as arcs. The customer side is simulated and not perfectly deterministic. Expected outcomes that describe the arc you want are stable across runs; outcomes that demand exact wording flap between Pass and No Pass.
  • Separate errors from failures. Simulations that hit the turn limit without resolving are marked as errors, not No Pass. A cluster of errors often points to an Action endpoint that is unreachable or returning errors, not a wrong answer.
  • Compare bots with duplicated sets. To A/B test a prompt or knowledge structure, duplicate the set for a second bot, change one thing, and compare run histories. See Comparing bots.

Automate it with the API

Teams that ship often run their regression check from CI or a scheduled job. Replays create separate replay conversations. They don’t overwrite the original conversation or apply any changes by themselves. See Replays for modes and model overrides, and API Keys for scopes.

Test Suite

Judges, scenarios, run history, and the run matrix.

Inbox

AI Steps, replays, feedback, and adding conversations to test sets.

Refine with AI

Turn a wrong reply into a reviewed fix.

Intent Rules

Suggested scenarios and manual test runs for each rule.