The Test Suite catches regressions before customers do. You build a test set: customer scenarios graded by an LLM judge, and run it every time you change a prompt, knowledge article, rulebook rule, or model. Pass rate trends over time, so a regression shows up as a number going down. Find it under Ship → Test Suite in the sidebar.
Test Suite page in the Fini Demo workspace showing one test set named Venmo Core Regression, its conversation and criteria counts, creation date, and the New test set button

The model

A test set is three things bound together:
  1. Agent coverage: every conversation in a test set must belong to the same agent. The Test Suite list shows sets for the currently selected agent; older sets with no agent link appear under Other · not assigned to a bot.
  2. Judge prompt: one LLM judge that grades every conversation in the set. Fini ships six pre-built judges for common evaluation dimensions, and you can also write your own.
  3. Scenarios: the customer situations to test. Each scenario is a customer-side opening prompt plus an expected outcome describing how the whole multi-turn conversation should resolve. Fini generates scenarios from your data sources; you review and edit them before the first run.
At run time, Fini evaluates each stored conversation in the set, and the judge grades each conversation Pass or No Pass against the prompt.

Pre-built judges

Fini ships six judges out of the box, each covering a common evaluation dimension. Pick one when you create a test set and edit the prompt freely, or start from a blank judge with Write your own.
Step 1 of the test set wizard showing the six pre-built judges in a 3-column grid (Knowledge consistency, Escalation behavior, Tone and empathy, Tool use correctness, Safety and boundaries, Goal resolution) plus a Write your own option, with Escalation behavior selected and its judge prompt loaded into an editable textarea below

Knowledge consistency

Given a customer message and the agent reply, judge whether the agent answered using the correct knowledge article. Catches drift when the agent retrieves the wrong source or hallucinates around the gap.

Escalation behavior

Judge whether the agent escalated to a human at the right moment. Passes when escalation matches an explicit ask, a cancellation threat, or a compliance concern; fails on unnecessary or missed escalation.

Tone & empathy

Judge the agent’s tone, whether it sounds like your brand, acknowledges frustration, and matches the emotional register of the conversation.

Tool use correctness

Judge tool calls, whether the agent invoked the right Action at the right time, passed the right arguments, and handled the result correctly.

Safety & boundaries

Judge whether the agent stayed within scope. Off-topic deflections, declining legal or medical advice, refusing to discuss competitors, anywhere the agent should know when not to engage.

Goal resolution

Judge whether the conversation resolved the user’s actual intent, whether the customer left understanding the answer, completing the action, or being routed to the right human. The most general-purpose judge.
The judge prompt is fully editable. The pre-built versions are good starting points; tighten them to reflect your team’s specific bar for what counts as a pass.
Edit the judge over time, not just the scenarios. As you learn what “good” looks like for your bot, your judge prompt should sharpen too. Fini versions the judge automatically, you’ll see judges v1, judges v2, judges v3 in run history so you can tell whether a pass-rate change came from a better agent or a stricter judge.

Creating a test set

The wizard is four steps. Before you start, pick the bot you want to test from the agent picker. The Test Suite list is filtered by that selected bot, and newly created sets are linked to the bots represented by the conversations in the set.
1

Pick a judge prompt

On Step 1, pick one of the six pre-built judges or click + Write your own. The judge prompt drops into an editable textarea below, refine the wording to match your team’s standard for what counts as a pass. The whole prompt is shown to the LLM judge for every scenario in the set, so be specific about edge cases.
2

Pick data sources

Step 2 is where scenarios will come from. Three categories of sources, all composable:
  • Knowledge: two checkboxes for Help Center articles and Rule book scenarios. These ground the expected outcomes in your actual documented behavior. Check both when you want scenarios that test what your knowledge base says should happen.
  • Upload CSV: bring your own rows when you have hand-authored scenarios already. Columns: prompt, expected (optional: tag).
  • Seed from inbox conversations: pull from real customer conversations to ground the generated scenarios in actual phrasings. Filter by tag (churn, refund, bug, howto, policy), resolved-by (Bot / Escalated / All), and time window (Last 7d / Last 30d / All time). The footer shows how many conversations match.
Below the sources, Additional instructions lets you steer scenario generation in plain language, “Focus on edge cases where users mention competitor names. Skip simple greetings.”
Step 2 of the wizard showing Help Center articles and Rule book scenarios checked, an Upload CSV drop zone, and the Seed from inbox conversations section expanded with churn tag selected, All resolved-by, Last 30 days window, and 6 conversations matching
3

Review the generated scenarios

Fini generates one scenario per source it sampled, each as a CUSTOMER SAYS opening message plus an EXPECTED OUTCOME describing how the conversation should resolve. Each scenario has a stable ID (S01, S02, S03…) that persists across regenerations and run history, so a “S03 No Pass” comment in a team channel always points at the same scenario even after edits. Scenarios are grounded to their source, for example inbox · c-9116 for a conversation, knowledge · escalation policy §2 for an article section, so you can audit where each one came from.
Step 3 of the wizard showing four generated scenario cards, each with a Customer says message and an Expected outcome, plus Edit expected and Regenerate buttons per scenario
For each scenario:
  • Edit expected: refine the expected outcome to match what you’d actually accept as a pass.
  • Regenerate: ask Fini to produce a different scenario from the same source.
  • ×: remove the scenario from the set.
The expected outcome is prose, not a structured assertion. It describes the multi-turn arc you want, “Agent acknowledges frustration, offers retention path BEFORE escalation, then routes to billing-human if user re-confirms cancel.” The judge prompt reads the full simulated conversation against this expectation and decides Pass or No Pass.
4

Run and review

Step 4 is the launch panel: confirms the test set name, judge, scenario count, and sources. Click Run Fini and review to kick off the first run.
Step 4 of the wizard showing the Ready to run summary card with the test set name, judge, scenario count, and sources, plus the Run Fini and review launch button
Fini replays each scenario against your bot, simulating the customer’s responses across multiple turns until the conversation resolves. The judge then grades each conversation Pass or No Pass. The whole run typically takes under a minute for a small set; larger sets scale roughly linearly.

Scenarios are multi-turn

This is the part that surprises new users: each scenario isn’t a single customer message, it’s a full conversation that Fini simulates end-to-end. Fini plays the customer side based on the scenario’s prompt and expected outcome, your bot plays the agent side normally, and the dialogue continues until the conversation resolves (the agent succeeds at the user’s intent, escalates, hits a dead end, or the simulated customer’s patience runs out). Simulations cap at a maximum turn limit; conversations that hit the cap without resolving are marked as errors rather than No Pass. In the test set detail page, each conversation card shows the turn count, 8 turns, 14 turns, 11 turns, these are the actual back-and-forth exchanges Fini ran, not just the opening message. The judge grades the whole conversation, not just the first reply. A scenario can pass the first turn beautifully and still fail because the agent escalated unnecessarily on turn 4. That’s by design, the questions you actually care about (did the agent resolve the issue? did it stay safe? did it escalate at the right moment?) are properties of full conversations, not single messages.
Because the customer side is simulated, scenarios are reproducible enough for regression testing but not perfectly deterministic, the same scenario can produce slightly different conversation paths across runs. Write expected outcomes that describe the arc you want rather than exact wording, and trust the judge to evaluate semantic match.

A worked example: escalation behavior

A concrete walkthrough using the Escalation behavior, this week set from the screenshots above.
1

Pick the judge

On step 1, select Escalation behavior from the pre-built judges. The prompt drops into the textarea, start from the default and refine if your team has tighter rules.The default reads:
2

Pick data sources

Check Help Center articles (142 docs) and Rule book scenarios (21 rules) to ground expected outcomes in your documented escalation policy.Check Seed from inbox conversations and filter to TAG: churn, RESOLVED BY: All, WINDOW: Last 30d. 6 conversations match, these will seed scenarios with real customer phrasings around cancellation and frustration.Skip CSV upload, you don’t have hand-authored cases for this yet.Optionally add an instruction: “Cover both ‘user vents but doesn’t ask to escalate’ and ‘user explicitly demands manager’, those are the two failure modes I see most often.”
3

Review the six generated scenarios

Fini produces six scenarios, each grounded to its source. A few representative ones:Click Edit expected on any scenario to tighten the criterion, or Regenerate to draw a different scenario from the same source.
4

Run and review

Click Run Fini and review. Fini simulates each multi-turn conversation against your bot and grades each one with the Escalation behavior judge. The first run might pass 4 of 6; the failing two tell you exactly where your escalation logic needs work.

The test set detail page

After the first run, the test set detail page becomes your home base. It surfaces three layers of detail:

Latest run

A single pass-rate number, when the run completed, how long it took, and counts for passing, failing, and errored scenarios. The bar underneath is segmented by conversation result, so hovering a segment highlights the matching row in the matrix.

Run history

A pass-rate chart shows the most recent runs from oldest to latest. Completed runs render as percentage bars, active runs pulse, failed runs show a warning column, and All runs expands the full timestamped run log.

Run matrix

The latest completed run appears as a conversations-by-criteria matrix. Failed or errored conversations sort first, rows expand into judge reasoning and evidence, and each row links back to the source Inbox conversation.
The conversations grid on a test set detail page, showing a row of filter pills (All, Pass, No pass) above a multi-column grid of scenario cards, each labelled Pass or No Pass with a one-line summary of the conversation and the judge's reasoning snippet

Reading the run history trend

The trend across runs is what tells you whether the agent is actually improving. The pattern you want is a stable or rising pass rate as the agent changes. Falling pass rate without an intentional judge or scenario change means something regressed; click the affected run, expand the failing rows, and read which criterion flipped.

Reading the run matrix

The matrix answers the question “which conversation failed on which criterion?” at a glance. Columns come from the run’s actual criteria results; rows are conversations. A green check means the criterion passed, a red X means it failed, and a dash means it was not graded. When a column has failures, its header shows the pass count for that criterion. Use the All / Pass / No pass filter above the matrix to collapse the run to only the rows you need. Click a row to expand the full criterion-level reasoning and evidence, or click Open in Inbox to inspect the generated conversation with its AI Steps trace.

Running a test set

The test set detail page header has a Run test button. Click it any time the agent changes, a prompt update, a knowledge edit, a Rulebook rule change, an Action update, or a model swap. Fini opens a run dialog before queuing the test so you can choose how Rulebook logic should behave during the replay.
enum
Re-run the conversations using the same rule choices and saved rule results from the original replies. Behavior version selection is disabled in this mode.
enum
Re-evaluate the latest rules while reusing saved results for matching external actions. This is the recommended default because it avoids live side effects.
enum
Re-evaluate the latest rules and allow live external actions. Use this only when external side effects are acceptable for every conversation in the test set.
When you pick Simulate rule execution or Execute live actions, you can run against the latest published Behavior or choose a saved Behavior version. Each completed run records the mode and Behavior version used, so a pass-rate change can be traced back to the environment that produced it. The same button is also what you use after editing criteria or changing the conversations in the set. A run always does the full loop, replay the conversation + grade it, so the result is comparable across runs no matter what changed.
The default run takes the agent state at the moment you click Run test, so if you’ve just published a prompt change, the new run picks it up. If your agent is mid-rollout across channels, confirm the channel you care about reflects the latest config first.

Reading a failure

Open a failed run and expand any No pass row. You’ll see:
  1. The judge’s reasoning: a short paragraph explaining why this conversation failed against the judge prompt. “Agent escalated to billing-human on turn 2 without offering retention path first, despite the scenario expecting retention-then-escalate ordering.”
  2. The full simulated conversation: every turn, customer side and agent side, with the agent’s AI Steps trace attached to each agent message. AI Steps shows the chain of tool calls, knowledge retrievals, and routing decisions the agent made for that message, the same surface as the Inbox.
  3. The expected outcome: the scenario’s criterion, alongside the agent’s actual behavior, so you can see exactly where the divergence happened.
From there, fix the underlying issue at the source, adjust the Prompt, tighten an Article, refine a Rulebook rule, or update an Action’s inputs, then click Run test and watch the new pass rate land in history.

Adding scenarios from the Inbox

The Test Suite and the Inbox form a feedback loop. When you find a real production conversation that went badly, or one the agent nailed and you want to defend forever, click + Add from inbox on the test set’s Conversations grid. The selected Inbox conversation becomes a new scenario in the set: its opening customer message lands as CUSTOMER SAYS, and you author the EXPECTED OUTCOME describing what should have happened. From the next run onward, that production failure is a permanent regression check. This is the single highest-leverage habit with Test Suite: every time you fix something the agent got wrong, add the conversation to a test set before closing the loop. The set grows in lockstep with your understanding of how the agent fails.

When to run

  • On every meaningful change to the agent: a new prompt, an updated knowledge source, a Rulebook rule edit, an Action change, or a model swap. This is the high-leverage moment: catch the regression before it ships.
  • When you change what counts as a pass: edit the criteria or the conversation expectations, then run the set again so the next result reflects the new standard.
  • As a release-gate or pre-deploy check: many teams run the full set once before each release as a regression net, even if no individual change felt risky.

Comparing bots

A test set appears under each bot represented by its conversations. To compare bot A and bot B against the same scenarios, keep separate sets for each bot or duplicate the set and seed the duplicate with conversations from the comparison bot. You end up with comparable sets sharing the same judge and scenario pattern; the run history on each tells you which bot is improving faster. This is the pattern most teams use to A/B test prompt strategies, knowledge structures, or model choices, duplicate, change one variable on the new bot, re-judge both, compare.

Why a scenario isn’t passing

Generated expected outcomes are a starting point, not gospel. If the agent’s behavior looks correct but the scenario marks it No Pass, the expected outcome is probably too narrow. Edit the scenario, refine the criterion, then click Run all.
If multiple scenarios fail for the same reason and you disagree with the judge’s reasoning, the judge criterion likely needs tightening. Edit the criterion, then click Run all so the next run uses the updated standard.
The agent passes the first turn but fails by turn 6, usually because earlier turns set up state the later turns can’t recover from. Open the full conversation, find the turn where the path diverged, fix at that layer (often a prompt instruction or a Rulebook condition).
If a scenario invokes a Tool node whose Action endpoint is unreachable or returns an error, the conversation can’t complete normally and the scenario errors out. Either point your Actions at sandbox URLs for the test bot, or scope your judge prompts to behaviors that don’t require Action invocation.
Test sets are listed for the selected bot based on the set’s conversations. Switch the agent picker to the bot whose conversations seeded the set. Older sets without bot links appear under Other · not assigned to a bot.
The same scenario can produce slightly different conversation paths across runs. Most variation is small enough that the judge still arrives at the same verdict; if you see flapping between Pass and No Pass on identical runs, the expected outcome is probably too brittle, broaden it to describe the arc rather than exact wording.
Running a test set does not automatically dry-run your Actions. If a scenario triggers a Tool node that hits a real API, cancelling a subscription, updating an address, the side effect actually happens. Point your Action endpoints at sandbox URLs for any bot you use for testing, or scope your judge prompts to behaviors that don’t require Action invocation.