Fini (usefini.com) states a benchmark of up to 90% resolution at 99% accuracy across its deployments. In the product, resolution is measured precisely: AI resolution rate is the share of conversations whose status is Resolved by AI, and you can check accuracy yourself on your own traffic with Inbox feedback, Test Suite pass rates, guardrail activity, and CSAT. The two numbers answer different questions. Resolution rate tells you how much of your volume the agent closed without a human. Accuracy tells you whether what the agent said and did was correct. A high resolution rate with poor accuracy is a liability, especially in fintech, banking, insurance, and healthcare, so this page covers both and shows where each one lives in the product.
The benchmark is Fini’s stated figure across deployments, not a guarantee for every account. Your own numbers depend on your knowledge coverage, the workflows you automate, the Actions you connect, and how much of your traffic needs a human by policy. The rest of this page shows you how to measure your own.

What the headline numbers mean

To confirm (internal, remove before publish): How is the 99% accuracy figure measured (sample, grading method, who grades, time period, which deployments)? No methodology is published in the docs or on www. Until it is, this page describes 99% as a stated benchmark only.
To confirm (internal, remove before publish): www.usefini.com/trust-metrics lists “Resolution Accuracy” at 95%, while the homepage and llms.txt say 99% accuracy. Which figure is current, and are they the same metric? Also, the trust-metrics page calls resolution rate a “vanity metric”, which conflicts with the resolution-first position in these docs. Should that page be updated?

How resolution rate is computed in the product

Every conversation the agent handles ends in exactly one of three statuses. The status comes from the agent’s Output Tag Selection on the conversation (the Conversation Status category), which you can inspect in the AI Steps trace. The three statuses sum to 100% of conversations in the selected window. From them, Analytics derives three rates: Because deflection includes conversations that are still waiting on the customer, deflection rate is always greater than or equal to AI resolution rate. The gap between the two is your Waiting for Customer share. Resolution vs deflection walks through a worked example. Where you see each number:
  • KPI cards at the top of Analytics show Deflection rate and Human escalation rate with period-over-period change pills.
  • The Resolution rate trend view plots the daily resolution rate alongside the Waiting for Customer share.
  • The Conversation Status doughnut shows the full three-way split.
  • Knowledge performance and Intent rule breakdown show AI Resolve Rate and Escalated Rate per knowledge slice and per Rulebook intent rule.
  • The Get agent analytics API returns aiResolutionRate, humanEscalationRate, and the raw counts (resolvedConversations, escalatedConversations, waitingForCustomerConversations, totalConversations).
For a step-by-step procedure, see Measure your resolution rate.

Why a resolved conversation is not automatically an accurate one

Resolved by AI means the agent closed the conversation without a human. It does not, on its own, prove that every reply was correct. A customer can accept a wrong answer and leave. That is why Fini pairs resolution with accuracy signals you control, and why every number in Analytics traces back to individual conversations you can open in Inbox and audit reply by reply.

Assessing accuracy yourself

You don’t have to take any benchmark on trust. Fini gives you several independent ways to grade the agent’s accuracy on your own conversations. Use more than one: each catches a different kind of error.
To confirm (internal, remove before publish): In the product, AI CSAT is an AI-inferred CSAT tag on a conversation. On www and in llms.txt, “AI CSAT 4.9 vs human CSAT 4.7” appears to compare CSAT on AI-handled conversations against human-handled ones. Are these the same measure? If not, which one does the 4.9 vs 4.7 figure use?
The signals work as one loop: production traffic produces evidence of errors, you diagnose and fix them, and Test Suite locks each fix in so it can’t quietly come back.

A simple accuracy audit you can repeat

1

Draw a sample

In Inbox, select the agent and the last 7 days, set Fini Touched to Yes, and filter Conversation status to Resolved by AI. Pick a fixed number of conversations at random (for example, 50) rather than the ones you remember.
2

Grade every Fini reply

For each conversation, read the thread and open AI Steps on any reply you are unsure about. Mark each reply thumbs up or thumbs down. Add a Feedback note to every thumbs down that says what the correct answer was.
3

Compute your accuracy rate

Divide the replies you marked correct by the replies you graded. Record the result with the date range, the agent, and the sample size so the next audit is comparable.
4

Fix and lock in

Run Refine with AI on the wrong replies, approve the fixes you agree with, then add each conversation to a Test Suite so the error becomes a permanent regression check. Then mark the feedback as actioned, so the Feedback Actioned filter in Inbox shows which thumbs-downs are addressed.
Grade escalated conversations too, with a different question: should the agent have escalated? Unnecessary escalations lower resolution rate; missed escalations are an accuracy and compliance problem. The Escalation behavior judge in Test Suite automates this check for scenarios you define.

What these numbers do not tell you

  • Analytics does not grade correctness. It counts statuses, escalation reasons, and ratings. Correctness comes from your review, Test Suite, and guardrails.
  • Test Suite is a fixed sample. A passing set means the agent handles those scenarios correctly. It does not cover traffic you haven’t written scenarios for, which is why production review still matters.
  • Guardrails are not a fail-closed boundary. A check can report could not run and the original reply can still be sent. See Guardrails.

Resolution vs deflection

The exact definitions, a worked example, and why Fini leads with resolution.

Measure your resolution rate

Step-by-step in Analytics and through the API, including before-and-after comparisons.

Analytics

Every KPI card, chart, and breakdown on the Analytics page.

Test Suite

LLM-judged regression evals that track pass rate as you ship.