A rubric your own agents pass or fail on. Then the AI has to pass the same one.
Vendors publish automation rates and CSAT scores. Very few publish how the number was measured. This page is the methodology — the one competitors can't copy in a week, because it starts with your own agents, not with our model.
You've heard 91% from another vendor. Or 87%, or 94%. You know intuitively that these numbers can't all mean the same thing — but you also can't tell the sales rep across the table exactly what's wrong with the way they measure. That's the gap this page is written for.
The honest answer is that "AI resolved 91% of chats" is not a statement, it's a claim about a rubric. And most rubrics in this category are undefined, or defined by the vendor after the number came out. What follows is how we do it instead.
The most common answer to "how do you measure the AI?" is CSAT — the little thumbs-up players give at the end of a chat. That's the wrong measurement for three reasons.
First, CSAT measures emotion, not correctness. A player can rate a chat five stars after receiving a factually wrong answer they liked hearing, and one star after receiving a correct answer they didn't. Both happen daily.
Second, in iGaming the player's emotional state is dominated by things support doesn't control — a losing session, a delayed withdrawal, a KYC document rejected on the fourth attempt, a responsible-gambling limit doing exactly what it's meant to do. Support gets the bill for all of it.
Third, the CSAT response rate is small and self-selected. Angry players and delighted players fill it in; everyone else doesn't. You're not measuring the middle 80% at all.
CSAT is worth watching. It's not worth building your evaluation of the AI on it.
Before we score any AI response, we score a sample of your existing human responses using the same rubric we're going to score the AI with. This is the number that matters — the human baseline on your book, not a benchmark from someone else's.
Three details make this measurement usable rather than decorative:
The 66% at PIN-UP / RedCore is not a slur on those agents — it's the honest score of a well-run human desk on a strict rubric. That's the ceiling anything AI has to clear.
Once the baseline exists, the AI is switched on parallel to the human, not instead of them. Every incoming conversation goes to a human agent as normal. The AI generates a response too, in the background, and stops there. The player never sees it.
Now we have the like-for-like comparison a demo can never give you: same conversation, same context, same customer, human answer next to AI answer. Both get scored on the same rubric by the same reviewers. This is the number that matters — the AI's score on your live traffic, not on a curated evaluation set the vendor prepared.
Shadow mode runs for a defined volume of traffic per topic — typically one to two weeks of live conversations, more on lower-volume topics. Nothing goes to production until the AI's score on a topic exceeds the human baseline on the same topic, and the gap is statistically real, not a noise-level 1%.
A single go-live is the moment where vendor claims meet reality, badly. We don't do it. Instead, when a topic clears its threshold in shadow mode, we route only that topic's live traffic to the AI, with a per-topic quality gate and an automatic rollback:
The rollout ends when there's nothing left to escalate that isn't genuinely a human's job. Which is not "100% automation" and never should be.
"You start with a rubric your own agents pass or fail on. Then the AI has to pass the same one. Then it has to keep passing it."
The labelled evaluation set the whole thing is scored against is yours at contract end, in the same format we use to run it. It's not "our proprietary benchmark" that walks out of the room with us.
That's the part that compounds. Every future model change — ours, or anyone else's — has to pass the same regression test. The evaluation set doesn't age the way a model does, because it's grounded in your own agents' work on your own topics. It's the closest thing to an insurance policy against being locked into a vendor whose numbers you can no longer verify.
It also gives you a defensible answer for the regulator, the compliance function and the board. When someone asks "how do you know the AI is answering RG topics correctly?", the answer is not "the vendor said so". It's "here's the rubric, here's the sample, here's this quarter's score, here's last quarter's".
Free PDF. Same framework, applied as a scorecard you can take to any support-AI conversation.
30 minutes. We look at your topic mix, your escalation rules and what you already measure, and tell you honestly how much of this we can run on your book before a rollout.
Or write: hello@arctura.eu