Measurement · How we prove it works

The methodology behind 66% → 91%.

A rubric your own agents pass or fail on. Then the AI has to pass the same one.

Vendors publish automation rates and CSAT scores. Very few publish how the number was measured. This page is the methodology — the one competitors can't copy in a week, because it starts with your own agents, not with our model.

The question an operator can't quite ask out loud

You've heard 91% from another vendor. Or 87%, or 94%. You know intuitively that these numbers can't all mean the same thing — but you also can't tell the sales rep across the table exactly what's wrong with the way they measure. That's the gap this page is written for.

The honest answer is that "AI resolved 91% of chats" is not a statement, it's a claim about a rubric. And most rubrics in this category are undefined, or defined by the vendor after the number came out. What follows is how we do it instead.

Why CSAT is not the answer

The most common answer to "how do you measure the AI?" is CSAT — the little thumbs-up players give at the end of a chat. That's the wrong measurement for three reasons.

First, CSAT measures emotion, not correctness. A player can rate a chat five stars after receiving a factually wrong answer they liked hearing, and one star after receiving a correct answer they didn't. Both happen daily.

Second, in iGaming the player's emotional state is dominated by things support doesn't control — a losing session, a delayed withdrawal, a KYC document rejected on the fourth attempt, a responsible-gambling limit doing exactly what it's meant to do. Support gets the bill for all of it.

Third, the CSAT response rate is small and self-selected. Angry players and delighted players fill it in; everyone else doesn't. You're not measuring the middle 80% at all.

CSAT is worth watching. It's not worth building your evaluation of the AI on it.

Step 1 — Baseline against your own agents

Before we score any AI response, we score a sample of your existing human responses using the same rubric we're going to score the AI with. This is the number that matters — the human baseline on your book, not a benchmark from someone else's.

Three details make this measurement usable rather than decorative:

  • Sampling is by topic, not by agent. A stratified sample across your real chat categories — deposit help, KYC status, bonus inquiry, RG concerns, withdrawal timing, technical issues, complaint escalation. Otherwise the strongest agent skews the number and the hardest topic disappears.
  • The rubric applies to both a human and a machine. Correctness against policy, completeness, tone, safety on RG-adjacent topics, respect for the escalation rule. A rubric that only works on AI answers is a rubric that hides the comparison.
  • Inter-rater agreement is measured. Two reviewers score the same sample; the number we report is the score after they agree, plus the disagreement rate. This turns a subjective judgement into a defensible one.

The 66% at PIN-UP / RedCore is not a slur on those agents — it's the honest score of a well-run human desk on a strict rubric. That's the ceiling anything AI has to clear.

Step 2 — Shadow mode on live traffic

Once the baseline exists, the AI is switched on parallel to the human, not instead of them. Every incoming conversation goes to a human agent as normal. The AI generates a response too, in the background, and stops there. The player never sees it.

Now we have the like-for-like comparison a demo can never give you: same conversation, same context, same customer, human answer next to AI answer. Both get scored on the same rubric by the same reviewers. This is the number that matters — the AI's score on your live traffic, not on a curated evaluation set the vendor prepared.

Shadow mode runs for a defined volume of traffic per topic — typically one to two weeks of live conversations, more on lower-volume topics. Nothing goes to production until the AI's score on a topic exceeds the human baseline on the same topic, and the gap is statistically real, not a noise-level 1%.

Step 3 — Staged rollout, one topic at a time

A single go-live is the moment where vendor claims meet reality, badly. We don't do it. Instead, when a topic clears its threshold in shadow mode, we route only that topic's live traffic to the AI, with a per-topic quality gate and an automatic rollback:

  • Each topic ships with its own quality threshold — the RG-adjacent thresholds are stricter than the "how do I change my password" ones, because the cost of a wrong RG answer isn't the same as a wrong password-reset answer.
  • Quality is measured continuously in production, not just at rollout. A sample of live AI responses per topic per day is scored by the same rubric.
  • If a topic drops below threshold — model drift, a knowledge-base gap, a change in player behaviour — that topic is automatically routed back to humans until the drop is diagnosed. Not the whole system; just that topic.

The rollout ends when there's nothing left to escalate that isn't genuinely a human's job. Which is not "100% automation" and never should be.

The measurement pipeline, in one picture

"You start with a rubric your own agents pass or fail on. Then the AI has to pass the same one. Then it has to keep passing it."

Step 4 — What stays with you

The labelled evaluation set the whole thing is scored against is yours at contract end, in the same format we use to run it. It's not "our proprietary benchmark" that walks out of the room with us.

That's the part that compounds. Every future model change — ours, or anyone else's — has to pass the same regression test. The evaluation set doesn't age the way a model does, because it's grounded in your own agents' work on your own topics. It's the closest thing to an insurance policy against being locked into a vendor whose numbers you can no longer verify.

It also gives you a defensible answer for the regulator, the compliance function and the board. When someone asks "how do you know the AI is answering RG topics correctly?", the answer is not "the vendor said so". It's "here's the rubric, here's the sample, here's this quarter's score, here's last quarter's".

Buyer's guide

The twelve questions to run this same test on any vendor.

Free PDF. Same framework, applied as a scorecard you can take to any support-AI conversation.

Get the guide

Bring your baseline. We'll bring the rubric.

30 minutes. We look at your topic mix, your escalation rules and what you already measure, and tell you honestly how much of this we can run on your book before a rollout.

Or write: hello@arctura.eu