Home/Blog/Why Random Ticket Sampling Distorts Your Agent Rankings
Metrics

Why Random Ticket Sampling Distorts Your Agent Rankings

Support leaders make real decisions on QA scores: coaching allocation, bonus payouts, promotion decisions, performance improvement plans. If those decisions are based on 2% samples, they are based on numbers with confidence intervals wider than most of the differences between agents. This is not a philosophical objection. It is arithmetic.

What is statistical power and why does it matter for QA?

Statistical power is the ability of a measurement to detect an effect that is actually there. A QA score has statistical power when the number of observations is large enough that the noise in individual tickets does not swamp the underlying signal of an agent's true performance level.

Here is the shape of the problem. Assume rubric scores per ticket have a standard deviation of about 15 points around an agent's true mean, which is realistic for support QA. With 10 tickets sampled, the standard error of the mean is roughly 15 divided by the square root of 10, or about 4.7. The 95% confidence interval around the observed score is plus or minus 9.2 points.

That means an agent whose observed monthly score is 87 could have a true performance level anywhere from 78 to 96. The difference between an 87 and a 90 agent is well inside the noise. Any ranking based on that gap is a coin flip.

How much do agent rankings actually shift month to month?

More than most support leaders expect. When we look at the same 20 agents across two consecutive months on a 2% sample:

Ranking shift % of agents
Moved 0 to 2 positions 25 to 35%
Moved 3 to 5 positions 30 to 40%
Moved 6+ positions 25 to 35%

Roughly a third of agents shift more than five positions in a 20-agent ranking every month, with no underlying performance change. This is what sampling noise looks like in the wild.

The consequences ripple outward. Agents notice the shuffling and lose faith in the process. Managers construct explanatory narratives ("Alex had a rough month") that fit the noise. Bonus programs pay agents who were lucky and skip agents who were unlucky. None of this is bad intent. It is what happens when the measurement is under-powered.

Why does full-coverage scoring fix the math?

Because the sample is the population. When every ticket is scored, an agent's monthly average is their actual performance, not an estimate of it.

Concretely, an agent who handles 200 tickets a month gets 200 scored data points across each rubric dimension. The standard error of the mean drops to roughly 15 divided by the square root of 200, or 1.1. The 95% confidence interval collapses to about plus or minus 2 points.

The math changes what you can say about the numbers.

  • Ranking stability. Month-over-month rankings become stable, because the underlying scores are stable. Movement now signals actual behavioral change, not sampling luck.
  • Fair comparison. The difference between an 88 agent and a 91 agent is now statistically real, so promotion and bonus decisions can defensibly reference the number.
  • Trend detection. A 3-point score drop over 60 days is now a trend, because 3 points is outside the noise floor. On a 2% sample, that same 3-point drop is indistinguishable from variance.
  • Coaching precision. Each agent's per-dimension score reflects hundreds of conversations, so weaknesses in specific dimensions are real and coachable.

Statistical power is not an academic property. It is what makes the QA numbers something a support leader can act on without hedging.

What about "the sample is representative"?

Representative of what? A random 2% sample is representative of the average conversation. It is not representative of an individual agent's month, because at 8 to 12 sampled tickets per agent, the sample is too small to represent that agent's distribution.

The distinction matters. If your question is "what is our team's average tone score," a 2% sample answers it with reasonable confidence. If your question is "which of our 20 agents needs tone coaching," the same sample cannot answer, because the per-agent slices are too thin.

Most QA programs mix these two use cases and quietly generalize the aggregate confidence to individual conclusions. The math does not support that generalization.

How do sampled scores distort coaching allocations?

Two failure modes, both common.

  • Coaching the wrong agent. An agent gets flagged for tone coaching because one of their 10 sampled tickets happened to be a hostile customer where the agent was curt. The other 190 tickets that month, all fine on tone, were not sampled. The coaching lands on a behavior the agent does not actually have.
  • Missing the actual coaching moment. An agent has been quietly quoting a stale policy across 15 tickets this month. None of those 15 landed in the 2% sample, so the pattern is invisible. The coaching that should have happened does not, and the ticket reopens compound over the next quarter.

Both failures come from the same source: individual-level decisions made on aggregate-level power. Full-coverage scoring closes both by making per-agent samples equal to per-agent volume.

Are there cases where sampling is appropriate?

Yes. Three, specifically.

  • Team-level trend reporting. If you need to know whether tone scores went up or down between June and July across the whole team, a 2% sample supports the conclusion with reasonable confidence. This is the aggregate use case and it is legitimate.
  • Compliance audits for low-risk categories. For ticket categories where the downside of an unscored bad ticket is small, sampling is defensible on cost grounds. Refund policy quotes are not one of these categories. Password reset tickets probably are.
  • Very small teams. Below 20 agents, a 2% sample can be manually expanded to 8 to 15% by dedicating enough grader time. That coverage level starts to be statistically usable for per-agent decisions.

Above 20 agents, none of these justify individual-level decisions on 2% samples. The math does not support the practice at scale.

What changes with per-agent statistical power?

Concrete outcomes we see repeatedly in teams that switch from sampling to full coverage.

  • Bonus disputes drop. Agents accept scores that are backed by their whole month, not by 10 tickets. Formal bonus disputes drop 60 to 80%.
  • Ranking narratives fade. Managers stop constructing month-to-month explanations for random shuffling. Time saved goes to coaching.
  • Under-the-radar patterns surface. Agents who were quietly making the same policy error across many tickets get identified because the pattern is visible across all of their conversations, not the ten that were sampled.
  • Onboarding scores stabilize faster. A new agent's 90-day trend line is a real signal at week eight, not a noisy estimate at week 20.

The through-line is trust in the number. QA programs that ran on sampled scores were used cautiously, with implicit hedging. Full-coverage scores can be used directly, because the confidence interval is small enough to defend the decision.

The mistake to avoid

Do not use aggregate-level sampling to make individual-level decisions. A 2% sample tells you what your team looks like on average; it cannot reliably tell you which specific agents to coach on which specific behaviors, or which agents deserve top or bottom rankings. The choice is not between rigorous and lax QA; it is between measurement that supports the decisions you are making, and measurement that does not. Full-coverage scoring is not a nice-to-have. It is what makes per-agent decisions defensible.

qa-samplingagent-rankingsstatistical-powersupport-metrics

Frequently asked questions

Why does sample size matter so much for agent scores?

Because rubric scores have natural variance. Even a consistently strong agent will handle some tickets better than others. With 8 to 12 sampled tickets a month, you cannot separate the agent's true performance level from the specific tickets that happened to be sampled. Statistical power requires enough observations to overcome variance, and 8 to 12 is far below that threshold for the effect sizes support leaders care about.

What is a confidence interval on an agent score?

It is the range within which the agent's true performance level is likely to fall. With 10 sampled tickets, an observed score of 87 typically has a 95% confidence interval of roughly 79 to 95. That means the agent's true level could plausibly be anywhere in that range, and a difference between an 87 agent and an 84 agent is not statistically distinguishable.

Doesn't stratified sampling help?

It helps at the team level for identifying category-level trends, and it does not help at the agent level for stable rankings. Even with stratification, each agent's sampled tickets remain small, and the confidence interval on per-agent scores stays wide. Stratification improves what you can say about ticket categories; it does not improve what you can say about individuals.

How wide is the ranking distortion in practice?

In our own analysis and in support teams we have worked with, agent rankings on 2% samples shuffle roughly 30 to 40% of positions month over month. Agents ranked in the top 20% one month land in the middle 60% the next month, not because they got worse, but because the sample was different. Full-coverage scoring cuts month-over-month ranking noise by 70 to 80%.

Are there defensible uses of sampled scores?

Yes, at the team or category level. A 2% sample can reliably tell you the average tone score for chat tickets in July versus June. What it cannot do is tell you which specific agents deserve top or bottom rankings, or which agents to coach on which specific behaviors. Team-level trends are fine; individual-level decisions are not.

Score every conversation, not a sample

Kelanyn grades 100% of your tickets and chats against your own rubric, then turns misses into specific coaching moments per agent.

Request early access