Home/Blog/The Hidden Cost of Grading Only 2% of Support Tickets
Strategy

The Hidden Cost of Grading Only 2% of Support Tickets

Every support leader we talk to admits, quietly, that their QA program grades a sample and hopes the sample is representative. Two percent is the number that comes up most often. The math on that number gets worse the harder you look.

What is a 2% QA sampling rate actually covering?

A team handling 10,000 monthly conversations with a 2% QA rate grades 200 tickets. Spread across 20 agents, that is 10 tickets per agent per month. Split that further across four rubric dimensions and you have 40 scored data points per agent per month. That is not a performance signal. That is a performance rumor.

The coverage math is worse when you look at what the sample misses.

  • Serious errors. Roughly 3 to 5% of tickets contain a real error: wrong policy, wrong refund, missed escalation, incorrect product answer. A 2% random sample catches, on average, 2% of those errors. About 98 errors per 100 go unread, unlogged, and unaddressed.
  • New failure modes. When a product bug ships or a policy changes, agents start giving the same wrong answer across dozens of conversations before anyone notices. A 2% sample cannot detect a pattern; the sample size is too small to distinguish a wave from a blip.
  • Cross-channel bleed. Sampling programs typically over-sample email and under-sample chat, because chat volume is higher and grading harder. Your worst conversations are frequently in the channel you sample least.

The 2% rate is not a coverage rate. It is a plausible-deniability rate.

Why does the sample miss the worst conversations?

Because bad conversations are rare, and rare events do not populate small random samples. This is not a philosophical point; it is arithmetic.

Assume 4% of your monthly tickets contain a serious accuracy error. A 2% random sample of 10,000 tickets is 200 tickets. The expected number of serious errors caught is 200 x 0.04, which is 8. Out of the 400 errors your customers experienced that month, you saw 8. That is a 2% capture rate on the metric you most care about catching.

Now add a real-world wrinkle: the graders who do land on a serious error often flag it, escalate it, and pull it out of the sample for review. That is the correct behavior. But it means the observable sample skews slightly toward routine tickets, which further weakens your read on the tail. Your dashboard says quality is 92%; the customers who received the bottom 4% say otherwise.

How does sampling distort agent rankings?

Six to ten graded tickets a month is not enough to distinguish agent performance. With four rubric dimensions and modest natural variance, the confidence interval on an agent's monthly score is roughly plus or minus 8 points. That means an agent's "true" performance level of 88 can show up as anything from 80 to 96 depending on which specific tickets landed in the grader's queue.

The visible consequences.

  • Monthly rankings flip. The agent who was #3 last month is #14 this month, not because they got better or worse, but because their sample was different. Managers respond by chasing narratives that do not exist.
  • Coaching lands wrong. An agent scored low on tone in a single sampled ticket where they were curt with a hostile customer gets a coaching session on tone, when the same agent's other 199 conversations that month showed no tone issue.
  • Bonus programs invite gaming. When bonuses are tied to sampled scores and the sample is small, agents rationally optimize for how a random ticket might look, not for how their overall conversations go. This distorts behavior more than most support leaders realize.

Full-coverage QA fixes this by making the sample size equal to the population. When every ticket is scored, an agent's monthly number is their actual performance, not a lottery ticket.

What is the lag cost of manual QA sampling?

Manual grading runs 5 to 8 days behind the customer moment. A ticket closed on Monday gets sampled Friday, graded the following Wednesday, discussed with the agent in a 1:1 the week after. By the time the agent hears about a mistake, they have already made the same mistake three more times, and the original customer is either forgiven you or churned.

Compressing the lag is the highest-leverage change most QA programs can make. Coaching that lands within 24 to 48 hours of the conversation, while the context is still in the agent's head, changes behavior. Coaching that lands ten days later reads as archaeology.

What does full-coverage QA actually change?

We track four numbers when teams move from a 2% sample to scoring every conversation.

  • Reopen rate. Typically drops 20 to 30% inside one quarter. Not because agents suddenly got better, but because coaching started landing on actual misses instead of sampled ones.
  • CSAT stability. The monthly CSAT number stops jumping four points in either direction and settles into a narrower band, because rare-but-severe errors are now caught before they land in customer surveys.
  • QA specialist mix. Graders shift from reading random tickets to reviewing flagged tickets. The same headcount produces more coaching output because they stop reading the average and start reading the exceptional.
  • Agent trust. Counterintuitively, agents become more accepting of QA, not less, once every conversation is scored. Full coverage is fairer than sampling. One bad ticket in a 200-ticket month is one bad ticket, not a monthly identity.

Why do most teams stay on sampling anyway?

Three reasons, and only one of them holds up.

  • "We tried it and the scoring wasn't accurate enough." Fair concern with early tools. Modern support QA systems calibrate against your own graders and mark low-confidence scores for human review rather than reporting them as fact. The question is not whether the machine is perfect; it is whether the machine plus a human on the flagged 5% beats a human on a random 2%. It does.
  • "Our rubric is too nuanced to automate." Sometimes true. More often, the rubric has been elaborated to justify manual grading rather than because the nuance is measurable. If your rubric line requires the grader to intuit customer emotional state from three-word replies, the rubric line is measuring the grader, not the agent.
  • "Budget." The real reason. Automated QA costs money that a headcount-based QA program does not visibly cost, because the sampling program's real cost, missed churn and coaching lag, does not show up on the QA cost center's line.

The last one is the one to interrogate. What does 3 CSAT points cost you? What does a 25% higher reopen rate cost? Those numbers, roughly estimated, dwarf the software cost of full-coverage QA in every account we have seen.

The mistake to avoid

Do not confuse sampling coverage with quality visibility. A 2% sample tells you the average conversation was fine. It cannot tell you which of your 200 agents just told 40 customers the wrong refund window, because that pattern lives in the 98% you did not read. Full-coverage scoring is not a fancier version of sampling. It is a different operating model, one where humans review flags rather than draw straws.

qa-samplingsupport-qacx-metricsagent-fairness

Frequently asked questions

Why is 2% the typical QA sampling rate?

Because manual grading takes 10 to 15 minutes per ticket, and a QA specialist can review about 30 to 40 tickets a day. On a team handling 50,000 monthly conversations, one grader covers roughly 800 tickets, or 1.6%. Add one or two graders and you get to 2 to 4% coverage. Above that, you are hiring headcount that produces coaching capacity, not more sampling.

Isn't a random sample statistically representative?

It is representative of the average conversation. It is not representative of the worst. Bad conversations, incorrect refund promises, missed escalations, are the tail of the distribution, and small samples miss the tail by design. A 2% sample of 10,000 tickets has roughly a 15% chance of catching any given serious error, which means most serious errors go unsurfaced.

Can we just weight the sample toward risky ticket types?

Stratified sampling helps but does not solve the problem. It requires knowing in advance which conversations are risky, which is exactly what you would use QA to figure out. Stratifying by ticket type, cancellation, refund, billing, catches obvious risks and misses the emerging ones, like a new bug that agents are answering incorrectly across the whole queue.

How do you know your agent rankings are wrong?

Two symptoms. First, month-over-month rankings flip agents who should be stable, because 6 graded tickets is not enough to distinguish an 88 agent from a 92 agent. Second, when you spot-check unsampled tickets from top-ranked agents, you find quality problems the sample missed. If either is true, your rankings are largely measuring sampling luck.

What does closing the coverage gap actually change?

Three measurable outcomes. Reopen rate typically drops 20 to 30% within a quarter because coaching lands on the actual misses. CSAT stabilizes because you catch policy errors before they escalate. And agent trust in QA improves, because one bad ticket in a 200-ticket month reads as one bad ticket, not a career-defining sample.

Score every conversation, not a sample

Kelanyn grades 100% of your tickets and chats against your own rubric, then turns misses into specific coaching moments per agent.

Request early access