Home/Blog/How to Build a Support QA Rubric That Actually Predicts CSAT
Playbooks

How to Build a Support QA Rubric That Actually Predicts CSAT

Most support QA rubrics measure what is easy to observe rather than what actually causes churn. The result is a scorecard that gives every agent a 92 while your CSAT slips two points a quarter and nobody can say why.

A rubric that predicts CSAT is a different artifact. It is short, weighted against real outcomes, calibrated before it hits the floor, and re-baselined every quarter. This is what that looks like.

What is a support QA rubric supposed to do?

A support QA rubric is a scoring framework for individual customer conversations. It exists to answer two questions: which agents need which specific coaching, and which conversations put the customer relationship at risk. If a rubric is not answering both, it is a compliance artifact, not an operational one.

Most rubrics degrade into compliance artifacts because they are built to be easy to grade rather than useful to act on. A line that reads "Did the agent thank the customer at close?" is easy to grade, hard to disagree about, and almost completely uncorrelated with whether the customer stays. Meanwhile "Did the agent give the correct answer?" is harder to grade, invites disputes, and correlates directly with churn. Rubrics that optimize for grader speed lose the signal that matters.

What dimensions belong in the rubric?

Six to eight, no more. Anything longer than eight loses inter-rater agreement fast, because graders start compensating for one dimension by moving another. The core four show up on almost every high-signal support QA program.

Dimension What it measures Typical weight
Accuracy Was the information given correct against your knowledge base or policy 25%
Resolution Did the customer's underlying problem actually get solved 25%
Process adherence Did the agent follow required steps, escalation, disclosures, ID verification 15%
Tone Empathy, calm, appropriate register for the situation 10%
Efficiency Time to resolve, unnecessary back-and-forth 10%
Personalization Referring to prior context, not repeating what the customer just wrote 5%
Compliance Regulated language, refund policy citations, data handling 5%
Ownership Whether the agent closed the loop or dumped the ticket 5%

Accuracy plus resolution should be at least 50% of the total weight. Those two dimensions correlate most strongly with 30-day reopen rate in every support dataset we have seen. If your rubric weights tone and closing lines above accuracy, you are optimizing for a scorecard, not for outcomes.

How do you weight dimensions correctly?

Start with the weights above, then tune from evidence, not opinion. After the first full quarter of scoring, correlate each dimension's per-agent score against two customer signals: post-conversation CSAT and 30-day reopen rate. Then adjust.

  • Correlation above 0.4. This is a load-bearing dimension. Keep the weight, consider raising it by 5 points.
  • Correlation between 0.2 and 0.4. Meaningful signal. Hold the weight.
  • Correlation below 0.2. The line is measuring something that does not move the customer. Drop it, or downgrade to a binary flag that only surfaces when violated.
  • Negative correlation. Rare but real. Usually a sign the definition is inverted. Rewrite it or cut it.

The mistake to avoid here is weighting the rubric to reflect what the team believes matters rather than what your own data shows matters. Do the correlation. Even a rough one, run once a quarter, beats intuition.

How should each dimension be scored?

Binary flags where the answer is clear, 1 to 5 scales where judgment is required. Never 1 to 100. The precision is fake and it destroys inter-rater agreement.

  • Accuracy. Binary. The information was correct, or it was not.
  • Compliance. Binary. Required disclosure was made, or it was not.
  • Resolution. Three point. Fully resolved, partially resolved, unresolved.
  • Tone, personalization, ownership. Five point. With clear anchor examples for score 1, 3, and 5.
  • Process adherence. Binary per required step, then averaged.

The rubric document should include anchor examples for every non-binary dimension. Anchors are what keep 15 graders scoring the same conversation within a point of each other. Without anchors, "empathy" becomes a projection test.

How do you calibrate a rubric before rolling it out?

Calibration is the step most teams skip and then wonder why agents keep disputing scores. It takes about two weeks and it is not optional.

  1. Assemble a calibration set. Pull 25 to 50 conversations that span your ticket mix: billing, product bugs, cancellations, onboarding, escalations.
  2. Score independently. Three graders score every conversation without discussing. This is your baseline inter-rater agreement.
  3. Measure agreement per dimension. Anywhere below 75% agreement, the definition is fuzzy. Rewrite it.
  4. Discuss and re-score. Bring graders together, walk through disagreements, sharpen definitions and anchors.
  5. Re-score fresh conversations. After definition edits, rescore a new 25-conversation set. Repeat until agreement is above 85% on binary lines and 75% on scored lines.

Only then does the rubric go live. Agents on the floor will not accept scores from a rubric that graders themselves cannot agree on, and they should not have to.

Should you score every conversation or a sample?

Every one, if you can. A 2% random sample is what most teams do because manual grading forces the compromise, but sampling produces two failures. First, the worst conversations, the ones with an incorrect refund promise or a missed escalation, almost never land in a 2% sample. Second, per-agent scores become statistically unstable, so one bad ticket can tank an agent's month.

Automated scoring closes this. When every ticket and chat is graded within hours of closing, the flagged 5% are surfaced daily, and QA specialists review flags rather than sample randomly. Human judgment stays central; the machine just does the reading.

How often should you re-baseline the rubric?

Quarterly for structure, continuously for definitions.

  • Weekly. Track dispute rates per dimension. Anything above 8% dispute rate needs a definition sharpening.
  • Monthly. Refresh anchor examples using real conversations from that month, so the rubric stays grounded in what the queue actually looks like.
  • Quarterly. Re-run the correlation analysis against CSAT and reopen rate. Adjust weights, drop weak dimensions, propose new ones.
  • Annually. Full structural review. Bring in customer research and churn interviews to decide whether the rubric still reflects what the business considers quality service.

If your rubric has not changed in a year, you are either operating an unusually stable product, or you are not looking hard enough.

The mistake to avoid

Do not build a rubric to make grading feel objective. Build it to change agent behavior in ways your customers can feel. That means fewer dimensions than you want, weighted against evidence you can point to, calibrated until graders agree, and re-baselined quarterly. A rubric that produces stable 92 scores across every agent every month is not measuring quality; it is measuring the rubric.

support-qaqa-rubriccsatagent-coaching

Frequently asked questions

How many dimensions should a support QA rubric have?

Six to eight. Fewer than six and you miss real behaviors; more than eight and inter-rater agreement collapses because graders start weighting dimensions inconsistently. The four that show up in almost every high-signal rubric are accuracy, resolution, tone, and process adherence. Add two or three domain-specific dimensions like empathy on cancellation flows or compliance on financial questions.

How do you know if a rubric line item is actually predictive?

Correlate each dimension's score against CSAT and 30-day reopen rate across a rolling 90-day window. Any dimension with correlation below 0.2 is noise. If closing salutation quality shows zero correlation with anything, it does not belong in the rubric, no matter how much the training team likes measuring it.

Should agents see their own QA scores?

Yes, with the conversation attached and the specific message that missed. Aggregate scores without examples produce defensiveness. Conversation-level scores with the actual quote and the rubric definition produce coaching conversations. Agents also need a dispute path, or the rubric becomes a control tool rather than a coaching tool.

How often should the rubric change?

Quarterly re-baseline, with continuous small edits allowed to definitions. Big structural changes, adding or removing dimensions, changing weights by more than 10 points, should batch to once a quarter so agents can adapt. If you change the rubric every month, agent scores become uncomparable across time and coaching loses its grounding.

What is inter-rater agreement and what number is good enough?

Inter-rater agreement is the percentage of times two graders assign the same score to the same conversation on the same rubric line. Above 85% agreement on binary lines and above 75% on scored lines is the working threshold before rolling a rubric live. Below that, the rubric is measuring the grader more than the agent.

Score every conversation, not a sample

Kelanyn grades 100% of your tickets and chats against your own rubric, then turns misses into specific coaching moments per agent.

Request early access