QA scores and customer outcomes — correlation is not a coaching plan
QM data is rich in relationships that look causal: low scores and high complaints, long handle time and weak resolution, new agents and low empathy. The temptation is to turn each correlation into a coaching plan. The disciplined analyst pauses first: is this causation, reverse causation, or a third variable driving both?
The recurring confounders
Agent tenure is the most persistent. New agents have lower scores and longer AHT and lower FCR and lower CSAT — so any correlation among those metrics partly picks up the tenure effect. Contact mix is next: an agent carrying a hard-mix workload looks worse on raw scores than one with an easy mix, and the apparent gap often doesn’t survive stratification. Customer segment, time-of-day patterns and the manager effect all follow the same logic.
And one confounder lives inside the QM programme itself: sample bias. If the scoring sample over-represents complaints, then ‘low scores correlate with complaints’ is true by construction and means nothing. Every relationship built on a skewed sample is skewed with it.
The claims to scrutinise
‘Low scores cause complaints’ — partially true, usually inflated. One operation found a headline correlation of r = −0.58 between agent scores and 7-day complaint rates; stratifying by tenure removed about 40% of it, contact mix another 20%. The remaining relationship was real but much smaller than the headline — and the original ‘coach the low-scorers’ plan would have spent most of its effort on a confound.
‘Coaching improves scores’ deserves the same scrutiny. The agents coached most are usually the agents already struggling — a selection effect — and struggling agents partly improve on their own through regression to the mean. Without a comparison group, the coaching gets credit for the recovery that was coming anyway. The same applies to ‘this team is excellent because the TL is great’: possibly true, possibly the easier queue, the more experienced agents at reallocation, or a generous rater.
The before-claiming-causation checklist
Five questions before any causal claim. Temporal sequence — does X precede Y, or could the arrow run the other way? Plausible confounders — what third variable could explain both? Control — can you stratify or regress to isolate the relationship? Mechanism — does the causal story make operational sense? Intervention — has anyone changed X and watched Y move?
In QM, the answer to the fifth question is unusually available: the coaching experiment. A small-scale, controlled intervention tests the causal claim before the operation invests at scale. If you believe coaching cadence drives new-agent empathy scores, lift the cadence to weekly for a sample of TLs and compare the next quarter — that settles what no amount of correlation can.
Handling a correlation responsibly
The disciplined sequence when an analyst finds a relationship: name the correlation with specifics and effect size; list the plausible confounders out loud; control where the data allows and report how much of the relationship survives; state the limitations honestly; and propose the intervention or controlled comparison that would settle the question.
That last step is what separates analysis from opinion. ‘We see X; the causal reading depends on Y; here is what we would need to know’ beats false confidence every time — and it converts the finding into a testable next step rather than an immediate, possibly wasted, coaching investment.
Why this discipline pays
The cost of skipping it is concrete: coaching effort routed to relationships that don’t survive scrutiny, agents coached for patterns that were actually mix or rater effects, and a QM function whose recommendations stop being trusted because the last three didn’t hold up. Confounded analysis isn’t neutral — it actively misallocates the operation’s scarcest development resource.
Done well, the discipline runs the other way. The function that names its confounders, controls for them, and tests its causal claims with small experiments becomes the operation’s trusted source of causal insight — the place leadership goes to ask not just what is happening but what would change it.
The closing principle
A correlation is a question, not an answer. Name the confounders, control for what you can, and propose the small experiment that would settle it — the smaller, sharper intervention that survives scrutiny beats the big confident one that doesn’t.
See also
- Small samples, big noise when a QA difference is real
- coaching-from-qa-results
- qa-scores-and-workforce-planning