Net Promoter Score vs. CSAT vs. CES: Which Metric Answers What?
Use Net Promoter Score when the question is whether customers would advocate for the broader relationship, CSAT when the question is whether a defined experience met expectations, and Customer Effort Score when the question is whether a defined task felt easy. They are complementary signals, not rival answers to one question. None explains why a score moved, and none can replace observed behavior such as retention, repeat use, referrals, or issue resolution.
Three metrics, three questions
The cleanest distinction is grammatical. NPS asks about recommendation intent. CSAT asks about satisfaction with something named. CES asks about the effort or ease involved in doing something named. If the object and decision cannot be written clearly, choosing a familiar acronym will not rescue the survey.
| Metric | The question it answers | Typical scope | Standardization |
|---|---|---|---|
| NPS | Would these customers recommend the company, product, or service? | A broader relationship, brand, or product | Fixed 0–10 grouping and subtraction formula |
| CSAT | Were these customers satisfied with this specified experience or outcome? | A product, service, event, or interaction | Commonly a five-point top-two-box percentage, but implementations vary |
| CES | How easy or difficult was this specified task or interaction? | A workflow, purchase, onboarding step, or support resolution | Question, scale, polarity, and aggregation can vary |
Net Promoter Score: recommendation intent at the relationship level
The Net Promoter System methodology uses a 0-to-10 likelihood-to-recommend question. Responses of 9 or 10 are Promoters, 7 or 8 are Passives, and 0 through 6 are Detractors.
NPS = percentage of Promoters − percentage of Detractors
Here is an illustrative example, not company data. Among 100 respondents, 45 are Promoters, 35 are Passives, and 20 are Detractors. The calculation is 45% − 20% = 25, so NPS is 25, not 25%. Passives do not enter the subtraction, but they remain in the 100-response denominator used to calculate both percentages. The possible score runs from −100 to 100.
One customer does not “have an NPS.” That customer supplies a rating. NPS is the aggregate produced from a defined respondent population. It is useful as a compact advocacy signal across the company, a product, or a segment, but the score alone does not reveal the experience that caused the rating.
CSAT: satisfaction with a named experience
CSAT usually asks how satisfied a customer was with a specified product, service, interaction, or event. Under the common five-point convention documented by Qualtrics, 4 means satisfied and 5 means very satisfied:
CSAT = (responses rated 4 or 5 ÷ all valid responses) × 100
In a separate illustrative survey, 72 of 100 respondents select 4 or 5 after a support interaction. CSAT is (72 ÷ 100) × 100 = 72%.
Some systems average the raw satisfaction ratings instead. That can be valid, but it is a different statistic. Label a top-two-box percentage as a percentage and a mean as an average; do not compare the two as though the shared CSAT label makes them equivalent.
Customer Effort Score: friction in a named task
CES measures how much effort a customer perceived while resolving an issue, completing a purchase, using a workflow, or getting an answer. Qualtrics places it close to the interaction being assessed. Unlike NPS, CES does not have one universally fixed question-and-calculation package.
One common convention is a mean ease rating:
Mean CES = sum of valid response values ÷ number of valid responses
In a third illustrative survey, 100 ratings sum to 540 on a seven-point scale whose labels run from difficult to easy. Mean CES is 540 ÷ 100 = 5.4 out of 7; higher means easier only because the survey defined the scale that way. Another program may reverse the polarity or report the percentage of favorable responses instead.
The three illustrative results—NPS 25, CSAT 72%, and CES 5.4 out of 7—cannot be ranked against one another. Their numbers have different denominators, scales, and meanings.
The same customer can give three sensible answers
A customer can be willing to recommend a product, satisfied that a support issue was resolved, and frustrated by how many steps resolution required. A high NPS, high CSAT, and poor CES would not contradict one another in that case. Each score describes a different layer of the experience.
That separation is what makes the metrics operationally useful:
| Decision | Best starting metric | What the score still cannot tell you |
|---|---|---|
| Is broad customer advocacy strengthening or weakening? | NPS | Which product, service, price, or brand factor caused the change |
| Did a support, onboarding, or purchase experience meet expectations? | CSAT | Whether the process was easy or whether the whole relationship is healthy |
| Did customers encounter friction in a specific workflow? | CES | Whether the outcome itself was satisfactory or the customer will recommend the company |
| Why are customers churning or expanding? | None on its own | Causation; connect feedback with behavior and account context |
Use NPS when the unit of management is a relationship-level signal. Use CSAT when an owner can act on a defined experience. Use CES when a team can remove steps, handoffs, repetition, waiting, or ambiguity from a defined task. If the intended action is unclear, collect qualitative or behavioral evidence before adding another score.
Write the measurement contract before sending the survey
The metric name is only the first line of the specification. A defensible trend needs a stable measurement contract:
- Object: the company, product, support case, onboarding flow, purchase, or another precisely named experience.
- Eligible population: who can be asked, which lifecycle state they must be in, and who is excluded.
- Trigger: the event or relationship stage that makes the question relevant.
- Timing: how soon after that trigger the request is sent and how repeat requests are suppressed.
- Question and scale: exact wording, endpoint labels, scale direction, and available nonresponse choices.
- Calculation: valid-response rules, top-box thresholds or averaging method, rounding, and treatment of missing data.
- Reporting slice: time window, product, plan, segment, channel, geography, or journey stage.
- Response owner: who reviews the reason behind the score and what action they can take.
Changing any of those fields can move the number without changing the customer experience. A support-triggered NPS and a relationship survey sent to all active customers may use identical arithmetic while observing different populations at different moments. Treat a method change as a break in the series, not as an improvement or decline.
Show response count and score distribution beside the headline. Two samples can both produce NPS 20: one could contain 30% Promoters, 60% Passives, and 10% Detractors; another could contain 60% Promoters and 40% Detractors. The net is identical, but the second population is far more polarized. The distribution preserves an operating clue that the single number discards.
The same caution applies to response bias. Qualtrics notes that people with very positive or very negative experiences may be more likely to answer a CSAT request. Report the result as a score among respondents under the documented method, not as a fact about every customer.
A “good” score requires a valid comparison
There is no context-free number that makes NPS, CSAT, or CES good. A benchmark becomes useful only when the comparison is close enough to answer a real question.
Bain’s own NPS benchmark guidance frames comparison against competitors and distinguishes overall NPS from journey- or episode-level views. Qualtrics describes CSAT benchmarking as inexact because businesses and products differ. CES is harder again when programs use different wording, ranges, directions, and calculations.
Before accepting an external benchmark, check the industry, market, customer population, object being rated, survey mode, timing, wording, and calculation. If those are unavailable, the number is context, not a target. A consistently measured internal trend is often the more defensible baseline, provided the response count and customer mix are visible.
NPS is a signal, not proof of growth
NPS is often called a loyalty metric, but its direct observation is stated recommendation intent. Actual recommendations, renewals, expansions, contraction, and churn are separate behaviors. Do not turn a survey response into a claim about revenue without testing the link in your own data.
A peer-reviewed empirical investigation in the Journal of the Academy of Marketing Science found mixed evidence across studies of NPS and sales growth. The reviewed direct studies did not confirm the stronger claim that NPS is a superior growth predictor to other customer-mindset metrics, and differences in research design limited broad conclusions.
The practical response is not to discard NPS. It is to validate it. For each measured cohort or segment, examine whether the score and its reasons move with later retention, expansion, referrals, product use, or other outcomes the business actually cares about. Keep correlation and causation separate. A score can be a useful early signal without being the mechanism that created the outcome.
Use all three without creating survey noise
A compact feedback system can give each metric one job:
- Run NPS as a relationship or product-level pulse for a clearly eligible population at consistent lifecycle moments.
- Trigger CSAT after a defined experience whose owner can improve the outcome.
- Trigger CES after a task where friction is a design or service decision.
- Store the raw rating, open-text reason, trigger, customer context, and operational outcome—not only the aggregate score.
- Apply suppression rules so one customer is not repeatedly asked to rate overlapping experiences.
- Route low scores and recurring themes to an owner with authority to investigate and respond.
Do not average NPS, CSAT, and CES into a synthetic customer-health score merely because all three are survey metrics. If a health model is needed, retain each input separately, state its weighting and time window, and validate the combined model against an observed outcome. Otherwise the blend hides the precise question each metric was selected to answer.
Let disagreement point to the next investigation
When the metrics move in different directions, use the pattern as a hypothesis—not a verdict.
High NPS with low CSAT suggests that broad goodwill may be masking a weak recent experience. Inspect the named touchpoint and its open-text reasons.
High CSAT with weak NPS suggests that one interaction met expectations while the broader value proposition, product fit, price, or relationship may still be weak. Look beyond the service team.
High CSAT with poor CES suggests that customers reached a satisfactory outcome through too much work. Examine transfers, repeated information, wait time, and avoidable steps.
Strong CES with low CSAT suggests that the process was easy but the result failed expectations. Removing more steps will not fix the underlying outcome.
These are diagnostic starting points. Confirm them with comments, journey events, support records, product behavior, and commercial outcomes before deciding what caused the pattern.
Sources
- Bain & Company, Net Promoter System, “Measuring Your Net Promoter Score”
- Qualtrics, “What Is CSAT and How Do You Measure It?”
- Qualtrics, “What Is Customer Effort Score (CES) and How Do I Measure It?”
- IBM, “What Is a Customer Effort Score?”
- Bain & Company, Net Promoter System, “Bain Certified Net Promoter Score Benchmarks”
- Journal of the Academy of Marketing Science, “The Use of Net Promoter Score (NPS) to Predict Sales Growth: Insights from an Empirical Investigation”
Continue the evidence path
Related reading
Related
Customer Satisfaction Metrics Explained: CSAT, NPS, CES, and What Each Measures
Place NPS beside CSAT and CES to choose the measure that actually matches the customer question.
Next step
Customer Feedback Surveys That Produce Decisions, Not a Backlog of Opinions
Build a sampling and follow-up process that turns the score into evidence instead of a standalone benchmark.