Customer Satisfaction Metrics: CSAT vs. NPS vs. CES
A customer can be pleased with a support outcome, frustrated by the work it took to get there, and still willing to recommend the vendor. If a dashboard treats those answers as versions of the same thing, the apparent contradiction becomes a reporting problem. In reality, the survey asked three different questions.

CSAT measures satisfaction with a named experience. CES measures the perceived ease or effort of a task. NPS records stated willingness to recommend a company, product, or service. None is the universal customer metric. The useful choice is the score that can change a decision your team actually owns.
That makes metric selection less about finding the most famous acronym and more about matching the question to the intervention. A support leader deciding whether case resolutions meet expectations needs a different signal from a product team trying to remove setup friction or an executive team tracking recommendation intent across a customer population. The arithmetic matters, but the decision comes first.
Match the metric to the decision it can change
The three metrics belong under the broad heading of customer experience, but they are not interchangeable. Qualtrics distinguishes CSAT as satisfaction with a product or service, NPS as a wider recommendation-intent question, and CES as the difficulty of completing a task. The distinction becomes practical when each score is attached to a situation narrow enough to investigate.
| Metric | The question it answers | Best-fit decision | Typical scope | What it does not establish |
|---|---|---|---|---|
| CSAT | Did this named experience meet the customer’s expectations? | Whether to improve an outcome such as support resolution, onboarding, delivery, or a product interaction | Usually a recent transaction or defined experience | Long-term loyalty, renewal, or the amount of work required |
| CES | How easy or difficult was this task? | Where to remove steps, waiting, transfers, repetition, or confusion | A specific workflow or service interaction | Whether the eventual outcome was satisfactory or the relationship is strong |
| NPS | How likely is this respondent to recommend the company, product, or service? | How recommendation intent changes for a defined population or segment | A broader relationship, brand, product, or service | An observed referral, renewal, expansion, or causal explanation |
Use CSAT when the team can name the experience being judged. “How satisfied are you with the resolution of this support request?” is much more useful than “How satisfied are you?” The narrower version tells the respondent what to consider and tells the team where to look if the score changes. CSAT can also follow an onboarding milestone or delivery, but the customer must have had enough time to judge the result. Qualtrics recommends surveying soon after a discrete interaction while allowing more time for products or services that need days or weeks of use before they can be assessed.
Use CES when the work itself is the problem. “It was easy to invite a teammate” can expose a permissions issue, unclear interface copy, or an unnecessary approval. “The company made things easy” blends too many encounters to point anywhere. Qualtrics defines CES around the effort needed to resolve an issue, fulfill a request, buy or return a product, or answer a question, and recommends asking immediately after the interaction being assessed. A short reason question can then tell the team which part felt difficult.
Use NPS when the decision concerns a wider relationship or market signal and the population is explicit. The standard question asks how likely someone is to recommend the organization, product, or service. That is a statement of intent. It is not evidence that the person made a referral, and it cannot substitute for renewal, retention, adoption, or expansion data when those behaviors are the real outcome.
This boundary matters because NPS is often asked to carry a claim bigger than its question. A peer-reviewed study in the Journal of the Academy of Marketing Science found that prior research on NPS and sales growth was mixed and that none of four earlier academic studies confirmed NPS as a superior predictor. Its own results found predictive value under specific conditions, using five years of data from seven large US sportswear brands—not a universal result for every industry or B2B program. The study’s authors argue that population, use case, horizon, and research design affect what NPS can support. Treat the score as recommendation intent unless separate evidence connects it to a business outcome in your setting.
The practical answer is often to use more than one metric at different journey points. That does not mean asking all three after every event. A subscription business might use CES after workspace setup, CSAT after a support resolution, and NPS at a stable relationship milestone. Each score then retains a clear object and cadence. Surveying every customer after every touchpoint creates fatigue and three numbers that are tempting to average into a meaningless “happiness” score.
Do not combine CSAT, CES, and NPS into one composite: they use different constructs, scales, and time horizons, so the blend hides which experience needs attention.
Calculate each score without changing its meaning
The formula is part of the metric. A label on a dashboard cannot rescue a calculation that uses different groupings, denominators, or scale direction. Before comparing teams or periods, make sure they are producing the same statistic.
The examples below are illustrations, not measured company results. Each starts with 100 valid responses so the arithmetic is visible.
| Metric | Common calculation | Illustrative inputs | Illustrative result |
|---|---|---|---|
| CSAT | Satisfied responses divided by valid responses, multiplied by 100 | 72 respondents select 4 or 5 on a five-point satisfaction scale | 72% CSAT |
| NPS | Percentage of Promoters minus percentage of Detractors | 45 Promoters, 35 Passives, 20 Detractors | NPS of 25 |
| CES | Sum of valid ratings divided by valid responses | Ratings total 540 on a seven-point easy-to-difficult scale | 5.4 out of 7 |
CSAT: state whether you report a percentage or an average
A common CSAT survey uses a five-point scale ranging from very dissatisfied to very satisfied. The common “top-two-box” calculation counts ratings of 4 and 5 as satisfied: satisfied responses / all valid rating responses × 100. Qualtrics documents both the five-point question and the top-two-box formula.
Suppose 72 of 100 valid respondents choose 4 or 5. CSAT is 72 / 100 × 100 = 72%. If the same set of ratings has a mean of 4.1, that average is also a legitimate summary of the responses, but it is not the same statistic. Calling both figures “CSAT” without displaying the calculation makes two dashboards look comparable when they are not.
The denominator should contain valid answers to the rating question. A skipped response is not a zero, and a genuine “not applicable” response is not dissatisfaction. Keep those cases out of the score, then show their counts or rates beside it. A spike in non-applicable answers may itself reveal that the survey trigger reached people who could not judge the experience.
Also keep the response distribution. Two teams can both report 70% top-two-box CSAT while one has mostly 4s and few 1s, and the other has a polarized mix of 5s and 1s. The headline percentage conceals that difference. Distribution, verbatim reasons, and operational outcomes supply the texture the aggregate removes.
NPS: Passives stay in the denominator
The Net Promoter System’s published method uses an 11-point likelihood-to-recommend scale from 0 to 10. Scores of 9 or 10 are Promoters, 7 or 8 are Passives, and 0 through 6 are Detractors. NPS is the percentage of Promoters minus the percentage of Detractors.
With 45 Promoters, 35 Passives, and 20 Detractors among 100 valid responses, the calculation is 45% − 20% = 25. Passives are not subtracted, but they still belong in the denominator used to calculate the two percentages. Removing them would inflate both groups and change the score.
NPS is reported as a score, not a percentage. Its theoretical range is −100, when every valid respondent is a Detractor, to +100, when every valid respondent is a Promoter. That wide range does not make a one-point movement meaningful by itself. A change can reflect sampling variation, a new survey channel, a different customer mix, or a real shift in sentiment. The score needs its valid response count, population, collection window, and method beside it.
Be equally precise about the object in the question. Recommendation of the company is not necessarily recommendation of one product, and recommendation of a product is not the same as satisfaction with the last service case. A business that changes the object but leaves one continuous chart has changed what the line means.
CES: publish the scale direction with the result
CES is less standardized. Some surveys ask how easy a task was; others ask how much effort it required. Some use agreement statements, numeric ranges, or labeled categories. IBM describes CES as an average—the sum of ratings divided by the number of responses—and notes that organizations choose different ranges. That flexibility makes a bare number unusually risky.
If 100 ratings sum to 540 on a seven-point scale where 1 means very difficult and 7 means very easy, the mean CES is 5.4 out of 7, and higher is better. Reverse the wording to “How much effort did this require?” and a higher number may now be worse. A report that says “CES rose from 5.1 to 5.4” is unintelligible without the exact question, endpoints, and polarity.
The task must be narrow enough to fix. A low ease score for “using our platform” could refer to login, setup, permissions, navigation, performance, or support. A low ease score for “adding a new billing administrator” points to a tractable workflow. Pair it with an optional open-text reason and observed measures such as completion, abandonment, repeated contact, transfers, or elapsed time. CES tells you that customers perceived friction; those other signals help locate it.
Build a survey that remains comparable
A score becomes a trend only when the rules that produce it remain stable. Change the respondent population, question wording, trigger, channel, or denominator, and the line may move even if the customer experience has not.
The American Association for Public Opinion Research advises keeping wording, framing, and survey methodology as similar as possible when measuring change over time. It also warns that a change in survey mode can mimic a shift because people may answer differently in a phone conversation than in a private web survey. Customer programs face the same problem when they move a question from email to an in-app prompt or from users after an event to account sponsors on a schedule.
Fix the question, trigger, population, and calculation before launch
A workable measurement plan can fit on one page. It needs enough detail for another analyst to reproduce the score and enough business context for a team to know what happens when it moves.
- Name the decision. State what a product, service, success, or leadership team might change after seeing the result. If no plausible action exists, the survey is probably collecting decoration.
- Choose the construct. Ask whether the decision needs satisfaction, perceived effort, or recommendation intent. This choice determines CSAT, CES, or NPS—not the metric already present in the reporting tool.
- Write the complete instrument. Preserve the exact question, object being rated, response labels, numeric values, scale endpoints, and direction. Pretest unfamiliar wording with people who resemble the intended respondents; AAPOR recommends cognitive or qualitative pretesting to learn how respondents interpret questions.
- Define eligibility and timing. Specify who can receive the survey, which event triggers it, how soon it appears, how repeated invitations are limited, and which reporting period receives late answers.
- Lock the calculation. Define a valid response, treatment of skips and “not applicable,” grouping rules, denominator, aggregation, rounding, and the minimum information shown with the score.
- Plan the breakdowns. Decide which customer tiers, regions, products, roles, or journey types matter before results arrive. Show the response count for every segment; a large total sample does not make a small subgroup precise.
- Connect the next observation. Decide which comments, interviews, tickets, usage events, resolution measures, renewals, or referrals can test the story suggested by the survey.
- Set an action rule. Define what level or pattern prompts investigation, who examines it, and what comparison will show whether a change helped.
Version this plan with the questionnaire. When a flawed question must change, mark a break in the series or run old and new versions in parallel long enough to see whether the instrument itself changes responses. AAPOR describes a split-ballot experiment for this problem: comparable respondents receive different versions while other factors remain as similar as possible. Without that bridge, a cleaner new question may be worth adopting, but it should not masquerade as an uninterrupted trend.
In B2B, decide whether a person or an account is the customer
The economic buyer, daily user, administrator, and executive sponsor can all answer honestly and still disagree. A user may find setup difficult while a buyer remains satisfied with the commercial outcome. An account may have one sponsor response and twenty user responses. If every person receives equal weight, the larger account can dominate the aggregate; if every account receives equal weight, a single answer may stand in for many people.
There is no universally correct unit. The choice must follow the decision. A product team diagnosing a workflow may reasonably analyze user-level CES. A customer-success leader evaluating account health may want role-specific views and an account-level rule. Keep both levels when they answer different questions, but do not silently average them into one population.
Response coverage matters as much as response count. AAPOR notes that very low response can miss entire types of respondents and bias survey estimates. In a B2B program, compare invitees and respondents by tier, tenure, role, geography, product, and recent support activity where those attributes are available and appropriate. If unhappy churned customers can no longer receive the in-product survey, a rising score among active users may coexist with a worsening overall outcome.
Report the score as a result among respondents under a documented method. That is a more defensible statement than treating it as the literal opinion of every customer.
Read combinations as clues, not conclusions
There is no context-free “good” CSAT, NPS, or CES. A benchmark is useful only when its instrument and population resemble yours. Match the metric definition, question wording, scale, direction, journey stage, trigger, channel, market, and reporting window before using an external figure as a target. Qualtrics calls CSAT benchmarking inexact because products and businesses differ; CES varies even more when programs choose different ranges and polarities.
The American Customer Satisfaction Index is another important boundary. ACSI is not a brand name for a one-question top-two-box CSAT percentage. ACSI calculates its satisfaction index as a weighted average of three survey questions within a multi-equation econometric model. An ACSI benchmark and your internal CSAT may both concern satisfaction, but their numbers are not directly interchangeable.
Internal comparisons are usually stronger when the instrument stays fixed and the team can connect changes to events. Even then, the aggregate is a clue. Split it by the segment and journey differences that could change the decision, and keep the number of responses visible. A five-point increase concentrated among newly onboarded administrators tells a different story from the same increase spread across every role.
When multiple metrics coexist, their disagreement is useful:
| Pattern | What it may mean | What to inspect before acting |
|---|---|---|
| High CSAT, weak CES | Customers reached a satisfactory outcome through too much work | Repeated information, transfers, waiting, avoidable steps, and time to completion |
| Strong CES, low CSAT | The process was easy, but the result did not meet expectations | Resolution quality, missing capability, product fit, or incomplete outcome |
| Strong NPS, weak recent CSAT | Broader goodwill may coexist with a poor touchpoint | The named incident, affected segment, prior relationship, and open-text reasons |
| Good transactional scores, weak NPS | Individual interactions may work while broader value or trust is weak | Adoption, renewal, price-value perceptions, relationship history, and respondent mix |
These interpretations are hypotheses. A smooth workflow can produce good CES even when the product fails to deliver the desired result. A customer may recommend a tool but choose not to renew because budgets changed. A dissatisfied respondent may remain because switching costs are high. Survey scores cannot tell those causal stories alone.
The dashboard should therefore show what someone needs to challenge the obvious story: exact question and version, scale direction, eligible population, valid responses, response rate where it can be calculated, distribution, collection window, channel, key segment counts, and relevant behavioral outcomes. Keep comments available for diagnosis, but do not let a memorable quotation outweigh the response pattern or observed behavior.
Trend interpretation also requires restraint. For a probability sample, sampling uncertainty shrinks as the sample grows, but subgroup results have uncertainty based on their own smaller counts. AAPOR’s margin-of-sampling-error guide stresses that the familiar margin of error applies to probability samples, not ordinary opt-in surveys. Most transactional customer surveys are closer to a census invitation with nonresponse or an opt-in mechanism than to a simple random sample. A large response total does not erase selection, wording, coverage, or mode effects.
Turn the score into a testable operating loop
The cleanest customer measurement program starts small. One question is tied to one journey moment and one decision. The team holds the instrument steady long enough to understand its baseline, collects a reason without forcing a long questionnaire, and connects the response to the relevant event. Only then does it look for a recurring driver it can change.
For example, weak CES after support resolution may coincide with repeat contacts and transfers. The response does not prove transfers caused the effort rating, but it identifies a testable path. The service team can examine cases, change routing for one issue type, and compare the same CES question plus transfer and repeat-contact rates before and after. If CSAT improves while CES does not, the outcome may be better even though the journey is still hard. The disagreement prevents a premature victory lap.
The same discipline protects NPS. Track recommendation intent for a consistent population, but measure referrals, renewals, adoption, and expansion directly when those are the decisions at stake. If NPS rises while renewal falls, do not choose the more flattering number. Split the data by customer role, tenure, tier, and product experience, then investigate the population in which the relationship breaks.
Closing the loop does not mean contacting every low scorer with a scripted apology. It means using a response at the right level. An urgent unresolved case may justify individual follow-up; a recurring onboarding complaint calls for a product or process change; a one-point aggregate movement with a changed respondent mix may call for no intervention at all. The action should match the evidence.
When the team makes a change, compare like with like: same question, scale, trigger, population, channel, and calculation. Show behavioral measures beside the survey result. A score becomes useful when it narrows uncertainty enough to choose the next investigation or intervention—not when it merely rises.
Choose the question before building the dashboard
Start with one sentence: “We need to know ___ so that ___ can decide ___.” If the blank is whether a defined outcome satisfied customers, use CSAT. If it is whether a defined task felt easy, use CES. If it is how recommendation intent is moving across a defined population, use NPS.
Then write the survey and calculation beside the decision before anyone configures a chart. The dashboard comes last. That order makes it harder for a convenient number to become a substitute for the question the business actually needed answered.
Frequently asked questions
How many responses do CSAT, NPS, or CES need?
There is no universal minimum; the count must support the precision and segment decisions you intend to make. In a simple random sample at 95% confidence, the maximum margin of sampling error is about ±9.8 percentage points at 100 responses, ±4.9 at 400, and ±3.1 at 1,000. Those calculations do not cover nonresponse, wording, coverage, or selection bias, and AAPOR says the conventional margin of sampling error does not apply to opt-in, nonprobability surveys. Plan each reported subgroup using its own response count, and audit who was invited but did not answer.
How should CSAT percentages be combined across teams or periods?
Pool the underlying counts only when the question, scale, eligibility, trigger, and calculation are the same. If Team A has 9 satisfied responses out of 10 and Team B has 45 out of 90, the combined CSAT is (9 + 45) / (10 + 90) = 54%, not the unweighted average of 90% and 50%, which would be 70%. Keep the results separate when the instruments or journey moments differ; arithmetic cannot make unlike measures comparable.
How should skipped and “not applicable” responses be handled?
Exclude a skipped or genuine “not applicable” answer from the valid-response denominator, and report its count or rate beside the score. Do not code it as zero. AAPOR recommends storing nonresponse separately from substantive choices such as “none of the above,” “don’t know,” or “prefer not to answer,” because they mean different things in analysis. Predeclare whether a partial survey counts when the primary rating is complete, and keep an answer such as “the issue was not resolved” inside the distribution rather than hiding it as missing data.
Is ACSI the same as CSAT?
ACSI and CSAT are different measures. A common internal CSAT is the percentage of respondents selecting the top two options on a five-point satisfaction question. ACSI uses a proprietary weighting of three satisfaction questions as part of an econometric model. Both concern satisfaction, but a score of 80 from one method is not automatically equivalent to 80 from the other; compare only after confirming the instrument and calculation.