In May 2021, Gartner predicted that more than 75% of organizations would abandon NPS as a measure of customer service success by 2025. It's 2026. NPS is still everywhere; Bain has estimated that around two-thirds of the Fortune 1000 use it. The obituary was wrong, and I think it was wrong for an interesting reason: the critics correctly diagnosed the disease and then blamed the wrong organ.\n\nThe disease is real. A relationship NPS score, on its own, tells you almost nothing actionable. It moves with sample bias, survey timing, even the weather of your last release. Executives stare at a two-point dip and demand answers nobody has. All of that criticism lands.\n\nBut the score was never supposed to carry the program. The score is a trigger. Its job is to sort respondents into rough emotional buckets in five seconds so that the next question can do the real work. And the next question is where nearly every NPS program I've looked at falls apart, because the next question is some variant of: \"What's the primary reason for your score?\"\n\n## The most expensive text box in your company\n\nThink about what that open text box is asking a customer to do. They just gave you a 6. They know why, but explaining it properly means reconstructing the incident, naming the feature, describing what they expected. That's a paragraph of unpaid labor into a box that has never once responded to anything they typed. So they write \"support was slow\" or \"too expensive\" or nothing at all, and your analyst is left reading tea leaves at scale.\n\nThen someone tags those fragments into themes, and the themes are useless in the specific way that should worry you: \"pricing\" was the top driver last quarter, it's the top driver this quarter, and nobody can say what about pricing, for whom, or what changed. The program produces a number executives argue about and a word cloud nobody acts on. Gartner wasn't wrong that this deserves to die. It just isn't the score's fault.\n\n## What a real follow-up looks like\n\nA follow-up worth the name reacts to the answer it just got. Customer says \"support was slow.\" The next question should be: slow how? Slow to first response, or slow to actually fix it? Which ticket? What were you trying to do while you waited? Two or three exchanges like that and \"support was slow\" becomes \"P1 tickets sit in a queue for a day because our EU customers file them at night.\" One of those is a shrug; the other is a staffing decision.\n\nHumans have always known this. Nobody running a customer interview hears \"it was fine\" and moves on. The only reason NPS surveys don't probe is that until recently, probing required a person, and you can't put a person behind 10,000 survey sends. Now you can put an AI interviewer that asks context-aware follow-ups behind every single response, and the economics of the whole program change. The score sorts; the conversation explains.\n\n## Running it without annoying people\n\nThe fear I hear most: \"customers will hate being interrogated after a one-click survey.\" Fair, if you do it badly. A few rules keep it honest. Probe detractors and passives harder than promoters; a 9 needs one question (\"what's the one thing we shouldn't break?\"), a 4 deserves a real conversation. Cap the exchange at two or three follow-ups unless the customer is clearly engaged. And let the interviewer read tone: someone typing one-word answers wants out, and pushing them is how you turn a 4 into a 2. We expose this as an explicit dial in conversation style and probing depth, because the right intensity for a churn-risk enterprise account is wrong for a free-tier user.\n\nThe volume objection has a similar answer. Teams worry they can't read 3,000 conversations, and they can't, but they don't have to. Transcripts are structured data now. We point the Insight Analyst at ours and ask things like \"what do the 4s and 5s from enterprise accounts say about onboarding?\" That question takes seconds to answer against interview transcripts, and it's unanswerable against a pile of two-word text-box fragments.\n\n## The bet I'd make\n\nIf your NPS program is on trial internally, don't spend the quarter debating replacement metrics. CSAT and CES inherit the exact same flaw the moment you bolt the same dead text box onto them. Keep the score, keep the trigger, and replace the follow-up with something that listens. Run it on one segment for a quarter and compare what reaches the roadmap. My bet is your issue was never the eleven-point scale; it's that for fifteen years the question after it never had a second question.