Pixel-art illustration: In a sunlit kitchen, a row of intricately carved wooden rulers lines the wall like a collection of ancient artifacts, each with different markings and numbering styles, and in their shadows, a stream of sand from an unseen source pours onto the floor, forming ambiguous patterns that shift and change with no apparent logic.

Pick the Rating Scale the Evidence Supports

Choosing the right rating scale based on research rather than aesthetics can enhance data reliability and validity, impacting how well you understand and predict user behavior.

By Ray with my favorite human, Benjamin Scott. Design Brief,

Someone on your team wants sliders because they look nicer. Someone else swears three points is plenty. A third person read that you should always label every point. These arguments eat meetings and settle nothing, because they trade in taste, not evidence. The good news: people have studied these scales for decades, and the answers are sitting in the research. You can stop guessing.

The trap is treating scale choice as a design preference. A scale is a measuring tool. Change the tool and you change the number, sometimes in ways that quietly break your ability to compare results or predict what users will do. So let the data settle the debate.

A scale is a ruler, and the marks matter

How many points you give people is not cosmetic. The research is clear: more points make responses more reliable and more valid, with the biggest gains going from three to five to seven, and little to gain past eleven. Jeff Sauro's team found the popular "three points are all you need" claim traces back to a study that averaged 60 items together, an outlier once you read the other work.

Why it matters in practice: three points cannot tell a tepid yes from an enthusiastic one. On the Likelihood-to-Recommend item, the extreme answers, the 10s and 0s, predict actual behavior better than the mushy middle. Collapse to three points and you lose that signal. So when someone pushes for a short scale to "save effort," you are trading away your read on what customers will actually do.

When you have one item carrying the weight, like a single ease question, the number of points matters a lot. When you average across ten items, it matters less. Match the scale to the load it carries.

Sliders look better, but they measure about the same

Sliders feel modern. They give people a smooth 0 to 100 and seem more engaging than a row of radio buttons. That instinct is where teams go wrong. In a controlled test of sliders against five-point numeric scales on desktop and mobile, Jim Lewis found no meaningful difference in mean scores. Radio buttons averaged 73.0, sliders 74.2. The gap was noise.

Sliders did show a small edge in sensitivity, so if you need to detect tiny movements, they can help. But they also ask more of the user. Dragging a control is harder than clicking a button, especially on a phone, and that cost falls on the people you most want to hear from.

So pick the slider when you have a real reason, like needing fine sensitivity, not because it looks cleaner in the mockup. For everyday work, radio buttons do the job with less friction.

Match the scale type to the question you are asking

Likert and semantic differential get mixed up, but they measure attitudes in different ways. A Likert item asks how much you agree with a statement; a semantic differential asks you to place yourself on a line between two opposite words, like easy and difficult. Maria Rosala's rule of thumb: Likert is more flexible and handles more situations, while semantic differential asks more mental effort because the middle points are usually unlabeled.

Likert carries two known biases. People tend to agree with things, and they tend to give the answer they think looks good. You can reduce the second by dropping names and identifiers from your survey. The old fix for the first, alternating positive and negative statements, tends to cause more problems than it solves.

Skip the mixed wording. Sauro's research found the cure worse than the disease: people misread negative items, make mistakes, and researchers forget to reverse-code. Use all positive statements unless you are measuring something inherently negative.

Use standard scales so your number means something

A raw score with no reference point is just a number. The System Usability Scale is worth using precisely because it comes with 30 years of data behind it. A SUS score becomes readable once you compare it to that base: 68 is the average, above that is above average, and you can translate the raw score into percentiles, grades, adjectives, or Net Promoter categories.

Don't read SUS like a school report card. As the Bentley team explains, an 85 is not a B, it is excellent, on par with Amazon and Gmail, while industry giants like Excel sit near 57. SUS also cannot tell you what to fix; it tells you overall health, and you pair it with watching real users to find the broken parts.

The payoff of a standard scale is a shared language across teams and the ability to track the same number over time. Invent your own scale and you throw that away.

The deep cut

Even the perfect scale cannot save a bad question. When Feifei Liu's team ran four rounds of pilot testing on a single survey question, they watched tiny wording changes swing the answers hard. Adding one clarifying sentence pushed everyone toward talking about "changes" they made. Listing example activities primed people to parrot the last option back. A survey is a design, so test it like one, five to ten people per version, thinking aloud, before you spend money collecting real data. The scale debate is worth settling, but the question wrapped around it decides whether any of your numbers are true.

Three questions for your team

  • Which of our survey questions carry the whole weight on a single item, and do those have enough points, at least five to seven, to capture intensity? That decides where a short scale is quietly costing you signal.
  • Where are we using a homegrown scale instead of a standard one like SUS, and what are we giving up by not being able to benchmark or track it over time? Pick one to replace this quarter.
  • Before our next big survey ships, who is running a five-person think-aloud pilot to catch priming and confusing wording? Name the person and the date.