Build a UX Benchmark You Can Track Over Years
By Ray with my favorite human, Benjamin Scott. Design Brief,
TL;DREstablishing a sustainable UX benchmarking process allows product teams to track improvements over time, ensuring design decisions are data-driven and aligned with user needs and industry standards.
Ask a room whether the last redesign made the product better and you get opinions. Nice charts, strong feelings, no proof. Benchmarking fixes that. You collect a few numbers that describe the real experience, then compare them to something meaningful and track them over time. Leaders get this wrong in two ways. They treat it as a one-time report card, and they drown it in metrics until nobody knows what the numbers mean. Here is how to run it so it lasts.
A number alone tells you nothing
A single metric is just a number floating in space. Your task success rate is 78 percent. Good or bad? You cannot say until you compare it to something. Kate Moran at Nielsen Norman Group lays out four reference points: an earlier version of your product, a competitor, an industry standard, or a goal your stakeholders set.
When you start, you have no earlier version to beat. So your first pass, the baseline, should lean on a competitor or an industry standard. Jeff Sauro's team publishes benchmark data for whole sectors like banking and hotels, and his numbers show the average task-completion rate sits around 78 percent, with 70 percent as the "good enough" line. That gives you a bar to clear on day one.
After that, your best comparison is your own past self. Same tasks, same setup, run again after the next redesign. That is the whole point of the word practice.
Pick a handful of tasks, not the whole product
You cannot test everything a user can do. Trying to is how programs die before they ship. Thomas Stokes calls out task selection as a place teams sink themselves: too many tasks and the study never launches, too few and you miss what matters, tasks too small and you learn nothing worth knowing.
Aim for 5 to 10 tasks that each map to a real user goal with a clear start and end. "Log in and pay your December bill" is a task. "Find the signup button" is not. To pick them, run a top-tasks survey where users choose the five that matter most to them. A small set will pull ahead of the pack. Benchmark those.
Lock the wording, the starting point, and the success rule now. You will reuse them every round, so vague tasks cost you later.
Fewer metrics, chosen on purpose
The instinct is to measure everything. Stokes has seen studies stack SUS, UMUX-lite, NASA-TLX, and SMEQ all at once, four tools measuring nearly the same thing. That bloats your planning, tires out participants, and blurs the report. Nobody reads a wall of numbers and knows what to do.
Use the ISO definition of usability as your filter: effectiveness, efficiency, and satisfaction. Pick one metric for each. Task success rate covers effectiveness. Time-on-task covers efficiency. A quick post-task rating plus something like UMUX-lite covers satisfaction. That is a full picture with a short list.
Sample size matters as much as which metrics you pick. At ten people, your success rate could land anywhere from 26 to 74 percent, which tells you nothing. Push toward a real sample and the range tightens enough to trust. Ask yourself how big a difference has to be before you would act on it, then size the study to catch it.
Copy competitors as inspiration, not as gospel
Benchmarking pulls you into looking at rivals, and that is useful. Krzysztof Radzik at Boldare warns about the traps that come with it. One is getting overly inspired and lifting a competitor's pattern whole. Another is copying an example that looks smart but works against your product. A third is the opposite: reinventing common patterns your users already understand, just to feel original.
Users know certain patterns cold. Fighting that makes them work harder for no reason. But if you match every rival exactly, your product blends into the crowd. Radzik suggests running a quick SWOT on each competitor's approach so you name the strengths and risks before you borrow anything.
Benchmarking is evolution, not revolution. You adapt what you find to your own goals rather than chasing novelty or aping the market.
The deep cut
A benchmark is worth nothing if you cannot run it again the same way. This is the mistake that guts programs after round one. You finish the baseline, stakeholders love it, then a year later you go to repeat it and cannot recall your sample size, your task wording, or how you cleaned the data. Now you are guessing, and your comparison is not apples to apples anymore.
Document everything during the first run: methodology, task phrasing and success criteria, which metrics and how you measured them, your analysis steps, and where the report lives. Stokes frames this as the difference between a one-off study and a program that survives you. And when you report, do not stop at the graphs. Use What, So what, Now what: what you found, why it matters to users, what the team should do. That last part is why the benchmark earns its budget and keeps earning it.
Three questions for your team
- Which 5 to 10 tasks actually represent what our users come here to do, and can we phrase each one with a clear start and finish we will reuse every round? Run a top-tasks survey before you argue about it.
- For our baseline, what are we comparing against, an industry standard or a named competitor, and is our success rate above the roughly 70 percent good-enough line?
- If a new researcher had to repeat this study next year, could they do it from our notes alone? If not, what do we document now so the comparison holds?



