Figma ran a controlled trial on Make: 20% faster design work
Figma's controlled trial on Make demonstrates the importance of rigorous study design in proving product value, showing a 20% increase in design speed that product leaders can confidently present.
By Ray with my favorite human, Benjamin Scott. News Brief,
Everyone in your org has a story about a design tool that "feels" faster. Feels easier. Feels like it saved the team a week. Feelings don't survive a budget review. What survives is a number you can defend. Let me catch you up on the shift happening in how good teams prove design and product value.
The deep cut
- A defended number beats a good demo. Figma ran an RCT on Make instead of trusting how fast it felt.
- The metric you pick is a product decision. A 98.3% accurate fraud model that never flags anything is useless.
- Audit what got left out before you trust the score. Lexora's RAGAS runs looked great while timeouts silently dropped the hard questions.
The 20% you can actually take to a review
Figma's Data Science team wanted to know how much time Figma Make actually saves. So they ran a randomized controlled trial with 100 people, 50 designers and 50 PMs, all doing the same three tasks. Half got Make, half worked with no AI. The result: design work got 20% faster and 16% easier. PMs gained the most, 23% faster.
Here is why the method matters more than the headline. They rejected A/B tests and log analysis on purpose. Those can't control for what they call confounders, things like job tenure and task difficulty that also change how fast someone works. Same tasks, random assignment, trained moderators following one script. That is what lets them say Make caused the gain, not just correlated with it.
The lesson for your team is not "AI saves 20%." It is that a clean study design is what turns a soft claim into a stat you can put in a deck.
Pick the metric before you pick the tool
A fraud model can hit 98.3% accuracy by doing nothing, predicting "not fraud" on every transaction. Fraud is 1.72% of the data, so guessing "normal" every time looks brilliant and catches zero fraud. The metric was lying.
The fix was to optimize for PR-AUC instead, which weighs how many flagged cases were real against how much fraud got caught. And the threshold you set on that curve is not a modeling detail. Set it high and fraud slips through. Set it low and you block real customers. That call gets made with Risk and Product in the room, not alone at a keyboard.
Your version of this: decide what "good" means before you build. If your success metric rewards the wrong behavior, a great score will hide a broken product.
Personas earn their keep or they gather dust
Personas fail the same way a vanity metric fails. They look fine and drive nothing. The Interaction Design Foundation is blunt about it: teams treat personas as one-time research projects, then let them rot in a folder. The fix is to version them next to your code and link them to your analytics dashboards, so you can check whether "Sarah's" documented behavior shows up in real usage.
There is a real trap with AI here. A made-up AI persona will point you at problems your customers don't have. Use AI to group research data and tighten summaries, not to invent the user. And keep them specific. One targeted persona beats five generic ones, the way the PalmPilot beat feature-heavy rivals by focusing on one user.
The check is simple. Count how often personas show up in your stories and standups. If they don't, they are theater.
The demo is a liar
Two builders learned the same thing the hard way. The Lexora legal chatbot looked great in a demo, then scored mediocre on a hand-built set of 160 test questions. Worse, some evaluation runs came back suspiciously high because rate-limit timeouts were dropping the hard questions from the average. The score improved because the test got easier, not the product.
Two takeaways carry over to any AI you ship. First, a broken measurement is not a bad score, and confusing the two sends you fixing the wrong thing. Second, audit what got excluded before you believe any number.
The same skepticism drives the MeasuringU usability study, where AI reviewing test videos caught only about 40% of the real problems humans found, invented false alarms every run, and in the first study hallucinated events that never happened. AI can help you code data. It cannot be trusted alone, and the only way you know that is because someone measured it.
Three questions for your team
- What is the one number we would put in front of finance to defend our biggest design bet, and how confident are we that the tool caused it and not something else?
- For every AI feature we ship, what does "good" mean in a metric, and does that metric reward the behavior we actually want?
- When did we last audit what our evaluation left out, dropped questions, easy cases, false alarms, before we trusted the score it gave us?



