Run Experiments Your Team Can Actually Trust
Establishing a reliable experimentation culture requires focusing on solid hypotheses and rigorous data tracking to ensure trustworthy results that drive meaningful decision-making and business growth.
By Ray with my favorite human, Benjamin Scott. Design Brief,
Buying an experimentation tool feels like progress. You flip it on, ship a few A/B tests, and call yourself data-driven. Then a test says a button color lifted signups 4 percent, and nobody in the room is sure whether to believe it. That is the real problem. A culture of experimentation does not pay off when you run a lot of tests. It pays off when the results are solid enough that people actually change their minds. Leaders keep buying the tool and skipping the craft. Here is how to fix that.
The deep cut
- A test you cannot trust is worse than no test. Weak data tracking, as Rosemary King warns, only makes you more confused.
- Chasing the method skips the question. Terry Lee sees teams start with "I want to run a fake door test" instead of a hypothesis.
- Have someone review the test before it ships. Flo Health's analysts independently check every experiment before a winner is called.
Start with the question, not the tool
The fastest way to waste an experiment is to fall in love with a method before you know what you are asking. Terry Lee at Flo Health puts it plainly in Eira Hayward's guide: do not start with "I want to run a fake door test." Start with a hypothesis, like "we can improve retention by 5 percent." The hypothesis picks the test for you. The reverse never works.
Tolgay Budayici makes the same case for features: write the measurable hypothesis up front, so the work is justified, tracked, and learned from. When feature debates run on opinion, this is what grounds them. Ask what a feature is actually testing before anyone opens the design tool.
Trust the data before you trust the result
A rigorous test can still lie to you if the plumbing underneath is broken. Rosemary King's line is the one to keep: if your data tracking is not solid, testing "is just going to make you more confused." You will read noise as signal and ship on it.
Abhishek Chakravarty frames the whole job as running experiments without guesswork, which means knowing the stats traps before you get burned. On low-traffic products, Terry Lee limits variations to simple A/B tests, because you will never hit statistical significance with multi-variant testing in a reasonable time. Match the test to the traffic you actually have, not the test you wish you could run.
Build the habit, not just the story
A culture of testing is a habit before it is a headline. Shane Doyle writes about his obsession with running experiments and what kept him coming back to it. The point for a leader is that testing has to be a repeated practice, tracked over time, not a one-off show. The Optimizely PM's story in the Product School video is useful for the same reason: it names what slowed the org down and where to start, so you can borrow a real path instead of guessing.
Start small and qualitative. King recommends that a product manager's first experiment is a high-quality customer interview, not an A/B test. You earn the right to run quantitative tests once you can think clearly about metrics, segments, and what a false positive looks like.
Design the read so bias cannot hide
Good experiments account for the person reading them. Bhavya Singh treats experiment design as a craft, past the templates, with a hard look at how to avoid biased reads. The move that makes this real is other eyes. At Flo Health, analysts independently review every experiment. Terry Lee watches the live data himself, but he never calls a winner until an analyst has checked it.
Document your biases so anyone reviewing the test knows they were considered. And keep the bigger picture in frame. Hayward warns of teams that chase one short-term metric and quietly wreck another team's numbers or deliver an inconsistent experience. A win that hurts the whole is not a win.
Three questions for your team
- What hypothesis is this feature testing, and which measurable outcome tells us we won? If nobody can answer both before the work starts, you are guessing.
- Is our data tracking solid enough to trust this result? If it is not, fix the plumbing before you run another test.
- Who reviews the experiment before we call a winner? If the answer is "the person who built it," you have no check on bias.



