Significance is not magnitude
A tiny lift can become statistically significant with enough traffic. Always inspect the absolute and relative effect alongside the test statistic.
The p-value is conditional
A p-value is not the probability that variant B is better. It measures how compatible the observed data are with the null model used by the test.
Confidence thresholds are conventions
A 95% confidence convention is common, but the right evidence threshold depends on the cost of being wrong and the decision being made.
Multiple comparisons matter
Testing many metrics or variants increases the chance that at least one result looks significant by chance. Pre-specification and appropriate correction help.
Pair statistics with judgment
Use statistical evidence with effect size, experiment quality, guardrail metrics and business context.
Related tools and guides
Have two options to compare?
Create a focused preference test and collect evidence before the decision gets expensive.
Create a survey