What Is a Stat Test?
A stat test, or more formally a hypothesis test, seeks to find out if an observed effect is real or if it is due to chance. To do this, we compare an observation from a sample to some other expectation. For example, we may observe in a study that women like chocolate more than men, and we could stat test this against the expectation that men and women like chocolate equally.
Stat testing allows us to identify reliable trends for an audience — it is structured to confirm whether the patterns we see in our sample are supported, or whether they're just due to chance.
The goal of stat testing is to determine whether a sample provides enough evidence to suggest a trend exists in the broader audience.
For example, with Idea-to-Idea stat testing:
Ask a question — "Is Idea 1's score significantly higher than Idea 2's?"
Collect data — survey or measure a group of people who represent the audience we care about (e.g., the broader group is Gen Z).
Run the test and interpret the results — check if the data supports a clear conclusion on our sample; that is, whether Idea 1's score is significantly higher than Idea 2's.
Explain It Like I'm 5
Stat testing an Idea Split is like a lemonade stand competition between two kids, Abigail and Bradley. In this competition, we want to find out if one kid sells significantly more lemonade than the other.
How does it work?
Imagine you're watching a lemonade stand competition between Abigail and Bradley. Both kids are selling lemonade to people passing by, but we don't know who sells more, so we count how many cups of lemonade each kid sells. This is similar to what the Idea Split stat test does: it looks at two ideas (like Idea A and Idea B) and checks if there's a big enough difference in the responses.
Before the competition starts, we guess that both kids will sell the same number of cups (our starting guess, or "null hypothesis") — but we're open to being proven wrong (the "alternative hypothesis"). After the competition, we compare the number of cups Abigail and Bradley sold. If the test shows a big enough difference, we can say our starting guess was wrong and that one kid sells significantly more lemonade.
What Is a Confidence Level?
Confidence level is how confident we are that the answers are representative of the broader audience. Testing at a 95% level means we'll only find significance when there is less than a 5% chance the difference is random — in other words, we are 95% confident the difference in results is non-random.
95% = Highest certainty
90% = Balanced
80% = More exploratory
What Is a Significant Difference?
A significant difference means that the data from our sample provides enough evidence to suggest the trend is true for the broader audience.
How Do We Decide What Is Significant?
Question — We start with a question about the audience: is a trend seen in the sample likely to be true of the whole audience?
Evidence — We use different stat tests for different types of data (e.g., proportions tests, t-tests). For idea-to-idea stat testing, we use a custom test. The goal of any stat test is to determine if the data gives us enough evidence to imply a trend exists in the broader audience.
Results —
Significant difference: the data provides enough evidence to suggest the trend is true of the broader audience.
Non-significant difference: the data does not provide enough evidence to make a judgement about the broader audience.
Significance vs. Actionability
Significance and actionability are not the same thing.
Significance: the result would be unexpected if there was no trend in the broader audience.
Actionability: the result is meaningful and/or impactful.
A result can be statistically significant but not meaningful enough to drive a business decision.
How Can Something Be Significant, but Not Actionable?
Example: A client is exploring a new package design. In an Upsiide study, the old design scored 76, and the new design scored 78.
Assume this difference is significant. However, it's likely not actionable — consumers seem to prefer the new design, but they like the old one nearly as much. In other words, it may not be worth the investment to move forward with the redesign.
Factors Influencing Trend Detection
Three main factors make it easier to find trends:
Sample size — collecting more respondents makes underlying trends more consistent.
Consistency of the scores — when respondents agree on an idea, trends are easier to spot.
Size and stability of the difference — larger differences among idea scores make it easier to spot trends, and this is reinforced when two idea scores have a high correlation, which makes the differences between them more consistent.
The Possibility of Counter-Intuitive Results
The reliability of the difference — not just its size — drives significance tests:
A small difference in score may be significant if the sample is large and the scores are consistent or trend together.
A large difference in score may not be significant if the sample is small, respondents have mixed opinions on the ideas, or the ideas do not trend together. This is why, in filtered or smaller-sample views, an idea with a smaller score gap to the benchmark can be flagged significant while an idea with a larger gap is not — significance reflects the reliability of the difference, not just the raw score distance.
Frequently Asked Questions
Why do we keep saying “implies” and “suggests”?
We still aren't 100% sure! Significant differences are stated when the evidence we have passes some threshold — commonly either 90% or 95%.
One way to think about this: there is still a 10% or 5% chance the trend seen in the data is due to chance, not because a trend exists in the whole audience. This means we still aren't 100% sure — we didn't survey every single person in the audience group.
Saying a significant difference “implies” or “suggests” the trend exists in the whole audience is a way of communicating this remaining uncertainty. A 5% chance is low, but it is not none — we can't be fully certain unless we found a way to survey the whole audience.
Can we use respondent weighting?
No — like normal Idea Score calculations, you cannot typically use weighting. If weighting is required, please reach out to the Advanced Analytics team to discuss whether it will work for your project.
Why is a small difference significant and a larger one not?
Idea-to-idea stat testing looks for reliable differences. If a result is unreliable, it could be a fluke of the sample (caused by chance alone) rather than a real trend in our audience.
Small differences can be significant when larger ones are not because we are looking at reliability, not size. A smaller difference could be more stable — this stability lends to reliability, which is what drives significance tests.
Key Takeaways
Statistical testing allows us to identify reliable trends for an audience.
Statistically significant differences mean the data provides enough evidence to suggest the trend is true of the broader audience.
Significance vs. actionability: a result can be statistically significant but not meaningful enough to drive a business decision.
Four main factors make it easier to find trends: consistency of scores, sample size, size of the difference, and stability of the difference.
