Two groups answered differently.Is the gap real?

You ran a survey, split the answers two ways, and one group said yes more often. Before you build a quarter around that gap, run it through a two-proportion z-test. This tells you the p-value, both rates, and whether the difference is a signal worth acting on or noise dressed up as an insight.

Try

Your two groups

How many people in your first group answered the question.

How many of them gave the positive / target answer.

How many people in your second group answered the question.

How many of them gave the positive / target answer.

The verdict

Significant difference

p-value

0.0

Probability this gap is just chance.

Difference

9.8%

Group B rate minus Group A, in points.

Group A rate

46.0%

Group B rate

55.8%

Yes-rate, side by side

Group A
46.0%
Group B
55.8%

p-value 0.0018 · z-score 3.12 · needs p < 0.05

Sensitivity · would it survive a stricter bar?

The same gap at each confidence level

ConfidenceNeeds p <Verdict
90%0.10Significant
95%0.05Significant
99%0.01Significant

Your p-value is 0.0018. A gap that clears 90% but not 95% is a lead to confirm, not a finding to ship — the stricter the bar it survives, the more confidently you can build on it.

Your move

Significance tells you it's real. I tell you what to do about it.

Group A 46.0% vs Group B 55.8% (+9.8 pts) — p = 0.0018, significant at 95%.

A p-value settles whether the gap is noise. It doesn't tell you which finding is worth a campaign, a price change, or a new segment. Bring me the survey behind these numbers and I'll show you, free, which differences are actually worth building on — and the move I'd make first.

Plain English

A gap on a chart is not a finding.

When two survey groups answer a yes/no question at different rates, the difference can come from one of two places: a real underlying difference between the groups, or pure sampling luck. With small samples, luck does a lot of the talking. Flip a fair coin 20 times in two rooms and you'll often see one room land 60% heads and the other 45% — no difference exists, the dice just rolled that way.

A two-proportion z-test separates those two cases. It asks: if the two groups were actually identical, how often would random sampling alone produce a gap this big or bigger? That probability is the p-value. A small p-value means a coincidence this large would be rare, so the difference is probably real. A large p-value means you can't rule out noise.

This calculator runs that exact test on your numbers. Enter each group's size and how many said yes, pick your confidence level, and it returns the p-value, each group's rate, the gap in percentage points, and a verdict you can actually act on — significant, or could be noise.

The formula

z = (pB − pA) ÷ √( p̄·(1−p̄)·(1/nA + 1/nB) ), where p̄ = (yesA + yesB) ÷ (nA + nB)

Group A: 230 of 500 said yes (46%). Group B: 290 of 520 said yes (55.8%). Pooled rate p̄ = 520/1020 = 51%. Standard error = √(0.51·0.49·(1/500 + 1/520)) ≈ 0.0313. z = (0.558 − 0.46) ÷ 0.0313 ≈ 3.12, giving a two-sided p-value ≈ 0.0018. At 95% confidence (needs p < 0.05) that gap is significant — not noise.

How to read the p-value

The p-value is the probability of seeing a gap at least this large if the two groups were truly identical. Lower means more surprising under 'no difference,' so more likely a real effect. You compare it against your confidence threshold — these are the conventional bars:

p-valueWhat it means
p < 0.01Strong evidence. A gap this large would almost never happen by chance — clears the 99% bar.
p < 0.05Statistically significant. The standard bar (95% confidence) — unlikely enough to be chance that you treat it as real.
0.05 ≤ p < 0.10Borderline. Clears 90% but not 95% — a lead to confirm with more data, not a result to ship.
p ≥ 0.10Not significant. You can't rule out sampling noise — the gap could easily be luck.

Conventional significance thresholds — not a published dataset. Pick your bar before you see the data, and pair the verdict with the actual gap in points.

Your result came back 'not significant.' Now what?

01

The gap looks big but p is high.

Your samples are too small. A 10-point gap on 60 people per group is well within the range of random noise. Collect more responses before you trust the direction, let alone the size.

02

Both rates are near 50%.

Variance is highest around 50%, so you need more responses to detect a difference there than at the extremes. Don't read a near-50/50 split as 'no difference' on a thin sample.

03

It's significant at 90% but not 95%.

That's a borderline read, not a green light. Treat it as a lead to confirm, not a result to ship. The stricter the bar it clears, the more confidently you can build on it.

04

Huge sample, significant, tiny gap.

With enough responses, even a one-point difference goes 'significant.' Significant means real, not large. Ask whether a gap that small is worth acting on at all.

Run surveys that actually settle the question

01Power it before you launch

Estimate the sample you need to detect the gap you care about. Most 'inconclusive' surveys were underpowered from the start, not unlucky.

02Pick the confidence bar up front

Decide on 90, 95, or 99% before you see the data. Choosing the threshold after the fact is how noise gets promoted to insight.

03Compare clean, separate groups

The two groups must be independent and mutually exclusive. Overlapping or self-selected segments break the math and the conclusion.

04Watch the base sizes, not just the percentages

55% vs 46% reads the same whether it's on 50 people or 5,000 — but only one of those is trustworthy. Always show the n.

05Keep the question binary and unambiguous

This test compares two yes/no rates. A double-barreled or leading question contaminates both groups equally and still gives you a clean-looking, meaningless p-value.

06Don't peek and stop early

Repeatedly checking and stopping the moment it goes significant inflates false positives. Set the sample size, collect it, then test once.

07Report the effect, not just the verdict

Always pair 'significant' with the actual gap in points. A decision needs the size of the difference, not only its existence.

08Replicate before you bet the roadmap

One significant survey is a strong hint. A second one that agrees is a finding. Big calls deserve confirmation.

The vocabulary

p-value
The probability of seeing a gap at least this large if the two groups were truly identical. Smaller means the difference is harder to explain by chance alone.
Statistical significance
A result is significant when its p-value falls below your chosen threshold (e.g. 0.05 at 95% confidence) — unlikely enough to be chance that you treat the difference as real.
Two-proportion z-test
The test for comparing two yes/no rates. It pools the groups, estimates the noise, and scores how many standard errors apart the two rates sit.
Confidence level
How sure you want to be before calling a difference real. 95% confidence means accepting a 5% chance of a false alarm.
Pooled proportion (p̄)
The combined yes-rate across both groups, used to estimate the standard error under the assumption that the groups are the same.
Sampling noise
Random variation between samples drawn from the same population. It's why two identical groups rarely return the exact same rate.

Survey significance questions, straight answers

Run a two-proportion z-test: it returns a p-value, the probability of seeing a gap this big if the groups were actually identical. If that p-value is below your threshold (0.05 for 95% confidence), the difference is statistically significant — unlikely to be chance. This calculator does that for you: enter each group's size and yes-count, pick a confidence level, and read the verdict.

A calculator tells you what. A call tells you what to do about it.

Send me the account behind these numbers. I'll tell you straight where the money's leaking and what I'd fix first — free, and you keep it whether you hire me or not.