Pivix Logo
Back to glossary

Sample Size

Sample size is the number of visitors or observations included in each variant of a test, large enough to detect a meaningful difference with acceptable reliability.

Key takeaways

  • Sample size determines test reliability and validity in CRO.
  • Larger samples needed for small effect sizes and high confidence.
  • Power analysis calculates necessary sample size pre-test.
  • Low traffic can make reaching needed sample size impractical.
  • Large sample size doesn't fix flawed test design.

In depth

Sample size refers to the number of observations or participants used in a statistical test. In conversion rate optimization (CRO), it typically involves the number of visitors interacting with different test variants. The sample size is crucial as it impacts the reliability and validity of the test results, helping ensure that any observed differences between variants are statistically significant and not due to random chance.

The required sample size for a test is influenced by several factors, including the baseline conversion rate, the minimum detectable effect size, and the desired confidence level and statistical power. Larger sample sizes are necessary to detect smaller effect sizes or when aiming for higher confidence. Conversely, smaller sample sizes may suffice for larger effects or lower confidence levels, but they risk producing less reliable results.

In practical application, determining the appropriate sample size involves performing a power analysis before launching a test. This calculation helps set expectations and timelines for running the test. In a quiz funnel using Pivix, calculating the sample size beforehand ensures that the test runs for an adequate period and reaches a reliable conclusion without premature decisions.

Sample size considerations have limits. In cases of low traffic, achieving the necessary sample size can be impractical, making some tests infeasible. Additionally, large sample sizes do not necessarily guarantee meaningful insights if the test design is flawed. Misinterpretations can arise if the underlying assumptions of the test, such as randomization or independence, are violated.

Example in practice

Suppose that before testing a new lead-capture step, an analyst computes that detecting a 1.5-point lift on a 9% baseline at 95% confidence needs about 8,400 visitors per variant. With 2,000 weekly quiz starts, the test would run roughly four weeks, so they schedule it accordingly instead of stopping early.

How to measure it

The effectiveness of a sample size in a test can be assessed by checking if the results reach statistical significance. This involves calculating the p-value of the observed differences between test variants; a p-value lower than the chosen confidence level (e.g., 0.05) indicates significance. If the test does not reach significance, it may suggest the sample size was too small.

Another way to measure sample size effectiveness is through the confidence interval, which provides a range within which the true effect size likely falls. A narrow confidence interval indicates a more precise estimate, often achieved with a larger sample size. If the interval is too wide, it suggests that the sample size is likely insufficient to provide reliable insights.

Common mistakes

One common mistake is not calculating the sample size before starting a test. Skipping this step can lead to tests running with insufficient data, resulting in unreliable conclusions. Teams should use a power analysis to determine the necessary sample size and plan their testing schedule accordingly, ensuring that they collect enough data to make informed decisions.

Another mistake is stopping the test too early, based on initial results. This often happens when teams see promising early data and decide to implement changes without reaching the calculated sample size. It's essential to wait until the predetermined number of observations is met to ensure that the results are statistically significant and not just a product of random variation.

Frequently asked questions

What happens if my sample size is too small?

An under-powered test produces noisy, unstable results that can falsely declare a winner or miss a real difference. You risk shipping changes that do not actually improve conversions.

How do I calculate the right sample size?

Use a power calculator that takes your baseline conversion rate, minimum detectable effect, confidence level, and statistical power. It returns the visitors needed per variant before you launch the test.

Can I run an A/B test on a low-traffic quiz funnel?

You can, but small effects may require more traffic than you generate in a reasonable timeframe. Calculate sample size first to confirm the test is feasible, or focus on bigger, bolder changes.

Why is sample size important in A/B testing?

Sample size is crucial in A/B testing because it impacts the reliability and validity of the results. An adequately sized sample helps ensure that any observed differences between variants are statistically significant and not due to random chance.

Can a test be reliable with a small sample size?

A test can be reliable with a small sample size if the effect size is large and the confidence level is low. However, this approach increases the risk of error and is generally not recommended for nuanced insights or high-stakes decisions.

What is a power analysis in the context of sample size?

A power analysis is a calculation used to determine the necessary sample size for a test. It considers factors like baseline conversion rate, effect size, confidence level, and statistical power to ensure the test is likely to detect meaningful differences if they exist.

Related terms

Turn glossary theory into qualified leads

Build a scorecard quiz funnel that qualifies and captures leads in minutes — no code required.

Start for free
  • No credit card
  • Free plan
  • Launch in minutes