Home About Me Bayesian Calculator Contact

Simulate A/B-test

This calculator allows you to model A/B test outcomes by simulating a synthetic user population. It is designed to show how unobserved differences between user segments impact the outcome of your A/B-test experiment.

By running multiple simulations under varying demographic assumptions, you can see exactly how your hypotheses impact the final results. Beyond simulation, the tool use AI to automatically summarize your findings.

Ultimately, it serves as a validation tool: by comparing real-world A/B test data against the predicted outcome based on your assumptions, you can see where they align or diverge. If your actual test results don't align with your initial assumptions, simply head back to the drawing board, adjust the parameters, and simulate a new scenario.

It’s the perfect sandbox to stress-test your hypotheses before you ever go live.


Instructions:

  • This tool simulates an A/B test using two variants: a control and a variant 1. The simulation models a population of n users, which you can adjust via the sample size setting below. In the control group, the baseline performance is determined by the "Metric baseline" parameter, which is also customizable.
  • You can configure up to six distinct user segments by defining how the new feature impacts the metric. Simply set the score to "Love" if you expect a positive reaction, "Neutral" for indifference, or "Hate" if you anticipate a negative response.
  • You can weight each segment to accurately reflect its proportion within your real-world user base.
  • Once configured, click Run Simulation to execute the synthetic A/B test and view the AI-generated summary of the results.
  • Be patient; the A/B simulations will take some time, and then AI will summarize the results for you.

Synthetic Users


Sample size

AI summary of the simulated A/B test using synthetic users.



The A/B test shows a statistically significant improvement in the metric for Variant 1 compared to the Control Group, with Variant 1 achieving a metric of 21.7 versus the Control Group's 19.9, resulting in a delta of -1.7 and a p-value of 0.0, leading to the rejection of the null hypothesis.

While the observed results for the Control Group (19.9) and Variant 1 (21.7) show a deviation from their respective ground truth values of 20.0 and 20.05, this deviation is consistent across both groups, with a difference of 19.7 for the Control Group and 19.75 for Variant 1.

The overall positive delta in Variant 1 is influenced by segment-specific contributions; Segment Hans-Ueli S. significantly boosts the delta with a contribution of 3.6, driven by a 24.0% weight and a substantial effect of 0.15.

Conversely, Segments Markus M. and Léa M., each with a 20.0% weight, negatively impact the delta, contributing -1.0 each with an effect of -0.05, thus partially offsetting the gains from other segments.

Segments Sofia R., Marco T., and Priya N. have no impact on the delta, as they contribute 0.0 to the overall delta, despite accounting for a combined weight of 36.0%.






Breakdown of the Delta

This plot illustrates the mechanics behind our A/B test results.

Because of randomness, observed values will always deviate slightly from the ground truth. As sample size decreases, this deviation (or noise) becomes significantly larger.


Distribution of Synthetic Population

This plot shows the assumed distribution of the synthetic population.