Home About Me Bayesian Calculator Contact

Simulate A/B-test

This calculator allows you to model A/B test outcomes by simulating a synthetic user population. It is designed to show how unobserved differences between user segments impact the outcome of your A/B-test experiment.

By running multiple simulations under varying demographic assumptions, you can see exactly how your hypotheses impact the final results. Beyond simulation, the tool use AI to automatically summarize your findings.

Ultimately, it serves as a validation tool: by comparing real-world A/B test data against the predicted outcome based on your assumptions, you can see where they align or diverge. If your actual test results don't align with your initial assumptions, simply head back to the drawing board, adjust the parameters, and simulate a new scenario.

It’s the perfect sandbox to stress-test your hypotheses before you ever go live.


Instructions:

  • This tool simulates an A/B test using two variants: a control and a variant 1. The simulation models a population of n users, which you can adjust via the sample size setting below. In the control group, the baseline performance is determined by the "Metric baseline" parameter, which is also customizable.
  • You can configure up to six distinct user segments by defining how the new feature impacts the metric. Simply set the score to "Love" if you expect a positive reaction, "Neutral" for indifference, or "Hate" if you anticipate a negative response.
  • You can weight each segment to accurately reflect its proportion within your real-world user base.
  • Once configured, click Run Simulation to execute the synthetic A/B test and view the AI-generated summary of the results.
  • Be patient; the A/B simulations will take some time, and then AI will summarize the results for you.

Synthetic Users


Sample size

AI summary of the simulated A/B test using synthetic users.



The A/B test shows a statistically significant improvement for Variant 1, with an observed metric of 21.7 compared to the Control Group's 20.0, resulting in a delta of -1.8 and a p-value of 0.0, leading to the rejection of the null hypothesis.

Despite the overall positive result, the segments reveal a nuanced picture of how Variant 1 performed across different user groups.

Segments Markus M and Léa M each contributed -1.0 to the overall delta, driven by a weight of 20.0% each and a negative effect of -0.05, indicating they underperformed with Variant 1.

Conversely, Segment Hans-Ueli S. significantly boosted the delta by 3.6, accounting for 24.0% of the weight and demonstrating a strong positive effect of 0.15.

Segments Sofia R., Marco T., and Priya N. had a neutral impact on the delta, contributing 0.0 each, with weights of 12.0% and an effect of 0.0, suggesting Variant 1 had no discernible impact on these user groups.

The observed difference between the Control Group's metric and its ground truth is 19.8, while for Variant 1, this difference is 19.85.






Breakdown of the Delta

This plot illustrates the mechanics behind our A/B test results.

Because of randomness, observed values will always deviate slightly from the ground truth. As sample size decreases, this deviation (or noise) becomes significantly larger.


Distribution of Synthetic Population

This plot shows the assumed distribution of the synthetic population.