All project evidence

AB Advisor

Completed portfolio build

A Bayesian experimentation workspace that turns posterior probability, expected lift, credible intervals, and expected loss into an evidence-based ship decision.

Problem & approach

Problem

An experiment needs a decision that accounts for uncertainty and downside.

Approach

Use posterior probability, expected lift, credible intervals, and expected loss to frame a ship decision.

Evidence

Explore the Bayesian workspace and tests in the repository.

Product decision review

Who it serves

Product and experimentation teams deciding whether an observed treatment effect is useful enough to ship without harming guardrails.

The decision this design supports

Frame the decision around uncertainty, minimum useful effect, and expected downside. The implementation supports primary, guardrail, and secondary metrics rather than treating every positive movement as a launch signal.

Implemented workflow

  1. Import experiment data and confirm user assignment, metric type, role, and direction.
  2. Set priors, minimum useful effect, and decision thresholds before interpreting the recommendation.
  3. Check assignment quality, posterior intervals, probability of improvement, and expected loss.
  4. Export a review report and retain the input and configuration so the decision can be reproduced.

Design tradeoff

Conjugate models make interactive analysis practical, but model choice and priors still matter. A simpler binary conversion analysis may be enough for one metric; revenue with many zeros requires a different model and more careful assumptions.

What to measure

Inspect posterior lift, credible intervals, probability of clearing the minimum effect, expected loss, and guardrail harm. Sample-ratio mismatch is a reason to investigate assignment before trusting downstream estimates.

Current scope

Bundled sample datasets are synthetic. The tool analyzes supplied data; it does not assign users or establish that an experiment was run correctly.

Next evaluation

Stress-test decisions under weak effects, guardrail regressions, missing observations, and prior sensitivity. Preserve reports for all scenarios, not just clear wins.

Inspect the implementation

Release history

GitHub releases are tagged versions. Implementation code and commits can exist without a published release.

Inspect repository and README · GitHub, new tab