Method
How we read an experiment
A repeatable sequence for A/B test analytics. Not a product. The same order every time, so a messy log cannot skip the unflattering checks.
1. Intake is a list, not a conversation
We will not start arithmetic until these sit on one page: hypothesis, primary metric formula, intended split, first-assignment timestamp, overlapping tests, and whether anyone watched live numbers. If the timestamp is “sometime in March,” the session waits.
2. Assignment before outcomes
Daily counts for each variant. A two-point drift is enough to pause. We also ask whether the unit you randomised (user, session, shop, account) is the unit in the spreadsheet. Mixing those is a common way a Malaysian marketplace test looks healthier than it is.
3. Windows that include the country
If the test ran through a festive period, we say so in the memo instead of averaging it away. Hari Raya and year-end sales are not a footnote when checkout is the primary metric.
4. One primary, a short guardrail list
Guardrails are pre-declared. A metric discovered while browsing the extract is labelled as exploratory. It may still be interesting; it does not get to veto or crown the test unless the brief already named it.
5. The memo has a verb
Ship, hold, iterate (with the element named), or stop. Interval estimates sit beside the point estimate. We stay on the working session while the team argues. The second reader’s comments are visible, not hidden in an appendix.
6. What we refuse
We will not polish a mid-flight metric change into a pre-registered success. We will not treat a recycled holdout as a control. We will not hurry a readout to match a launch date if the assignment file is incomplete.
If this sequence matches how you want a test closed, request a readout or look at the experiment readout itself.