
Open the visual model ↗
Start with a decision, not a button color
Write the decision the experiment will inform: whether to continue an offer, revise its promise, increase fulfillment capacity, or stop. Then state one learning question and the smallest contrast that can answer it. An offer combines audience, promise, price, term, placement, and fulfillment. Changing several at once may test a package, but it cannot reveal which component produced the difference.
Choose the method for the uncertainty. Interviews or usability sessions can expose confusion; a live offer can show observed action. The UK Government Digital Service advises planning around research objectives and questions, and notes that interview or usability rounds often use a handful of participants while surveys, benchmarking, and A/B tests need far larger samples for clear quantitative findings. A small publisher should therefore describe a tiny live comparison as directional, not as a precision experiment.
- Prewrite eligibility, exposure, primary outcome, cost limit, duration, and stop conditions.
- Count unique eligible readers, unique exposed readers, unique paid readers, and unique fully refunded readers consistently.
- Keep delivery quality equal enough that the comparison remains interpretable.
Create an experiment card
The card should fit on one page and be frozen before launch. Identify the operational owner and the person allowed to stop the test for service, trust, or capacity reasons. Record exclusions such as current paid members, people already shown another price, staff accounts, and addresses without the needed consent. Define exposure as the event that genuinely presented the offer, not merely membership in a mailing list.
| Field | Example entry | Why it matters |
|---|---|---|
| Decision | Continue a two-week briefing pilot? | Connects evidence to an action |
| Eligible audience | Active free readers meeting stated rule | Defines the denominator |
| Contrast | $24 pilot vs $30 pilot; same promise | Limits what changed |
| Primary outcome | Unique paid readers not fully refunded by review date | Prevents metric shopping |
| Capacity stop | Pause at 12 unique paid readers | Protects fulfillment |
| Review date | Seven days after delivery | Allows refunds and service checks |
Calculate descriptively and preserve uncertainty
The following numbers are hypothetical, with at most one qualifying purchase per reader. Group A has 42 unique exposed readers, four unique paid readers, and one of those four fully refunded, leaving three net paid readers: 7.1%. Group B has 39 unique exposed readers, five unique paid readers, and no full refunds, leaving five net paid readers: 12.8%. The descriptive difference is 5.7 percentage points after rounding.
That difference is not proof that B caused more purchases. The groups are small; eligibility, exposure, timing, or chance may explain it. Report raw counts beside rates. Also report cost and delivery evidence. If B's promise requires ten extra fulfillment hours, a higher descriptive rate may still be a worse operating choice.
| Measure | Group A | Group B | Interpretation |
|---|---|---|---|
| Unique exposed | 42 | 39 | Rate denominators |
| Unique paid readers | 4 | 5 | One qualifying purchase per reader |
| Unique fully refunded readers | 1 | 0 | Subset removed from net outcome |
| Net paid outcome | 3 / 42 = 7.1% | 5 / 39 = 12.8% | 5.7 pp descriptive difference |
| Extra fulfillment hours | 0 | 10 | Commercial tradeoff |
Decide with a result packet
Freeze a result packet containing the experiment card, audience query or rule, creative, delivery timestamps, counts, refunds through the review date, service incidents, costs, and deviations. A result can be unusable because tracking broke or one group received the wrong page. Say so. Do not repair missing evidence by choosing a more flattering denominator after the fact.
Choose one disposition: adopt as a limited offer, repeat because an operational failure obscured the test, revise the promise after qualitative research, or stop. If repeating, write why another test could change the decision. Preserve the first result rather than pooling unlike versions. Over time, a publisher can build better priors about audience response, but each offer still needs its own capacity and contribution check.
Continue the work
Sources & limits
This is an editorial pilot method, not a statistical test plan. Small descriptive differences should not be described as causal, significant, or generalizable without appropriate design and analysis.
- Plan user research for your service
Primary source · Publication date not stated · Primary source checked · 19 September 2026
Source claims and editorial judgments remain separate. Send a correction with the passage and supporting evidence.
