How arms work
In productionAn experiment splits one offer's traffic between its live version, the control, and one to three challengers, with the traffic weights you choose. A challenger is an ordinary version of the offer document, published as an arm: it goes through the same publish pipeline, and every visible option has to match on a real Shopify cart before it can take a visitor.
- Assignment is deterministic. The visitor id decides the arm, so a returning visitor lands on the same one. Resolve, cart and price all pick the version through one serving module, read from Postgres on every request, so the page and the cart never disagree about which arm a visitor is in.
- Server and browser see the same arm. Resolve and carts take the visitor id you pass, so a brand that renders on its own server keeps the id in a first-party cookie and sends the same id from the browser; the page it rendered and the cart then agree.
- A stale page cannot buy the wrong version. A cart call that sends the version the page showed, after the visitor has been moved to another (a paused test, say), is answered 409
offer_changedwith the offer to render instead. - Nothing is served before the store can charge it. The start writes every arm into the store's option map first. Until that write lands the experiment is
startingand every visitor gets the control; if Shopify refuses it, the experiment ends stopped and nothing was ever served.
Agents run the whole cycle through the API or MCP: create_experiment to start, get_experiment to read the report, pause_experiment and resume_experiment, stop_experiment to end it, and publish_offer to make the winner the live version.
The metric
Arms are judged on one number: profit per visitor. It is the realized contribution of the arm's first orders, in the current cost view and net of refunds, divided by the visitors the arm was shown to.
| Profit | Each first order's revenue and shipping charged, less refunds and every cost Offer Suite holds for its lines, shipping and payment processing. When a cost is corrected, the current view restates past orders, so the report always uses today's best costs. |
|---|---|
| Visitors | Distinct visitors whose page reported viewing the arm: a view_offer event with their visitor id. The Verified offer theme block sends it; a brand's own page posts it to the events endpoint with its public key. |
| Currency | One shop currency, the one most of the window's first orders were placed in. Money is never converted or added across currencies; orders in any other currency are left out and counted. |
| Left out | Renewal orders are not in the metric yet. Orders whose cart carried no visitor id stay in the arm's order totals and out of the metric. |
| Shown, not deciding | Conversion, buyers, average order value and refund percent sit beside the metric so you can see why an arm earns what it earns. |
Statistics
OF_DEV_KITS · v3 against v4 · profit per visitor, realized, current costs · specimen data
| Arm | Profit per visitor | Conversion | Buyers | Visitors | Traffic |
|---|---|---|---|---|---|
| v3control | $1.42 | 7.5% | 2,250 | 30,000 | 50.0% |
| v4challenger | $1.68 | 7.0% | 2,100 | 30,000 | 50.0% |
- A planned sample. At the start, the report sizes each arm from the control's last 30 days: the visitors it needs to detect the smallest lift you care about (10% by default) at 80% power. A control with no history gets no plan, and the report says the decision rests on the interval alone.
- An interval you can read every hour. The difference between each challenger and the control sits inside an asymptotic always-valid confidence sequence. It stays valid at every look, so checking the report hourly cannot manufacture a winner.
- One error budget for all challengers. The false-positive rate, 5% by default, is Bonferroni-split across the challengers, so adding challengers does not raise the overall chance of a false winner.
- A recommendation, never an action. Each comparison ends in one of four words: challenger wins, control wins, no difference, or keep running. Nothing acts on it: stopping and promoting are your agent's calls.
Guardrail
In productionEvery day at 06:17 UTC, and whenever an agent calls check_experiment, each arm of every running experiment is proved again: a real Shopify cart is built for each visible option and checked against what the option should charge.
- If any option of any arm no longer matches, or cannot be checked, the experiment pauses at once and every visitor is served the control.
- Carts already built keep the prices they were built at, so no shopper finds a different total at checkout.
- The pause reason names the failing versions and options. Fix the cause, publish the arm again, then resume: each visitor goes back to the arm they had.
Price tests
In productionA price test is an experiment whose challenger charges less. The challenger gives one-time lines a line_price below the catalog price, and Offer Suite's cart transform charges it as the product's own price in the cart and at checkout, with no strikethrough and no discount line. The shopper sees one price; the report judges whether the lower price earns more per visitor.
| Plans | Any plan where the cart transform is registered. On Shopify Plus it uses lineUpdate; on other plans the transform expands the line into itself at the fixed price. |
|---|---|
| Lines | One-time, non-gift lines. Subscription lines, gifts and lines already priced by a bundle or tier group cannot take a line price. |
| Direction | Price cuts only. A line price above the catalog price captured in the offer document is refused when the document is validated. |
| Catalog drift | If the live catalog price later falls to or below the line price, the transform leaves the catalog price in place and the verifier reports it. |
Evidence
A price test running on the dev store: the report from the dev API's experiment read, small sample and all, with each arm's charge taken from its version. It is a running example of the report, not a result.
Price test on the dev store · Running since 29 September 2026 · report window to 2 October 2026 · profit per visitor, realized, current costs, USD
| Arm | Charges | Profit per visitor | Conversion | Buyers | Visitors | Traffic |
|---|---|---|---|---|---|---|
| v1control | $69.95, catalog price | $5.70 | 7.7% | 1 | 13 | 50.0% |
| v2challenger | $59.95, line price | $5.85 | 9.1% | 1 | 11 | 50.0% |
The report's own notes
- Profit per visitor counts first orders in the current cost view, net of refunds; renewals are not in it yet.
- A visitor is exposed when the SDK renders the arm (a view_offer event); visitors who never saw the offer are not counted.
- The interval is an always-valid confidence sequence: reading the report every day does not inflate false positives. Stopping and promoting are your calls.
- The control had no profit history in the 30 days before the start, so no sample size could be planned; decisions rest on the interval alone.
“The SDK” here is the client the Verified offer theme block runs; a brand's own page can send the same view_offer event.
The captured figuresGET experiment and its arms' versions, dev API
{
"experiment": {
"id": "8189cf1b-3b0a-4689-9851-89db94fa7b85",
"offerId": "840c4c73-1abb-495b-91c9-f62eb65af1a9",
"status": "running",
"controlVersion": 1,
"arms": [
{
"weight": 0.5,
"version": 1
},
{
"weight": 0.5,
"version": 2
}
],
"falsePositiveRate": 0.05,
"minimumDetectableEffect": 0.1,
"plannedVisitorsPerArm": null,
"startedAt": "2026-09-29T13:05:58.005Z",
"pausedAt": null,
"endedAt": null,
"winnerVersion": null,
"report": {
"window": {
"start": "2026-09-29T13:05:58.005Z",
"end": "2026-10-03T01:30:14.562Z"
},
"currencyCode": "USD",
"arms": [
{
"version": 1,
"control": true,
"weight": 0.5,
"visitors": 13,
"buyers": 1,
"conversion": 0.07692307692307693,
"profitPerVisitorCents": 569.6923076923077,
"orders": 1,
"ordersWithoutVisitor": 0,
"revenueCents": 7795,
"refundedCents": 0,
"profitCents": 7406,
"averageOrderValueCents": 7795,
"refundPercent": 0,
"plannedVisitors": null
},
{
"version": 2,
"control": false,
"weight": 0.5,
"visitors": 11,
"buyers": 1,
"conversion": 0.09090909090909091,
"profitPerVisitorCents": 585,
"orders": 1,
"ordersWithoutVisitor": 0,
"revenueCents": 6795,
"refundedCents": 0,
"profitCents": 6435,
"averageOrderValueCents": 6795,
"refundPercent": 0,
"plannedVisitors": null
}
],
"comparisons": [
{
"version": 2,
"falsePositiveRate": 0.05,
"differenceCents": 15.307692307692264,
"lowerCents": -2361.1213497332824,
"upperCents": 2391.736734348667,
"radiusCents": 2376.4290420409748,
"relativeLift": 0.026870105320010725,
"decision": "keep_running"
}
],
"assumptions": [
"Profit per visitor counts first orders in the current cost view, net of refunds; renewals are not in it yet.",
"A visitor is exposed when the SDK renders the arm (a view_offer event); visitors who never saw the offer are not counted.",
"The interval is an always-valid confidence sequence: reading the report every day does not inflate false positives. Stopping and promoting are your calls.",
"The control had no profit history in the 30 days before the start, so no sample size could be planned; decisions rest on the interval alone."
]
}
},
"armCharges": [
{
"version": 1,
"firstChargeCents": 6995,
"catalogCents": 6995,
"purchase": "one_time",
"mechanism": "catalog_price"
},
{
"version": 2,
"firstChargeCents": 5995,
"catalogCents": 6995,
"purchase": "one_time",
"mechanism": "line_price"
}
]
}Price test, dev store, captured 2 October 2026. The dev store records no unit cost for this product, so profit here is revenue less fulfilment and payment processing.
What it does not do yet
- No automatic winner promotion: the report recommends, and the winner goes live when your agent publishes it with
publish_offer. - Stopping an experiment in the dashboard records a winner note only; it promotes nothing.
- One active experiment per offer, with at most three challengers.
- Price tests cut prices only; a line price can never be above the catalog price.
- Renewal orders are not in profit per visitor yet: the metric counts first orders.
- Orders whose cart carried no visitor id are counted apart from the metric.
- Offer Suite checks which lines a cart holds, not which arm the shopper saw, so a shopper who edits a cart's offer attribute to another arm's, on a cart holding that arm's lines, can get that arm's price. Every arm's price is one you published and verified.
Questions
Why not decide on conversion?
Because conversion can rise while profit falls. A deeper discount or a richer gift sells to more people and keeps less from each order, and a test judged on conversion would pick it. Profit per visitor counts both halves: how many visitors buy, and what each first order keeps after every cost and refund. The example report above shows the opposite case, a challenger that sells to fewer buyers and still earns more per visitor. Conversion stays in the report, beside the metric, but it does not decide.
Can I look at the results every day?
Yes. The interval is an always-valid confidence sequence, so reading the report daily or hourly does not inflate false positives. When the whole interval clears zero the report says so, even before the planned sample is reached.
What happens to shoppers when an arm breaks?
The guardrail pauses the experiment and every visitor is served the control version from their next request. Carts already built keep their prices. When the arm is fixed, published again and the experiment resumed, visitors return to the arms they had.
Does a price test need Shopify Plus?
No. It works on any plan where Offer Suite's cart transform is registered. Plus stores get the price through lineUpdate; on other plans the transform expands the line into itself at the fixed price. Either way the shopper sees the lower price as the product's own price.