Choose Between Conversion Rate, Revenue per User and Capped Revenue Before You Launch the Test

Choose Between Conversion Rate, Revenue per User and Capped Revenue Before You Launch the Test

Most e-commerce A/B tests get declared on whichever metric moved. That is how a team ships a variant that lifted conversion rate by 4% and lost money. The fix happens before launch: pick one primary metric, write down its cap and its minimum detectable effect, and treat everything else as a guardrail. Metric choice is not cosmetic. Microsoft's experimentation team measured revenue per user as their most skewed metric, and showed that capping it dropped the minimum users per arm from roughly 114,000 to 9,700 for the test to stay statistically well behaved (Kohavi et al., KDD 2014). This guide compares the five metrics teams actually argue about and gives you the arithmetic to choose between them.

The five candidates, and what each one really measures

User conversion rate. The share of users in the arm who placed at least one order. Binary, bounded, low variance. Blind to basket size, so it rewards anything that pulls cheap orders forward.

Revenue per user (RPU). Revenue in the arm divided by users in the arm, zeros included. Closest to the business, and most vulnerable to a single enormous order.

Capped revenue per user. The same metric with each user's revenue truncated at a ceiling chosen before launch. Trades a little bias for a large drop in variance.

Average order value (AOV). Revenue divided by orders. Never a primary metric in a user-randomized test: its denominator is an outcome of the test, so a variant that wins more small orders shows falling AOV while making more money.

Revenue per session. Another ratio whose denominator is not the randomization unit. Randomize users, analyse sessions, and a naive t-test gets the variance wrong: you need the delta method or a cluster-aware estimator (Deng, Knoblich and Lu, KDD 2018).

Compare them on six criteria

CriterionUser CVRRPUCapped RPUAOVRevenue / session
Denominator is the randomization unitYesYesYesNoNo
Traffic neededPredictableDepends on skewLowLowMedium
Sensitive to basket sizeNoYesPartlyYesYes
Survives one huge orderYesNoYesNoNo
Defensible to financeNoYesAlmostNoAlmost
Safe as a primary metricYesWith sizingYesNeverOnly with delta method

Work out whether revenue per user actually costs you traffic

The folklore says revenue metrics always need far more traffic than conversion rate. That is not a rule but a consequence of one number: the coefficient of variation (CV) of revenue per user, its standard deviation divided by its mean, across all users including non-buyers.

For 80% power and a 5% two-sided test, users needed per arm to detect the same relative effect is roughly 16·p(1−p)/(p·r)² for a binary rate and 16·CV²/r² for a continuous metric. Set them equal and the crossover is CV = √((1−p)/p):

Your user conversion rateCV of revenue per user below which RPU needs no extra traffic
1.0%9.9
2.0%7.0
2.4%6.4
3.0%5.7
5.0%4.4

Worked example, with illustrative figures rather than client data. A store converts 2.4% of users, mean revenue per user is 24 currency units, and 90 days of orders give a standard deviation of 240, so CV is 10. Detecting a 10% relative lift needs about 65,000 users per arm on conversion rate and about 160,000 on raw RPU. Cap the metric at the 99th percentile of order value and CV typically falls into the 4 to 6 range; at CV 6 the requirement is about 58,000, now cheaper than conversion rate.

A four-step procedure

  1. Name the decision, not the metric. Write the sentence "if this wins we roll it out to all traffic and expect X". The metric is whatever makes that sentence checkable.
  2. Eliminate ratios whose denominator the test can move. AOV goes to the guardrail list immediately. Revenue per session stays only if your tool does cluster-aware variance.
  3. Compute the CV and read the table. Under the threshold, use RPU. Over it, cap and use capped RPU. Still over, fall back to user conversion rate and put RPU on the guardrail list.
  4. Freeze the cap, the MDE and the runtime in writing before a single user is exposed. A cap chosen after you have seen the data is not a cap, it is an outcome.

How to set the cap

Take the 99th percentile of per-user revenue over a window at least as long as the planned test, covering the same seasonality, and record it in the test document. Report capped RPU as the primary result and uncapped RPU alongside it. If the two disagree in direction, a handful of very large orders is driving it: a finding to investigate, not a result to ship.

What changes in a high-inflation market

Turkey ran 31.51% annual and 1.84% monthly consumer price inflation in August 2026 (Presidency of Strategy and Budget, 4 September 2026, on TÜİK data). Teams here often assume this invalidates revenue-based test metrics. It does not, and the correction matters:

  • It does not bias the comparison. Both arms live through the same price changes at the same time. Concurrent control is exactly what protects you.
  • It does break your business case. Convert the lift to a percentage and apply it to a forward revenue plan. Never multiply a four-week absolute currency delta by 13.
  • It does break historical baselines. A CV computed on last year's orders is wrong for sizing today's test. Recompute on a recent window.
  • It does interact with your cap. A fixed currency cap set in January is a much tighter cap by December. Re-derive the percentile each quarter.

The same logic applies in any market with fast-moving prices or a volatile currency.

Guardrails: what you check but never declare on

A guardrail can veto a win but cannot create one. Declare on one primary metric and keep this list short, because every extra metric you would act on inflates your false-positive rate (Kohavi, Deng and Vermeer, KDD 2022). A workable default set:

  • Uncapped revenue per user, as the reality check on the cap
  • Orders per user, to catch a lift that is only basket-splitting
  • Return or cancellation rate on a lag window, because a checkout change that pushes bad orders through looks like a win for two weeks
  • Core Web Vitals on the changed template
  • Sample ratio, checked daily; a mismatch invalidates everything above

Where this breaks down

This procedure assumes user-level randomization, one test at a time on the affected surface, and a purchase that completes inside the measurement window. It gives nothing useful when:

  • The purchase cycle is longer than the test. Travel and high-consideration retail book weeks after the session. Use a leading indicator with a validated relationship to revenue, and accept that you are testing the indicator instead.
  • Traffic is below the threshold at any cap. Under roughly 10,000 users per arm per week, moderated usability work and a UX audit will tell you more than an underpowered test.
  • The change targets a subgroup. A metric measured on all users dilutes an effect confined to, say, instalment payers. Pre-register the segment as the population rather than slicing afterwards.
  • Revenue is recognised differently from what the tag fires. If GA4 purchase revenue and your order table disagree, fix that first.
  • Your conversion counting method is not what you think. GA4 counts a key event either once per event or once per session depending on how it was created, so the same behaviour yields five conversions or one (Google Analytics Help).

The CV table is a triage tool for choosing between metrics, not a replacement for a power calculation. Size the test on the metric you chose, with your own variance estimate.

Frequently asked questions

Marked for FAQPage schema.

Can I use two primary metrics if I require both to win?

You can, and it makes the test more conservative rather than less. Requiring both to clear significance lowers your power, so size for the harder one. What you cannot do is declare on whichever of the two happened to win.

Does capping revenue bias the result?

Yes, deliberately. A capped metric measures typical-order revenue rather than total revenue. That bias is the price of a usable variance, and it is acceptable because it applies identically to both arms. Report the uncapped number alongside it so the trade stays visible.

What percentile should the cap sit at?

The 99th percentile of per-user revenue is a reasonable default. If CV is still above the threshold there, try 97.5. Below roughly the 95th percentile you start discarding ordinary large baskets rather than outliers.

Why is average order value such a bad primary metric?

Because the test changes how many orders exist. A variant that converts more hesitant, lower-basket shoppers raises revenue and lowers AOV at the same time. AOV is a diagnostic for interpreting a revenue move, never the thing you decide on.

Can I switch the primary metric after the test starts?

No. Choosing the metric after seeing the data is the same error as stopping the test after seeing the data, and it inflates false positives the same way. If the chosen metric turns out to be wrong, record that, finish the run, and treat the result as exploratory.

How do I handle multiple currencies?

Convert every order to one reporting currency at a single rate frozen at test start, not at the daily rate. A moving rate injects variance that has nothing to do with your variant, and it hits arms unequally when their traffic mix differs.

What if conversion rate goes up and capped revenue per user goes down?

That is the most informative outcome in this guide. The variant wins more orders of lower value. Whether to ship depends on margin and repeat behaviour, so escalate it as a business decision rather than resolving it statistically.

Run this on your next test

Pull 90 days of per-user revenue, compute the coefficient of variation, and check it against the table above before you write your next test brief. If you would rather have metric definitions, caps and guardrails set up once and applied consistently across a whole experiment programme, that is what our conversion rate optimization service does.

Sources


Inci Dindar
Written by

Inci Dindar

With a background in Software Development, she led design, UX, and product teams across fintech, media, e-commerce, and travel. For the past three years she has consulted full-time on UX and design — including as an external UX consultant for a global management consultancy helping brands and startups with audits, conversion optimization, and design processes.


Related Articles

Switas As Seen On

Magnify: Scaling Influencer Marketing with Engin Yurtdakul

Check Out Our Microsoft Clarity Case Study

We highlighted Microsoft Clarity as a product built with practical, real-world use cases in mind by real product people who understand the challenges companies like Switas face. Features such as rage clicks and JavaScript error tracking proved invaluable in identifying user frustrations and technical issues, enabling targeted improvements that directly impacted user experience and conversion rates.