Define What You're Actually Optimizing For Before You Start Testing

It's entirely possible to run a technically correct A/B test, get a clear winner, and still have improved the wrong thing — because nobody ever explicitly defined what "better" actually meant before the test began. A test compares two options and picks a winner, but it can't tell you whether you picked the right thing to measure in the first place.

Why "Which One Wins" Isn't the Same as "Which One Matters"

Every test implicitly relies on a specific metric — a definition of success the test is optimizing toward — and if that metric is vague, poorly chosen, or simply assumed rather than stated, the test can produce a confident, statistically valid answer to a question that doesn't actually matter to the business. A button color test that only measures click-through rate, for instance, can produce a clear winner that increases clicks while quietly decreasing actual completed purchases, if the underlying metric was never explicitly tied to revenue in the first place.

Defining your metric properly before testing should involve:

  • Stating explicitly what outcome you actually care about, not just what's easiest to measure
  • Connecting that outcome as closely as possible to real business value, like revenue or qualified leads
  • Being specific enough that two people looking at the same results would agree on the winner
  • Confirming the metric can't be "won" in a way that's technically true but practically meaningless

A Simple Framework

  1. Before designing any test, write down the specific outcome that actually matters to the business
  2. Choose a metric that's as close to that real outcome as you can practically measure
  3. Confirm the metric can't be gamed by a change that improves the number without improving the actual goal
  4. Only then design the test itself, using that metric as the deciding factor

> Tip: If you're tempted to measure something because it's easy to track rather than because it's what genuinely matters, that's usually a sign your metric needs more thought — click-through rate is easy to measure, but revenue per visitor is usually what you actually care about, and the two don't always move together.

Example

Before: A test measuring only click-through rate on a call-to-action button, declaring a winner that increased clicks but, unmeasured at the time, actually decreased completed purchases.

After: The same test redesigned to measure revenue per visitor directly, revealing that the "losing" version by click-through rate was actually the stronger performer where it genuinely mattered.

Common Mistakes

  • Measuring whatever is easiest to track instead of what actually reflects business value
  • Assuming a clear test winner automatically means the right thing was improved
  • Never explicitly stating the metric before designing and running the test
  • Choosing a metric that can technically improve without the underlying goal improving at all

Once you've defined a genuine metric, tracking how it actually trends over the course of a test is essential to trusting the result. SeoWolf's Cohort Tool can help visualize that trend against the metric that actually matters.


A test can only ever tell you which option won by the measure you gave it — the real work happens before the test even starts, in deciding honestly what "winning" should actually mean.