Back to blog

Azulens Blog

How We Approach Incrementality Testing at Azulens

Incrementality testing sounds straightforward until you try to apply it across very different businesses. Some advertisers have dozens of measurable markets and years of history. Others have only a few regions, short campaign windows, retailer sales, marketplace data, or no geographic conversion data at all. At Azulens, we start with the decision the advertiser needs to make, then determine what can actually be measured and manipulated. From there, we design the experiment around the client rather than forcing the client into one testing methodology.

Different measurement situations flow through a central testing framework into geographic, rotating-market, or time-based experiment designs, all leading to an incremental business impact measurement.
Different data and measurement situations require different experiment designs, but they ultimately serve the same goal: measuring incremental business impact.

Incrementality testing has a deceptively simple question at its core:

What would have happened if we had not spent this media budget?

Answering that question well is much harder.

A retailer might have daily revenue for dozens of markets. A movie launch may have only a few weeks to generate demand. A business selling through Amazon may have limited access to conversion data. Another advertiser may have excellent historical sales data but only a handful of regions that can be targeted independently.

All of them want some version of the same answer: Did the advertising create enough additional business to justify the spend?

The way we get there can be very different.

That is why our approach at Azulens starts with the measurement situation rather than with a particular statistical method.

We first ask what can actually be measured

Before thinking about treatment groups, control groups or statistical models, we need to understand what the business can actually observe and change.

Can the final business outcome be measured geographically? At what level? Can advertising be changed independently in those same markets? How much historical data exists? How quickly should advertising affect the outcome? Can some markets receive no advertising at all, or can we only create a difference between higher and lower spend?

And most importantly, what decision does the advertiser want the experiment to support?

That last question has a major influence on the design.

Suppose a company would only scale a channel if incremental ROAS is above 1.0. Before spending money on the experiment, we want to know whether the proposed design has enough sensitivity to distinguish performance around that level.

Simply being able to run a test is not enough. The test should have a realistic chance of answering the business question behind it.

A strong geo setup gives us the most options

Some businesses are naturally well suited to geographic experiments.

Imagine a company with many independently targetable markets, stable sales patterns and a long history of daily revenue by geography. Advertising can be active in one group of markets and withheld in another.

In a situation like this, a geographic holdout can be a strong option.

Historical sales help us identify markets or groups of markets that have behaved similarly before the experiment. We can then examine how accurately the untreated markets would have predicted the treatment markets during historical periods.

This distinction matters because raw business volatility tells us surprisingly little about the actual sensitivity of an experiment.

A business might see revenue move 20 percent from one week to another. That does not mean advertising needs to create a 20 percent lift before we can measure it.

Much of that movement may be shared across markets.

If the control markets explain most of those changes, the unexplained treatment versus control difference can be much smaller than the volatility visible in the raw sales numbers.

For test planning, that residual uncertainty matters far more than the headline volatility of the business.

Three volatile market revenue series share similar movements, allowing control markets to explain much of the variation and leaving a smaller residual difference for measuring incremental lift.
High volatility in raw sales does not necessarily mean a test can only detect equally large effects; what matters is the unexplained difference between treatment and control.

Sometimes fewer markets can produce a better experiment

More geographic markets usually give an experiment more independent information.

But media budget is finite.

If a fixed budget is spread across too many large markets, the additional advertising pressure may become so small relative to existing demand that the resulting sales effect becomes difficult to identify.

Concentrating the experiment into a smaller set of suitable markets can sometimes create a stronger treatment signal.

This means test design becomes a balancing exercise. We look at the number of independent markets, historical predictability, baseline revenue, available media budget, treatment intensity and the amount of time available.

The largest possible geographic test is therefore not automatically the best one.

The useful design is the one that gives the advertiser enough sensitivity to answer the decision they actually care about.

A fixed media budget spread across many markets creates relatively weak treatment intensity, while concentrating the same budget in fewer suitable markets can create a stronger measurable signal.
More markets can provide more information, but with a fixed budget, concentrating treatment in fewer suitable markets can sometimes improve test sensitivity.

Few large regions require a different approach

Sometimes the outcome data are excellent while geography is limited.

A company operating across eight large regions has a very different testing problem from a retailer with fifty independent local markets.

With fewer regions, an unusual event in one market can have much more influence on the result. Historical matching and pre-test simulation therefore become especially important.

When there are only a handful of large, stable markets, repeating treatment changes over time can sometimes add useful information.

Instead of permanently assigning one location to treatment and another to control, media pressure can rotate across markets and periods according to a predefined schedule.

One market might receive higher pressure during one period, another during the next, followed by further planned rotations.

This allows the same market to be observed under different treatment conditions.

Whether that design makes sense depends heavily on how quickly advertising affects the outcome and how long the effect persists. If advertising in one period continues to influence sales several periods later, the next treatment condition can become contaminated.

In those situations, longer blocks, washout periods or a different experiment structure may be necessary.

Product launches come with a different kind of uncertainty

A new movie, game or product launch creates another challenge.

There may be very little history for the exact thing being advertised.

Previous launches can still tell us something about geographic sales behavior and likely volatility, but we should not treat them as if they provided the same level of certainty as years of stable sales for the same business.

For these cases, we rely more heavily on scenarios.

We can examine what happens under different assumptions about launch volatility, different market combinations, different campaign durations and different levels of media concentration.

This allows us to understand whether a test remains useful across a reasonable range of outcomes instead of relying on one overly precise forecast.

That distinction is important when setting expectations before launch.

A weaker historical basis does not necessarily mean that a good experiment cannot be run. It means we should be more careful about how precisely we describe its expected sensitivity in advance.

Geographic testing can help businesses with imperfect attribution

Many of the companies that benefit most from incrementality testing do not have clean platform level attribution.

They may sell through retailers, close customers over the phone, operate physical locations or sell through marketplaces.

The important question is whether the final business outcome can be connected to a geographic unit that can also be manipulated through media.

For a brand selling through retailers, the outcome might be regional sell through.

For a multi location business, it might be store or branch revenue.

For a company closing leads over the phone, it could be closed revenue assigned back to a geographic sales territory.

The sales channel itself does not determine whether a geographic incrementality test is possible. What matters is whether the outcome and the media treatment can be aligned at a useful geographic level.

When geography is unavailable, time becomes part of the experiment

Some advertisers cannot obtain geographic outcome data at all.

They may only have a national revenue series.

In that situation, we cannot create geographic treatment and control groups. If media pressure can be changed over time, the experiment can instead use controlled time periods.

Advertising might alternate between higher and lower intensity or between active and inactive periods.

These tests need to account for the fact that observations over time are connected.

Monday is related to Tuesday. December behaves differently from February. Promotions, holidays and external events can affect the whole business at once.

Conversion delays and advertising carryover can complicate the picture further.

For that reason, we treat time based experimentation as its own measurement problem rather than pretending that dates are simply another version of geographic markets.

We want to know whether the test is useful before it starts

One of the most important parts of our approach happens before any experiment goes live.

Where the available history allows it, we simulate how the proposed design would have behaved in the past.

We can first run placebo experiments during historical periods where no treatment occurred. A good design should normally produce estimated effects close to zero in those periods.

We can then introduce artificial treatment effects of different sizes and analyze them using the same method we intend to use for the real experiment.

This gives us an empirical view of the sensitivity of the proposed design.

One useful planning metric is the Minimum Detectable Effect, or MDE. It tells us approximately how large a true effect would need to be for the proposed experiment to detect it with a chosen level of statistical power.

The next step is connecting that statistical sensitivity to the advertiser's economics.

Suppose the planned treatment markets would generate €500,000 in revenue without the additional advertising.

The advertiser plans to invest another €30,000 and considers an incremental ROAS of 1.0 the minimum acceptable result.

That means the business needs the additional spend to generate €30,000 in incremental revenue. Relative to the €500,000 treatment baseline, this corresponds to a 6 percent commercially relevant lift.

If historical simulations indicate that the proposed experiment only has adequate power around a 10 percent lift, we learn something important before launching it.

The current test design is probably not sensitive enough to answer the advertiser's decision at the desired level of confidence.

Historical data is used for placebo tests and simulated lift scenarios to estimate test sensitivity, which is then compared with the commercially required effect before deciding whether to run or redesign the experiment.
Pre-test simulation connects statistical sensitivity with business economics: if a test can reliably detect about a 10% lift but the business decision depends on detecting 6%, the design may need to change.

We can then explore alternatives.

Maybe the media budget should be concentrated into fewer markets. Perhaps the experiment should run longer. A stronger difference between treatment and control may be possible. Another geographic combination may produce a more predictable counterfactual.

We can simulate those alternatives before committing the media budget.

And if the available geography, budget and timing still cannot produce a useful experiment, recommending against the test can be the right outcome.

The methodology can change while the business questions stay consistent

Different situations require different statistical approaches underneath.

A large geographic holdout, a test involving only a few regions and a national time based experiment should not all be analyzed in the same way.

From the advertiser's perspective, however, the final questions remain remarkably consistent.

How much incremental revenue did the treatment create?

What was the incremental ROAS?

How uncertain is that estimate?

How does the result compare with the business threshold agreed on before the test?

And what decision should follow?

This consistency matters to us at Azulens.

The complexity behind experiment design should not force advertisers to reinvent the measurement process every time their data situation changes.

Why we built Azulens around this process

In practice, incrementality testing often involves several disconnected pieces.

Historical sales may live in one system while media performance lives in another. Test planning happens in spreadsheets, notebooks or specialist statistical tools. Campaign execution needs to be checked while the test is running. Once the experiment finishes, the statistical result still needs to be translated back into a business decision.

We want those pieces to belong to the same measurement process.

With Azulens, the goal is to connect outcome data, media context and business economics so that an experiment can be evaluated before launch and interpreted against the same assumptions afterwards.

The exact methodology can adapt to the advertiser's situation.

The decision process stays consistent.

We start with what can be measured and what can realistically be manipulated. We determine the business effect that needs to be measurable. We design the experiment around those constraints and evaluate its expected sensitivity before the budget is committed.

And when the available data, geography, timing or budget cannot support a useful answer, the right recommendation may be to redesign the experiment or not run it yet.

For us, that is what makes incrementality testing genuinely useful.

Keep exploring

More from the Azulens blog

View all articles →