Incrementality Testing for Restaurant Promotions

The restaurant industry has long had a complicated relationship with discounts. They feel like generosity. They function as necessity. And they get measured like science, even when the methodology is closer to astrology.
The problem isn't that operators are running promotions. It's that they're reading the results as though the scoreboard is accurate, and in most cases, it isn't. Someone redeemed your offer, sales climbed during the window, and your attribution report flagged a win. But would those customers have come anyway? Would you have captured that revenue without spending a dollar? These aren't rhetorical questions. They are, in fact, the only questions that matter, and most promotional measurement systems don't bother asking them, which is precisely the gap that incrementality testing is designed to close.
Consider where the industry actually sits right now. According to 2024 survey data from the National Restaurant Association, 61% of restaurant operators offer loyalty programs, and 76% of limited-service restaurants reported measurable traffic increases attributed to those programs. A separate 2024 survey found that 51% of U.S. diners said a coupon or discount would convince them to try a restaurant they hadn't visited before; 30% said they'd refuse to try a place that offers no deals at all.
That's substantial demand pressure toward discounting, and operators have responded accordingly. But here's where the math gets uncomfortable: while airline and retail loyalty programs typically return around 1% of spending back to the customer, restaurant programs frequently return 10% or more, often to the most price-sensitive segment of the customer base. That gap is not sustainable at thin margins, particularly when the customers you're buying with those offers aren't the ones you're building a business around. A 2026 Phygital Index Report found that 45% of diners said their favorite restaurant chain had changed in the past year, up from 33% the year prior. If promotions are supposed to build loyalty, that number is something close to an indictment.
High promotional spend creates pressure to justify high promotional spend. Marketers optimize toward the metrics they can access. When those metrics conflate activity with causation, the optimization loop becomes self-reinforcing and quietly destructive. An estimated 172.6 million American consumers redeemed digital coupons in 2025, and digital redemption is individually trackable in ways paper coupons and broadcast media rarely were. Worth asking, though: is that infrastructure being used to measure what actually happened, or merely to document what was observed?
What Incrementality Testing Actually Measures, and Why the Definition Is Doing Real Work
The test answers one question: what portion of observed sales would not have occurred without the promotion?
Not "did sales go up?" Not "did customers respond?" Those questions have obvious answers and nearly useless implications. The operative word is without. The test measures the counterfactual, the parallel version of events where the promotion was never sent, sometimes called the potential outcome in causal inference literature. The gap between that counterfactual and actual results is the incremental lift. Stated plainly: Incremental Lift equals Treatment Group Results minus Control Group Results, divided by Control Group Results.
This is deliberately different from an A/B test, which compares two versions of something against each other. An A/B test tells you whether version A outperforms version B. An incrementality test tells you whether doing anything at all was better than doing nothing. Those are different questions, and the answer to one doesn't answer the other.
The distinction from attribution matters just as much. Attribution tells you who touched what before converting. Last-click attribution, also called last-touch attribution, assigns full credit to whatever touchpoint preceded a conversion. Incrementality testing tells you whether that touch caused the conversion. A customer who was going to visit on Friday, received your Thursday push notification, and visited on Friday will appear in an attribution report as a validated conversion. In an incrementality test, they're correctly identified as someone who would have come regardless. That's a meaningfully different business outcome, and confusing the two is how promotional budgets quietly expand year over year with nobody quite understanding why.
One boundary worth marking clearly: incrementality testing measures net lift. It doesn't, on its own, explain the mechanism behind that lift. Decomposing how much came from new customers versus menu switching versus time-shifted demand requires additional analysis layered on top. The test tells you the size of the effect. Decomposition tells you its origin. Don't ask the test to do both jobs at once.
The Three Experimental Designs Available to Restaurant Marketers
There is no single correct design. The right choice depends on what you're measuring, what channels you're working in, and honestly, how much operational disruption you're willing to absorb during the test window.
Customer-Level Holdout
This is the cleanest design when individual identity is known. Within a loyalty program, app, email list, or SMS audience, you randomly withhold the promotion from a defined subset of eligible customers. That withheld group receives nothing. Their behavior during the promotion window establishes your counterfactual.
A holdout of roughly 5 to 10% of the eligible audience typically provides enough statistical power without forgoing so much potential revenue that internal stakeholders start asking uncomfortable questions mid-test. Individual-level randomization minimizes contamination between groups; you know who received the offer and who didn't, which makes the comparison as clean as real-world conditions allow. This design is best suited for digital promotions where identity is already tracked.
Geo-Based Holdout
When individual identity isn't available, which is the case for broadcast media, out-of-home campaigns, and regional limited-time offers, you work at the market level. A matched set of geographic markets serves as the treatment group; another matched set runs dark, meaning the promotion doesn't appear there at all.
The critical phrase is "matched set." Demographic similarity between markets is a starting point, not a conclusion. Markets that look alike on population data can diverge sharply in category behavior, historical promotional response, and baseline conversion rates. Getting the matching wrong means your control group represents a different kind of restaurant market entirely.
The documented risk with geo-based holdouts is spillover. Customers cross market boundaries. Word of mouth travels. Social posts don't respect designated marketing areas. When contamination reaches control markets, the counterfactual degrades, and the test becomes directional at best.
Statistical Retrospective Matching
This is the method you use when someone forgot to build a holdout before the campaign launched. It happens more than anyone admits.
Post-hoc matching algorithms identify customers or markets from historical data that most closely resemble the promoted population and construct a synthetic control group, a technique sometimes formalized as the synthetic control method in causal econometrics. Less precise than a pre-formed holdout, but capable of extracting directional signal from campaigns already completed. The ceiling on accuracy is set by the quality of historical data and the matching methodology; unobserved differences between groups can introduce bias that no statistical technique fully removes.
Use this design when a pre-planned test wasn't feasible. Be transparent about its limitations before presenting results to anyone making consequential budget decisions based on them. Presenting retrospective matching results with the same confidence as a pre-planned holdout is the kind of thing that gets decisions made on bad information.
Building the Baseline: The Calculation That Everything Else Depends On
Why does the baseline matter so much? Because a flawed baseline makes nearly every redemption look incremental. And if every redemption looks incremental, you haven't solved the attribution problem; you've just dressed it up more convincingly.
The baseline is not the prior-period average. Comparing promotion week to the week before ignores seasonality entirely and misses the post-promotion dip where demand pulled forward from future weeks will eventually surface, producing a baseline that overstates true incremental demand. The baseline is not the non-promoted store average unless those stores were deliberately matched before the test began. Unmatched comparisons compound small pre-existing differences into results that appear meaningful but reflect composition error.
A defensible baseline requires a control group matched on geography, store format, historical sales velocity, and any external variable expected to affect the test period: a regional event, a competitive opening, a weather pattern.
Analytical approaches span a real spectrum here. Simple moving averages are fast and directional but can't distinguish trend from noise. Regression models account for seasonality and day-of-week patterns and represent a reasonable standard for most chain-level analyses. Machine learning forecasts incorporating price elasticity, weather, local events, and concurrent marketing activity are appropriate for large chains with sufficient historical data to train and validate a predictive model.
But the practical failure mode is similar at every level of sophistication. Overclaiming the baseline understates what would have happened anyway, which mechanically inflates the apparent lift. This is how a promotion that generated genuinely new visits in the single digits gets reported as a 30% lift event. The math checks out. The conclusion doesn't. I've sat in rooms where that exact scenario was used to justify doubling the promotional budget the following quarter.
The Four Distortions That Quietly Erode True Lift
Even a well-constructed test can misrepresent performance if it measures only the top-line lift number without examining what's happening underneath. The honest formula for overall incremental lift is: Lift in Promoted Item Sales, plus Halo Effects, minus Cannibalization, minus Pull-Forward. Ignore any one of those components and the reported number is misleading by precisely whatever was omitted.
Subsidized Loyal Customers
Your most frequent guests are also your most likely redemption cohort. They open your app, respond to push notifications, engage with your loyalty program. They were coming anyway. When they redeem a discount on a visit they would have made at full price, you haven't changed their behavior; you've reduced your margin on it.
This surfaces cleanly in the data when you segment redemption by pre-promotion visit frequency. If the heaviest redeemers are concentrated in your top-quintile visit cohort and that cohort shows no increase in visit rate during the promotion window, you're looking at deadweight loss, not lift.
Pull-Forward and the Post-Promotion Dip
A customer who would have purchased at full price next week buys this week on discount. The promotion window looks flattering. The following week develops a trough. Research on multi-year retail promotional data finds that a substantial portion of promotional revenue in many categories represents cannibalized future demand rather than net new demand. The mechanism applies directly to restaurant visit behavior, where occasions have natural spacing that promotions can compress but not eliminate.
The fix is straightforward: extend the measurement window past the promotion end date and treat the post-period with the same analytical rigor as the promotion period itself. A test that stops at the closing date is answering a narrower question than the one that actually matters.
Within-Menu Cannibalization
A discounted limited-time offer drives units on the promoted item while full-price equivalents in the same category quietly decline. Promoted-item lift looks strong; portfolio revenue is flat or negative. The honest metric is total category lift, or better, total portfolio margin impact.
This distortion is underappreciated because it requires looking at what didn't change rather than what did. The spike in promoted-item units is visible; the erosion of full-price equivalent orders requires specifically measuring the category around it. Most promotional reports never bother.
Halo Effects
The positive version of the same logic: a beverage promotion drives incremental food attach, or a dessert limited-time offer increases main-course check attachment. The promoted item looks marginal in isolation, gets cut, and so does the halo lift nobody measured.
Both cannibalization and halo effects argue for the same discipline: define the portfolio before the test begins and track it throughout. This sounds obvious. It is frequently skipped.
Design Mistakes That Invalidate Results Before the Data Is Collected
One argument holds that a flawed test is better than no test. I've heard it often. I don't buy it. A flawed test produces a number that looks authoritative and can't be corrected after the fact. No test produces a gap that at least everyone acknowledges. The flawed test is usually worse precisely because it forecloses the honest conversation.
Running tests too short is the most common error and the most resistant to correction, because early results often look clean and convincing exactly when the noise hasn't had time to develop into visible variance. Full purchase cycles, delayed behavioral responses, and external fluctuations all need time to surface.
Contaminated control groups are the second invalidating error and are frequently undetectable after the fact. In digital campaigns, the same user can appear in both treatment and holdout groups through platform identity resolution gaps. In geo-based tests, customers cross market boundaries or encounter the promotion through social channels. Once contamination occurs, no statistical adjustment fully restores the integrity of the counterfactual. You're estimating the damage, not correcting it.
Testing multiple variables simultaneously, creative and offer value at the same time, or launching an additional channel mid-test, makes it impossible to attribute observed lift to any specific cause. The result tells you something changed. It cannot tell you what. This is self-evident, yet it happens regularly when promotional calendars collide with test windows.
Then there's stopping measurement at the promotion end date. This is a structural choice that overstates promotional effectiveness, and in most measurement systems, it's the default.
A lift figure without confidence intervals is not a result. It is a number. Sample size and test duration required to achieve meaningful statistical power must be calculated and committed to before the test launches, not adjusted retroactively when early readings look disappointing.
What to Measure Beyond Sales Lift: Profitability, Retention, and Long-Term Customer Value
Sales lift is not the goal. It's a leading indicator, and a potentially misleading one, of whether the promotion was actually worth running.
The profitability formula is direct: Promotion Profitability equals Incremental Sales multiplied by Unit Margin, minus Trade Spend. This figure can be negative even when incremental lift is positive and statistically significant. A promotion that generates real new visits at margins that fail to cover the promotional cost is not a success. It is a well-measured failure. That's more useful than an unmeasured one, but it's still a failure.
Retention complicates the picture further. Customers acquired during promotional periods frequently exhibit lower purchase frequency and lower retention than cohorts acquired outside promotional windows. A promotion can appear profitable within the test window while generating a customer cohort whose lifetime value is materially below average. The measurement required to surface this: track purchase frequency, retention rate, and customer lifetime value (CLV) for the promo-acquired cohort against a matched baseline cohort at 30, 60, and 90 days post-acquisition. Most operators who started tracking this found that their promo-acquired cohort underperformed on repeat visits. Most restaurants also can't explain why their repeat visit numbers keep disappointing them.
Incremental ROAS, or iROAS, deserves attention as an operational metric precisely because standard ROAS doesn't. Standard ROAS includes conversions that would have happened without the promotion; incremental ROAS strips those out. A campaign with strong standard ROAS and weak incremental ROAS is efficiently capturing organic demand. That's not the same as generating it. Scaling that campaign doesn't grow the business; it subsidizes behavior that would have occurred anyway, at increasing cost.
Recall the churn data: 45% of diners changed their favorite restaurant chain in the past year, up from 33% the year prior. Promotions that generate visits but fail to build genuine preference don't interrupt that cycle. They accelerate it, training customers to engage only when a discount is present. The measurement system has to account for whether the promotion changed the customer relationship, not just whether it changed the transaction count.
Translating Test Results Into Decisions: What to Keep, Scale, or Kill
A test result without a decision framework is, at best, an interesting document. At worst, it's ammunition for whoever argues loudest in the room.
High incremental lift, positive promotion profitability, and strong post-promotion retention: this is a scalable promotion. Expand it to similar customer segments or comparable markets. Identify which segment characteristics correlate most strongly with lift and bias future targeting toward them.
High redemption rate with low or zero incremental lift means the promotion is subsidizing existing behavior. The structural fix is to restructure targeting to exclude high-frequency guests who were coming regardless, or to redesign the offer trigger so the reward attaches to a new behavior rather than a habitual one.
Positive short-term lift followed by a post-promotion dip that effectively erases it: pull-forward is the dominant mechanism. Space promotions further apart, or shift to frequency-based reward structures that don't concentrate demand in a single window.
Strong promoted-item lift with negative portfolio lift means cannibalization is real and the promoted-item view is misleading you. Reprice the promoted item, bundle it with complementary full-price items, or restrict the discount to configurations that don't compete directly with full-price equivalents.
That raises an important question about the broader picture. If these decisions are available to any restaurant marketer willing to design the test correctly, why isn't this standard practice? Part of the answer is organizational: the people who run promotions are often not the people who measure them, and the measurement system rarely forces a reconciliation between what was expected and what actually happened. Part of the answer is incentive structure: a high redemption number is easy to celebrate in a quarterly review, even when it reflects activity that cost more than it created.
The more honest answer, though, is that a legitimate incrementality test requires deliberately withholding a promotion from real customers, accepting that the holdout group's behavior represents a necessary cost of knowing the truth, and committing to a measurement window that extends past the moment when promotional excitement is still generating internal goodwill. Most promotional evaluation processes were never built to accommodate any of that. They were built to confirm, not to interrogate.
In an environment of compressing margins, eroding brand loyalty, and an increasingly deal-conditioned consumer base, the cost of mismeasurement compounds every cycle. Each flawed test produces a flawed decision. Each flawed decision sets the parameters for next year's promotional calendar. At some point the gap between what you think your promotions are doing and what they are actually doing stops being a measurement problem and starts being a business one. Those are harder to fix.


