Media Mix Modeling
Media mix modeling (MMM) is a statistical method that estimates how each marketing channel contributes to sales or conversions by analyzing aggregate spend and outcome data over time.
Key takeaways
- Adstock models carryover; saturation curves model the diminishing return on extra spend.
- Coefficients are only identifiable if channel spend actually varied across the history.
- The output splits the outcome into a baseline plus a per-channel contribution.
- Marginal return at current spend, not average return, guides the next budget decision.
- Aggregated inputs mean MMM cannot evaluate individual creatives, keywords or audiences.
In depth
An MMM regresses a business outcome, usually weekly sales or qualified leads, against columns of historical data: spend per channel, price, promotions, distribution, seasonality and outside factors such as weather or competitor activity. Two transformations do most of the work. Adstock spreads a week's spend forward to represent advertising that keeps working after it runs, and a saturation curve bends each channel's contribution so that additional spend adds progressively less. The fitted coefficients then split the outcome into a baseline and a per-channel contribution.
Model quality is set by variation in the inputs. If a channel's budget has been flat for two years, the model has nothing to learn from and its coefficient will be unstable no matter how much data you feed it. Deliberate variation, staggered launches and regional differences all improve identifiability. The trade-off runs the other way too: more variables raise the risk that correlated channels split credit arbitrarily, so modellers constrain coefficients with priors or merge channels that always move together.
Teams use the fitted response curves for planning rather than for reporting. Reading the marginal return at current spend for each channel shows where the next unit of budget earns most, which turns a quarterly planning argument into an arithmetic one. Picking the modelled outcome is a judgement call: modelling raw form fills lets cheap traffic dominate, while modelling qualified leads scored by a scorecard funnel ties the coefficients to prospects who actually fit, so the curves guide spend toward pipeline instead of volume.
MMM is coarse by construction. It works on weekly or monthly aggregates, so it cannot say which creative, audience or keyword worked, and it will not answer a question about a two-week test. It needs a long history, typically two or three years of consistent data, which rules it out for young companies and for channels launched last quarter. And because it is correlational, an omitted variable such as a price cut or a distribution change will be absorbed into whichever channel happens to correlate with it.
Example in practice
How to measure it
Judge the model before you judge the channels. Hold back the most recent months, predict them, and compare predicted to actual outcomes; a model that cannot reproduce a period it never saw should not set budgets. Check that the residuals show no pattern over time, and that the baseline share, the portion of outcomes not explained by media, is plausible for a brand of your age.
Then read the channel outputs as two separate numbers. Contribution is the share of outcomes the model assigns to a channel over the period; marginal return on ad spend is the extra outcome from the next unit of budget at today's level. A channel can be large in contribution and poor at the margin, which is exactly the case for a saturated one. Plan on the margin, report on contribution.
Common mistakes
The first mistake is feeding the model a decorative spend history. When every channel has run at a constant weekly budget, the regression cannot separate them and produces confident-looking coefficients built on almost no information. Before commissioning a model, deliberately vary budgets by region or by week for a period, and record every promotion, price change and site outage, because unrecorded events end up attributed to whichever channel was live at the time.
The second is reading the model as a verdict rather than a prior. Teams cut a channel because its coefficient came back low, without checking whether that channel had any usable variation or whether its effect lands outside the modelling window. Use the model to generate ranked hypotheses, then confirm the two or three biggest reallocation decisions with a holdout test before moving budget, and refit the model each quarter as new weeks arrive.
Frequently asked questions
How is media mix modeling different from attribution?
Attribution tracks individual user journeys and credits touchpoints at the click level, while MMM works on aggregated spend and outcome data over time. MMM is privacy-resilient and captures offline and brand effects that click-based attribution misses.
How much data does media mix modeling need?
As a rule of thumb, two to three years of weekly observations, so the model sees several seasonal cycles and enough budget movement to separate channels. Shorter histories can work if spend varied sharply or if regional data multiplies the observations. What matters more than length is variation: a long flat history contains little usable information.
Is MMM better than multi-touch attribution?
They answer different questions. MMM works on aggregates and covers offline, brand and privacy-restricted channels, but cannot see individual journeys. Multi-touch attribution follows tracked users at a granular level but goes blind wherever identifiers are missing. Mature teams run both plus periodic holdout tests, and treat persistent disagreement between them as a signal to investigate rather than a tie to break.
What is adstock in an MMM?
Adstock is the transformation that spreads a period's advertising effect into later periods, capturing the fact that an ad seen this week can produce a conversion next month. It is defined by a decay rate: a high rate means the effect fades quickly, a low one that it lingers. Getting adstock wrong shifts contribution between fast and slow channels.
Can a small company run media mix modeling?
It is possible but often not the best use of effort. With few channels, short history and low weekly volumes, the model has little to separate and confidence intervals swamp the estimates. Smaller teams usually get further with holdout tests on one channel at a time, and revisit MMM once spend is spread across enough channels to make manual comparison unreliable.
How often should a media mix model be refreshed?
Refit quarterly as new weeks of data arrive, and rebuild the specification whenever the business changes shape: a new pricing model, a new market, or a channel that did not exist before. Between refits, use the existing response curves for planning but treat any channel whose recent spend has moved far outside its historical range as out of scope.
What does the baseline in an MMM represent?
The baseline is the portion of the outcome the model does not attribute to media: demand from brand equity, existing customers, distribution and seasonality. A very large baseline suggests media is doing less than the team believes, or that a driver is missing from the model. Tracking how the baseline moves over years is itself a measure of brand building.