How to Run Media Mix Modeling Without a Data Science Team: An 8-Step Guide (2026)
You can run media mix modeling without a data science team by using a free open-source tool, Google's Meridian or Meta's Robyn, on two to three years of weekly spend and sales data, limited to a handful of channels. Meridian is Python-based, so someone has to run it. The real work is data preparation and calibration, not statistics.
Media mix modeling (MMM) has a reputation as an enterprise-only exercise: six-figure vendors, quarterly decks, a team of econometricians. That reputation is out of date. The modeling code is now free. What has not changed is that the method is unforgiving about inputs. A model fed thin or messy data produces confident-looking nonsense.
This guide is the practical version. It walks through eight steps, includes a readiness scorecard you can screenshot and score in ten minutes, and tells you when the honest answer is “not yet.” If your current measurement is mostly last-click, read our comparison of last-click, data-driven and MMM attribution first; this post assumes you have decided MMM is worth trying.
Step 1: Decide what question the model must answer
MMM answers one question well: how should the budget be split across channels? Google's own documentation is explicit that MMM is a macro tool that works at the channel level, and it advises against campaign-level data. If you want to know which ad set or which creative wins, use platform reporting and a proper testing framework instead.
Write your question down in one sentence before you touch any data. A good one reads: “If we move 10% of budget from channel A to channel B, what happens to total sales?” A bad one reads: “What is working?” The first can be modeled. The second cannot.
Step 2: Score your readiness before you build anything
Most failed MMM projects fail on inputs, not on software. Score yourself against the table below. It takes ten minutes and saves weeks. The thresholds for history come from Meridian's documentation; the remaining rows reflect the data problems that break models in practice.
NAMEX MMM Readiness Scorecard (2026)
| Check | Ready (2) | Partial (1) | Not ready (0) |
|---|---|---|---|
| Weekly history | 3+ years | 2 to 3 years | Under 2 years |
| Outcome data (KPI) | Weekly sales or leads from your own system | Platform-reported conversions only | No consistent weekly record |
| Spend records | Weekly spend by channel, reconciled to invoices | Platform exports, some gaps | Spend not tracked by channel |
| Channels to model | 3 to 6 meaningful channels | 7 to 9 | 10+ or many tiny channels |
| Budget variation | Spend has moved up and down by channel | Small changes only | Flat budgets, no variation |
| Controls | Pricing, promotions, seasonality logged | Some logged | None recorded |
| Experiments | At least one holdout or geo test run | Planned | None and none planned |
| Owner | Named person to run and refresh quarterly | Outside help for the run | Nobody |
Budget variation deserves a note. A model learns from change. If you have spent the same amount on the same channels for three years, there is nothing for it to learn from, and no vendor can fix that with clever code.
Step 3: Assemble one weekly dataset
Meridian needs three core inputs: a KPI that can be summed across time, media data (spend, impressions or clicks) by channel and week, and control variables that could otherwise be mistaken for media effects. Weekly granularity is the sweet spot Google recommends, because daily data is noisy and monthly data throws away variation.
Build a single spreadsheet with one row per week. Columns: total sales or leads, spend for each channel, and a handful of controls such as a promotion flag, price changes, and a holiday marker. Missing media weeks are filled with zero. Missing KPI or control values are interpolated. Google's guidance on this is specific, and getting it wrong distorts the model quietly.
Pull spend from invoices where you can, not just from platform dashboards. Reconciling the two is tedious and it is the single most valuable hour in the project. If you have not audited your account data lately, the six-step wasted ad spend audit is a good warm-up.
Step 4: Cut the number of parameters
This is the step most teams skip. Google frames sample size as data points per model parameter and aims for roughly 10 to 15. Its worked example uses 12 channels, 6 controls and 8 knots, which is 26 parameters. With two years of weekly data (104 points) that gives only 4 points per parameter, which is too few for reliable estimates.
Do the arithmetic on your own data. Three years of weekly national data is 156 points. At 10 points per parameter, that supports about 15 parameters in total. Counting each channel, control and time knot as one parameter, as Google's example does, a sensible build is five channels, three controls and six knots. Two years of data cuts that ceiling to about 10.
The fix is to group. Combine small channels into a single line (for example, all paid social under one heading) and drop anything under a few percent of spend. You lose granularity and gain a model you can trust. For a view of how spend is typically distributed, see our 2026 digital ad spend by channel benchmarks.
Step 5: Pick the tool that matches your team
You have four realistic routes. None of them removes the need for judgment; they differ in who does the coding and what they cost.
| Route | Software cost | Coding needed | Main strength | Main risk |
|---|---|---|---|---|
| Google Meridian | Free, open source | Yes, Python | Google reach and frequency data; GeoX experiment calibration | Google data advantage can bias results toward Google channels |
| Meta Robyn | Free, open source | Yes, R | Long-running, documented, built for smaller advertisers | AdExchanger reports Meta has scaled back the engineering team |
| Paid vendor platform | Subscription or project fee | No | Managed setup and refreshes | Many are built on Meridian or Robyn code; you pay for packaging |
| Incrementality tests only | Media cost of the holdout | No | Direct causal evidence for one channel at a time | Slow; cannot cover every channel at once |
Meridian is the practical default for most advertisers because of its documentation and its experiment tooling, and Google added Meridian integration for Analytics customers in 2026 according to AdExchanger. Be aware of the criticism in that same reporting: an expert described the tool as free in the sense that a puppy is free. It gets its value from Google's own data, and it needs customization for each business, which rarely happens. Treat its output on Google channels with a raised eyebrow, and test it.
If nobody on your team can run a notebook, hire someone to run the model once and hand you the assumptions in plain language. That is a few days of contractor time, not a headcount.
Not sure your channel mix is even a good candidate? Request a free audit of your channel mix and we will tell you whether modeling is worth the effort or whether a simpler test would answer your question faster.
Step 6: Calibrate with at least one experiment
A model built purely on history can mistake coincidence for cause. If you raised spend on a channel every December, the model cannot cleanly separate the channel from the season. An experiment breaks the tie. Meridian's GeoX lets you run geo-level tests and feed the result back into the model as an experimental prior, or use the test as a standalone causal read.
The minimum viable version is a geo holdout on your largest channel: switch it off or down in a set of comparable regions for several weeks, keep everything else steady, and compare outcomes. For CTV and streaming specifically, our step-by-step CTV attribution guide covers holdout design and the CTV advertising benchmarks give you cost context.
Step 7: Read the output as ranges, not answers
MMM gives you an estimated return for each channel, a saturation curve showing where extra spend stops paying, and an effect delay. Every one of these comes with an uncertainty band. If the band for a channel runs from a strong return to nearly none, the honest report is that the model cannot tell yet.
Run three sanity checks. First, do the channel rankings roughly agree with your last-click data, and where they differ, is there a reason (upper-funnel channels rarely get click credit)? Second, do the baseline sales look plausible against your brand's history? Third, does the model behave when you drop a week or a channel? If small changes flip the conclusions, simplify the model.
Step 8: Turn the result into one budget move
Do not rebuild the whole plan. Pick the biggest gap between where money goes and where the model says returns are strongest, and move a modest slice, say 10% of that channel's budget, for a defined period. Treat it as the next experiment. This is our rule of thumb, not a Google or Meta recommendation, and it keeps the downside small while the model earns your trust.
Then set a refresh rhythm. Quarterly is realistic for a small team: append the new weeks, rerun, compare to the last version. If your streaming spend is sizeable, our streaming ad CPM benchmarks show what prices to expect when you shift budget there, and our targeting services and programmatic services pages describe how we plan the resulting buy.
Common mistakes to avoid
Modeling too many channels for the data you have is the most common error, followed by using platform-reported conversions as the KPI, which bakes each platform's own attribution bias into the model. Others: skipping control variables, changing the model setup after seeing results until it says what you hoped, and treating a single run as permanent truth.
The bottom line
MMM without a data science team is realistic when you keep the scope small: three years of weekly data if you can get it, five or six channels, one experiment, and a person who owns the quarterly refresh. If you cannot score ten on the readiness table, fix the data or run incrementality tests first. That is not a failure. It is the order the work has to happen in.
Frequently asked questions
Can a small team really run media mix modeling without a data scientist?
Yes, within limits. Meridian and Robyn are free, but Meridian is a Python library, so someone comfortable running notebooks has to execute it. The harder part is assembling two to three years of clean weekly data and keeping the channel count low. A marketer who can handle spreadsheets and follow documentation can manage that with outside help on the modeling run.
How much data do I need for a media mix model?
Google's Meridian documentation recommends at least two years of weekly data for geo-level models and three years for national-level models. It also frames the target as data points per model parameter, aiming for roughly 10 to 15. Two years of weekly national data is 104 points, which supports only a small number of channels and controls.
Is media mix modeling better than attribution?
They answer different questions. Attribution tracks user-level paths and misses channels that do not click, such as CTV and DOOH. Media mix modeling works on aggregate spend and outcomes, so it can see those channels, but it is slow, needs history and produces ranges rather than exact answers. Use it for budget allocation and use platform data for in-flight optimization.
Do I need to run an incrementality test as well?
It is strongly advisable. A model built only on observational data can confuse correlation with cause, especially when budgets moved together with seasonality. Meridian includes GeoX for geo experiments that can calibrate the model. Even one clean holdout on your biggest channel gives the model an anchor and gives you a sanity check on its output.
Next steps
Want a second opinion on whether your channel mix is ready for modeling? Book an intro call with Ryan or request your free audit. We work with no minimum spend and no retainer fees, and we will tell you plainly if a simpler test is the better first move.
Sources: Google for Developers, Meridian documentation (data requirements and FAQ, retrieved 30 September 2026); AdExchanger, reporting on Meridian and Robyn (2026); Robyn paper on arXiv (2403.14674); IAB, State of Data 2026 (published 2 February 2026).