How we audit your MMM

A technically sound model that nobody trusts is still a failed measurement programme. We start with how the MMM supports decisions, then stress-test whether the outputs deserve that trust.

Decision use and trust come first

A common pattern: the MMM exists, the decks look polished, and yet no one really trusts it. Marketing plans around it. Finance discounts it. Leadership asks for the model, then decides with judgement, last year’s mix, or platform ROAS instead.

That is not a side issue. If outputs are not used, the organisation is paying for theatre. If they are used without trust, budgets move on shaky ground. Our audit asks both questions explicitly:

  • Is the model used? Which decisions does it inform (or claim to inform): annual allocation, in-year reforecasts, channel cuts, scenario planning?
  • Is it trusted? By whom: marketing, finance, agencies, leadership? Where does scepticism come from, and is it justified?
  • What would make it decision-ready? Technical fixes, clearer caveats, better governance, or an honest “do not use for X” boundary.

We treat “shelfware MMM” as a first-class finding. Sometimes the priority is restoring credibility so decisions can use the model. Sometimes it is stopping people from over-claiming what the model can do.

Stacks we regularly review

Google Meridian

Bayesian hierarchical MMM. We examine prior choices, geo / national structure, media transforms, whether outputs are used within the model's identification limits, and whether stakeholders actually rely on those outputs.

Meta Robyn

Nevergrad-optimised MMM. We check hyperparameter search design, solution clustering, spend constraints, stability under reasonable perturbations, and whether “best” models are trusted enough to change spend.

PyMC-Marketing

Bayesian MMM in the PyMC ecosystem. We review model structure, prior elicitation, MCMC diagnostics, and how channel effects are communicated so decision-makers can trust (or correctly discount) them.

Proprietary / vendor models

We work with the commissioning company, not the vendor. Where code is unavailable, we audit documentation, data dictionaries, contribution reports, experimental calibration evidence, and how the business is expected to use the numbers.

The checklist

  • Decision use: which planning and budget decisions the model is meant to support, and whether those decisions actually use it
  • Trust & governance: who believes the outputs, who does not, and whether caveats match how results are presented
  • Specification: business question, identification, controls, hierarchy
  • Assumptions: priors, constraints, omitted variables, structural choices
  • Data: coverage, leakage, seasonality, price/promo treatment
  • Adstock & saturation: transforms and parameter plausibility
  • Validation: holdout, sensitivity, stability across seeds/runs
  • Calibration: agreement with geo-lift / incrementality (optional)

Four steps, three weeks

Standard turnaround three weeks. Two-week rush available where capacity allows, typically +25%.

1

Intro + scope

Free 20-min fit call, then optional fixed-fee scoping (£750) credited to the audit.

2

You share

Model code or spec, data, and experiment results under NDA.

3

We audit

Decision use and trust first, then the technical checklist; reproduction if scoped.

4

Report + walkthrough

Verdict on trust and decision-readiness, ranked issues, and a findings call.

Discuss how the model is (or is not) used

Tell us what you run, who trusts it, and which decisions hang on it. We'll confirm fit and scope on a short call.