← Back to the Meta Ads Guide
Measurement Guide

Paid Social Measurement, the honest version

How to tell what your paid social is actually doing, from the platform's flattering numbers to the tests that tell the truth.

By Jonas Sluijs. I spent years on the platform side, in leadership roles at Meta, where I led the Benelux business, and at Snap as Head of Northern Europe. I managed more than $500 million in ad spend across clients like Booking, Disney and Takeaway.com, and I watched the best of them run hundreds of lift and geo tests. I saw, over and over, the gap between what a dashboard reported and what a holdout proved. This is the version of measurement I wish more advertisers had. Not a tool that promises the truth, but a way to climb toward it and always know which rung you are standing on.

Reading time: about 25 minutes. Jump to any rung using the links below.

Your ROAS is a story the platform tells about itself

Your ROAS is a story your ad platform tells about itself, and it is a generous one. Since iOS 14 broke the tracking that made last-click look trustworthy, the number in Ads Manager and the number in your bank account quietly drifted apart. The platform still reports a return, but it counts sales it did not cause, credits view-throughs nobody remembers, and claims conversions that another channel also claims. Every platform marks its own homework, and every platform gives itself a good grade.

This matters because you make real decisions on those numbers. You scale what the platform flatters and kill what it cannot see. If the flattery is uneven, and it always is, you are quietly moving budget toward whatever games attribution best, not toward whatever grows the business. Most advertisers never find out which of those calls were wrong.

The fix is not a better dashboard. It is understanding that no single number is the truth, and knowing how to climb toward one.

Measurement is a ladder, not a metric

Measurement is a ladder, not a metric. Each rung sits closer to the truth and costs more to reach, in money, in time, or in the scale you need to pull it off. You climb as you grow. You do not skip to the top, and most businesses do not need to.

  1. Platform metrics. Ads Manager ROAS, CPA, CTR. Fast, free, most biased.
  2. Web analytics. GA4 and clean tracking. A second, independent view.
  3. Blended, or MER. Total revenue over total spend. Cannot be gamed by attribution.
  4. Post-purchase surveys. Ask the customer. The cheapest incrementality proxy.
  5. Lift tests. Holdouts and geo tests. Causal proof, real cost.
  6. MMM. Top-down modeling for allocation across everything. Needs scale.
  7. Triangulation. All of the above, cross-checked. The expert's actual answer.

The mistake nearly everyone makes is treating rung one, the platform's own numbers, as the answer. It is not the answer. It is the fastest, cheapest, most flattering view, and it is useful for exactly one job: steering the campaign day to day. For the question that actually matters, is this spend making me money I would not otherwise have, you have to climb. The rest of this guide is each rung: what it measures, what it hides, and when to reach for it.

Rung 1: the platform's own numbers

This is what everyone starts with: ROAS, CPA, CTR, CPM, frequency, all reported inside Ads Manager. It is genuinely useful, for one thing. It is your steering wheel, not your speedometer.

What it is good at. Relative, same-day reads. Which of my three ads has the strongest hook. Whether this ad set is fatiguing. Where cost is drifting. For in-platform optimization, the platform's own signal is the right signal, because that is the signal the algorithm is optimizing against too.

What it hides. Three things inflate reported ROAS. Attribution windows, where the default seven-day-click and one-day-view credits touches that barely influenced the sale. Cross-platform double counting, where Meta and Google both take credit for the same order. And the big one, incrementality: the platform happily counts the customer who was going to buy anyway.

How to set up Ads Manager so it lies less.

Use this rung to rank creative and spot fatigue. Do not use it to decide whether the channel made money. It cannot answer that, and it will always say yes.

Rung 2: a second opinion (GA4 and the plumbing)

Two imperfect measurements that disagree tell you more than one confident number ever will. GA4 is your second opinion, and the tracking underneath it is what makes every rung below less wrong.

GA4 for paid social. Tag your links, define your key conversions as events, and use the paid-social channel grouping. GA4 sees the session and the cross-channel path the platform cannot, so it is a useful check on the platform's self-report. It is not gospel either, it has its own gaps around consent and cross-device, but a second lens is the point.

UTM discipline is the most under-rated job in measurement. A strict, consistent scheme for source, medium, campaign and content is the difference between a clean paid-social bucket and one that is half wrong. Inconsistent UTMs quietly poison every report you build on top of them. Write the taxonomy down, enforce it, never freestyle a tag.

The plumbing: CAPI and match quality. Send events server to server through the Conversions API so ad blockers and opt-outs do not erase them, and push your event match quality up with email, phone, external ID and click ID. Higher match quality means the platform optimizes on a truer picture, and your reported numbers move closer to reality. Clean UTMs and server-side events do not make the numbers true. They make them less wrong, which is the whole game at this rung.

Rung 3: the number that doesn't lie (MER)

Blended is the first number on the ladder that cannot be gamed. It is total revenue divided by total ad spend, the marketing efficiency ratio, and because it does not care which channel claims a sale, attribution tricks cannot touch it.

Two companions make it sharper. New-customer CAC, total spend over new customers, strips out the returning buyers the platform loves to claim as fresh wins. Acquisition MER uses new-customer revenue only, for the same reason.

Why the board should look here. It is the one marketing number that ties straight to the P&L. If your MER is healthy while platform ROAS looks poor, you are probably profitable and mis-measuring. If platform ROAS looks great while MER sinks, you are buying sales you would have gotten for free. For most small and mid-sized advertisers, MER watched against a target over time is a better decision tool than any per-channel ROAS.

Its limit. Blended is honest but blunt. It moves with seasonality, promotions and brand, and it cannot tell you which channel to cut. It tells you the quarter is working or it is not. To know which lever caused it, you climb.

Rung 4: the cheapest incrementality you're not using

Before you spend a euro on a lift test, ask your customers. A post-purchase survey is one question at checkout, how did you hear about us, and it is the most under-used measurement tool in the whole stack.

Because it is completely independent of any pixel, it catches the channels tracking misses, the podcast, the friend, the billboard, and it exposes the uncomfortable cases where a top-performing paid channel is really just harvesting demand that something else created. Compare the survey mix to your platform-reported mix and the gap is your attribution bias, roughly quantified for the price of a form.

Any single answer is noisy, people forget and misremember, but at scale the aggregate is stable and directionally true. Tools like Fairing or KnoCommerce make it a one-click checkout question. Run it before anything fancier, because it is the cheapest read on incrementality that exists and almost nobody bothers. Your customers will tell you what your pixel cannot, if you just ask them.

Rung 5: the real thing (lift tests)

This is the rung where you stop estimating and start proving. Incrementality is the sales you got because of the ad, minus the sales you would have gotten anyway. Reported ROAS counts both. Incremental ROAS counts only the first, and the gap between them is often enormous, especially on retargeting and branded search, which mostly bill you for customers who were already on their way.

The methods, from most accessible to most rigorous:

What you get is an incrementality ratio, incremental over reported, and a true iROAS. Once you know your retargeting is twenty percent incremental and your prospecting is eighty, you can finally allocate on truth instead of last-click. The recurring lesson from running these at scale is uncomfortable and consistent: the line with the best reported ROAS is usually the least incremental, cannibalization hides comfortably inside a great-looking dashboard, and the only honest way to learn what a channel is worth is to turn part of it off and watch what happens. If you have never run a holdout, you do not know what your ads are worth. You know what the platform says they are worth.

How to actually design a test

Most "just run an A/B test" advice is dangerous, because it skips the design that decides whether the result means anything. A sloppy test does not give you a weaker answer, it gives you a confident wrong one.

The discipline is the credibility. Decide what result would change your mind before you run the test, or you will simply confirm what you already believed.

Rung 6: MMM for people who aren't Google

Media mix modeling is a top-down statistical model that takes all your spend, every channel, seasonality, promotions and baseline demand, and estimates each channel's contribution to sales. It uses no user-level data, so it survives every privacy change, and it is the right tool for one question: how should I split budget across everything at once.

The honest cutoff: you are probably not big enough yet. MMM needs years of weekly history, real spend across several channels, and enough variation to model. Below roughly seven figures of annual spend, it overfits and hands you a confident story that is mostly noise. If that is you, stay on MER, surveys and lift tests, and come back to MMM when you have the scale and the history. Saying so cuts against a whole industry that wants to sell you a model, and it is still the right advice.

If you are big enough. You no longer have to buy a six-figure consultancy to start: Meta's Robyn and Google's Meridian are open source. The move that separates a useful model from an expensive horoscope is calibration. An MMM on its own is correlational and can be badly wrong, so you feed it the causal ground truth from your geo and lift tests to anchor it. Calibrating the model with experiments is the single biggest quality lever in modern MMM. Watch the saturation curves and carryover, the diminishing-returns and lingering-effect assumptions, because that is where an MMM earns its keep and also where it is most easily fudged.

Rung 7: the expert answer (triangulation)

No single method is the truth. The people who measure well do not crown a winner, they triangulate, and use each tool for the job it is honest about.

When these agree, act with confidence. When they disagree, that disagreement is the most valuable signal you have, because it is pointing straight at the measurement that is lying, and finding out which one is how you actually learn your account. Three imperfect views, cross-checked, beat one confident number every time. Trust the triangle, not the dashboard.

Measurement by business model

The right rung and the right metric depend on what you sell. Pick the scorecard that matches how the business makes money, then measure that, not whatever the platform hands you. We build a full plan for each business type, but the measurement shape is:

Connect your reporting to AI

A lot of measurement time dies in the reporting layer: exporting, pivoting, and eyeballing rows. That is exactly what a language model is good at, and almost nobody has wired it up well yet.

The simple version: export your Ads Manager or GA4 data, by CSV or API, and hand it to an AI model with a precise job. Not "analyze this," which returns mush, but a specific ask:

The guardrail matters. An AI model will confidently invent a story from noise, the same failure mode as an underpowered test. Give it the decomposition framework, ask it to show its reasoning, and treat it as a fast analyst you still supervise, not an oracle. The value is speed and a second read, not outsourced judgment. This is also where paid.social is heading: measurement you can talk to.

The whole thing in five lines

Get those right and you stop paying for sales you would have made anyway, and start seeing the ones you would not.

Want a second pair of eyes on your measurement setup before you scale spend? Book a free intake call.

Common questions

What is incrementality in advertising?

Incrementality is the share of sales an ad actually caused, on top of the sales you would have gotten anyway. Reported ROAS counts both, so it overstates value, especially on retargeting and branded search. Incremental ROAS counts only the sales that would not have happened without the ad, measured with a holdout or geo test.

Why is my ROAS higher in Ads Manager than in reality?

Because the platform credits itself generously. It counts view-through conversions, claims sales another channel also claims, and includes customers who would have bought without the ad. Ads Manager ROAS is a steering metric, not a measure of profit created. Blended MER and a holdout test tell you the real figure.

What is the difference between ROAS and iROAS?

ROAS is platform-attributed revenue over spend. iROAS, incremental ROAS, is only the revenue that would not have happened without the ad, over spend. iROAS is almost always lower, and the gap is largest on lower-funnel tactics like retargeting that mostly harvest existing demand.

What is MER in marketing?

MER, marketing efficiency ratio, is total revenue over total ad spend across all channels. Because it ignores which channel claims each sale, it cannot be gamed by attribution, which makes it the honest board-level number. New-customer CAC, spend over new customers, strips out returning buyers for a cleaner acquisition read.

How do I run an incrementality test on Meta?

The two main methods are a platform Conversion Lift, where ads are withheld from a random control group, and a geo test, where you split regions into test and control and change spend only in the test regions. Geo tests are harder to game. Decide your minimum detectable effect and holdout size first, run for at least one full purchase cycle, and guard against contamination between groups.

Do I need media mix modeling (MMM)?

Probably not until you are spending well into seven figures a year across several channels with a few years of history. Below that, an MMM overfits and produces a confident story that is mostly noise. Smaller advertisers should rely on blended MER, surveys and lift tests. When you do run MMM, calibrate it with experiments so it reflects causal reality, not just correlation.

How do I measure paid social after iOS 14 and cookie loss?

Stop treating any single tracked number as truth. Send events server-side through the Conversions API with strong match quality, cross-check the platform against GA4, and anchor everything to blended MER against your P&L. For causal reads that survive privacy changes, use holdout and geo tests, and at scale media mix modeling, none of which rely on user-level tracking.

What is a post-purchase survey and does it work?

It is a single checkout question, how did you hear about us, answered by the buyer. Because it is independent of any pixel, it catches channels tracking misses and reveals when a paid channel is really harvesting demand another channel created. Any single answer is noisy, but at scale the pattern is stable and directionally true, making it the cheapest incrementality proxy that exists.

How long should an incrementality test run?

At least one full purchase cycle, and ideally two, so delayed conversions land and a new creative gets past its novelty bump. Too short and you measure the click instead of the customer. Avoid running across a confounding event like a major sale unless you are measuring that period, and make sure both groups have enough conversions to reach statistical power.

See measurement by business type Book a free intake call →