Paid Social Measurement, the honest version
How to tell what your paid social is actually doing, from the platform's flattering numbers to the tests that tell the truth.
By Jonas Sluijs. I spent years on the platform side, in leadership roles at Meta, where I led the Benelux business, and at Snap as Head of Northern Europe. I managed more than $500 million in ad spend across clients like Booking, Disney and Takeaway.com, and I watched the best of them run hundreds of lift and geo tests. I saw, over and over, the gap between what a dashboard reported and what a holdout proved. This is the version of measurement I wish more advertisers had. Not a tool that promises the truth, but a way to climb toward it and always know which rung you are standing on.
Reading time: about 25 minutes. Jump to any rung using the links below.
- Your ROAS is a story the platform tells about itself
- Measurement is a ladder, not a metric
- Rung 1: the platform's own numbers
- Rung 2: a second opinion (GA4 and the plumbing)
- Rung 3: the number that doesn't lie (MER)
- Rung 4: the cheapest incrementality you're not using
- Rung 5: the real thing (lift tests)
- How to actually design a test
- Rung 6: MMM for people who aren't Google
- Rung 7: the expert answer (triangulation)
- Measurement by business model
- Connect your reporting to AI
- The whole thing in five lines
- Common questions
Your ROAS is a story the platform tells about itself
Your ROAS is a story your ad platform tells about itself, and it is a generous one. Since iOS 14 broke the tracking that made last-click look trustworthy, the number in Ads Manager and the number in your bank account quietly drifted apart. The platform still reports a return, but it counts sales it did not cause, credits view-throughs nobody remembers, and claims conversions that another channel also claims. Every platform marks its own homework, and every platform gives itself a good grade.
This matters because you make real decisions on those numbers. You scale what the platform flatters and kill what it cannot see. If the flattery is uneven, and it always is, you are quietly moving budget toward whatever games attribution best, not toward whatever grows the business. Most advertisers never find out which of those calls were wrong.
The fix is not a better dashboard. It is understanding that no single number is the truth, and knowing how to climb toward one.
Measurement is a ladder, not a metric
Measurement is a ladder, not a metric. Each rung sits closer to the truth and costs more to reach, in money, in time, or in the scale you need to pull it off. You climb as you grow. You do not skip to the top, and most businesses do not need to.
- Platform metrics. Ads Manager ROAS, CPA, CTR. Fast, free, most biased.
- Web analytics. GA4 and clean tracking. A second, independent view.
- Blended, or MER. Total revenue over total spend. Cannot be gamed by attribution.
- Post-purchase surveys. Ask the customer. The cheapest incrementality proxy.
- Lift tests. Holdouts and geo tests. Causal proof, real cost.
- MMM. Top-down modeling for allocation across everything. Needs scale.
- Triangulation. All of the above, cross-checked. The expert's actual answer.
The mistake nearly everyone makes is treating rung one, the platform's own numbers, as the answer. It is not the answer. It is the fastest, cheapest, most flattering view, and it is useful for exactly one job: steering the campaign day to day. For the question that actually matters, is this spend making me money I would not otherwise have, you have to climb. The rest of this guide is each rung: what it measures, what it hides, and when to reach for it.
Rung 1: the platform's own numbers
This is what everyone starts with: ROAS, CPA, CTR, CPM, frequency, all reported inside Ads Manager. It is genuinely useful, for one thing. It is your steering wheel, not your speedometer.
What it is good at. Relative, same-day reads. Which of my three ads has the strongest hook. Whether this ad set is fatiguing. Where cost is drifting. For in-platform optimization, the platform's own signal is the right signal, because that is the signal the algorithm is optimizing against too.
What it hides. Three things inflate reported ROAS. Attribution windows, where the default seven-day-click and one-day-view credits touches that barely influenced the sale. Cross-platform double counting, where Meta and Google both take credit for the same order. And the big one, incrementality: the platform happily counts the customer who was going to buy anyway.
How to set up Ads Manager so it lies less.
- Build a custom column set you read every time: outbound CTR (not link CTR, which is inflated by comment taps and see-more expands), CPM to reach a thousand people, frequency, the three-second view rate as a hook read, and your true north event at your target cost.
- Know your attribution window, and compare seven-day-click against one-day-view to see how much of your "performance" is view-through you should not trust.
- Use breakdowns (placement, age, gender) to diagnose, never to fragment your ad sets into pieces too small to learn.
- Save it as a custom report so you and your team read the same numbers the same way, week after week.
Use this rung to rank creative and spot fatigue. Do not use it to decide whether the channel made money. It cannot answer that, and it will always say yes.
Rung 2: a second opinion (GA4 and the plumbing)
Two imperfect measurements that disagree tell you more than one confident number ever will. GA4 is your second opinion, and the tracking underneath it is what makes every rung below less wrong.
GA4 for paid social. Tag your links, define your key conversions as events, and use the paid-social channel grouping. GA4 sees the session and the cross-channel path the platform cannot, so it is a useful check on the platform's self-report. It is not gospel either, it has its own gaps around consent and cross-device, but a second lens is the point.
UTM discipline is the most under-rated job in measurement. A strict, consistent scheme for source, medium, campaign and content is the difference between a clean paid-social bucket and one that is half wrong. Inconsistent UTMs quietly poison every report you build on top of them. Write the taxonomy down, enforce it, never freestyle a tag.
The plumbing: CAPI and match quality. Send events server to server through the Conversions API so ad blockers and opt-outs do not erase them, and push your event match quality up with email, phone, external ID and click ID. Higher match quality means the platform optimizes on a truer picture, and your reported numbers move closer to reality. Clean UTMs and server-side events do not make the numbers true. They make them less wrong, which is the whole game at this rung.
Rung 3: the number that doesn't lie (MER)
Blended is the first number on the ladder that cannot be gamed. It is total revenue divided by total ad spend, the marketing efficiency ratio, and because it does not care which channel claims a sale, attribution tricks cannot touch it.
Two companions make it sharper. New-customer CAC, total spend over new customers, strips out the returning buyers the platform loves to claim as fresh wins. Acquisition MER uses new-customer revenue only, for the same reason.
Why the board should look here. It is the one marketing number that ties straight to the P&L. If your MER is healthy while platform ROAS looks poor, you are probably profitable and mis-measuring. If platform ROAS looks great while MER sinks, you are buying sales you would have gotten for free. For most small and mid-sized advertisers, MER watched against a target over time is a better decision tool than any per-channel ROAS.
Its limit. Blended is honest but blunt. It moves with seasonality, promotions and brand, and it cannot tell you which channel to cut. It tells you the quarter is working or it is not. To know which lever caused it, you climb.
Rung 4: the cheapest incrementality you're not using
Before you spend a euro on a lift test, ask your customers. A post-purchase survey is one question at checkout, how did you hear about us, and it is the most under-used measurement tool in the whole stack.
Because it is completely independent of any pixel, it catches the channels tracking misses, the podcast, the friend, the billboard, and it exposes the uncomfortable cases where a top-performing paid channel is really just harvesting demand that something else created. Compare the survey mix to your platform-reported mix and the gap is your attribution bias, roughly quantified for the price of a form.
Any single answer is noisy, people forget and misremember, but at scale the aggregate is stable and directionally true. Tools like Fairing or KnoCommerce make it a one-click checkout question. Run it before anything fancier, because it is the cheapest read on incrementality that exists and almost nobody bothers. Your customers will tell you what your pixel cannot, if you just ask them.
Rung 5: the real thing (lift tests)
This is the rung where you stop estimating and start proving. Incrementality is the sales you got because of the ad, minus the sales you would have gotten anyway. Reported ROAS counts both. Incremental ROAS counts only the first, and the gap between them is often enormous, especially on retargeting and branded search, which mostly bill you for customers who were already on their way.
The methods, from most accessible to most rigorous:
- Conversion Lift (platform holdout). The platform withholds your ads from a random control group and compares. Easy to run, but the platform is grading its own homework, so treat a platform-run lift result as a strong hint, not a verdict, and be most skeptical when the number is most flattering.
- Geo or matched-market tests. Split regions into test and control, change spend only in the test regions, and compare the difference. Platform-agnostic, hard to game, and the gold standard for the questions that matter most: is this channel incremental at all, and what happens if I cut twenty percent. Meta's open-source GeoLift makes the market matching and the analysis rigorous. It needs enough geographic spread and volume to work.
- PSA or ghost tests. The control group sees a placebo, a public-service ad, instead of nothing, which controls for the auction dynamics. Cleaner and more expensive, less common.
- Rolling holdouts. Hold back a slice of your addressable audience on an always-on channel and measure the delta over time.
What you get is an incrementality ratio, incremental over reported, and a true iROAS. Once you know your retargeting is twenty percent incremental and your prospecting is eighty, you can finally allocate on truth instead of last-click. The recurring lesson from running these at scale is uncomfortable and consistent: the line with the best reported ROAS is usually the least incremental, cannibalization hides comfortably inside a great-looking dashboard, and the only honest way to learn what a channel is worth is to turn part of it off and watch what happens. If you have never run a holdout, you do not know what your ads are worth. You know what the platform says they are worth.
How to actually design a test
Most "just run an A/B test" advice is dangerous, because it skips the design that decides whether the result means anything. A sloppy test does not give you a weaker answer, it gives you a confident wrong one.
- Hypothesis first. One specific, falsifiable claim, for example "cutting retargeting fifty percent will not reduce total sales by more than five percent." No hypothesis, no test.
- Minimum detectable effect. The smallest lift worth acting on. A smaller effect needs more sample and more time. Decide it up front, because it sets everything else.
- Power and sample size. You need enough conversions in both arms to separate signal from noise. Underpowered tests are the norm and they produce confident nonsense. As a rough gut check, if you cannot gather a few hundred conversions per arm in a sensible window, you cannot reliably measure a small effect.
- Holdout size. A bigger holdout gives more statistical power and costs more foregone sales. Ten to twenty percent is a common range for lift.
- Duration. At least one full purchase cycle, ideally two, so delayed conversions land. Too short and you measure the click, not the customer.
- Contamination. Make sure test and control do not bleed into each other through the same user, geographic spillover or cross-device exposure. This ruins more tests quietly than anything else on this list.
- Novelty and seasonality. A new creative gets a honeymoon, and a test run across a big sale measures the sale. Run long enough to pass the novelty bump, and avoid or account for confounding events.
- One change at a time. Change the spend, not the spend and the creative and the landing page together, or you will not know which one moved the number.
The discipline is the credibility. Decide what result would change your mind before you run the test, or you will simply confirm what you already believed.
Rung 6: MMM for people who aren't Google
Media mix modeling is a top-down statistical model that takes all your spend, every channel, seasonality, promotions and baseline demand, and estimates each channel's contribution to sales. It uses no user-level data, so it survives every privacy change, and it is the right tool for one question: how should I split budget across everything at once.
The honest cutoff: you are probably not big enough yet. MMM needs years of weekly history, real spend across several channels, and enough variation to model. Below roughly seven figures of annual spend, it overfits and hands you a confident story that is mostly noise. If that is you, stay on MER, surveys and lift tests, and come back to MMM when you have the scale and the history. Saying so cuts against a whole industry that wants to sell you a model, and it is still the right advice.
If you are big enough. You no longer have to buy a six-figure consultancy to start: Meta's Robyn and Google's Meridian are open source. The move that separates a useful model from an expensive horoscope is calibration. An MMM on its own is correlational and can be badly wrong, so you feed it the causal ground truth from your geo and lift tests to anchor it. Calibrating the model with experiments is the single biggest quality lever in modern MMM. Watch the saturation curves and carryover, the diminishing-returns and lingering-effect assumptions, because that is where an MMM earns its keep and also where it is most easily fudged.
Rung 7: the expert answer (triangulation)
No single method is the truth. The people who measure well do not crown a winner, they triangulate, and use each tool for the job it is honest about.
- Platform metrics for daily steering: which creative, is it fatiguing, where is cost drifting.
- Experiments, lift and geo, for causal reads: is this channel incremental, what happens if I cut it.
- MMM for allocation across everything, calibrated by those experiments.
- MER as the constant reality check against the P&L.
When these agree, act with confidence. When they disagree, that disagreement is the most valuable signal you have, because it is pointing straight at the measurement that is lying, and finding out which one is how you actually learn your account. Three imperfect views, cross-checked, beat one confident number every time. Trust the triangle, not the dashboard.
Measurement by business model
The right rung and the right metric depend on what you sell. Pick the scorecard that matches how the business makes money, then measure that, not whatever the platform hands you. We build a full plan for each business type, but the measurement shape is:
- Ecommerce. MER and iROAS, watched on new-customer CAC rather than the blended ROAS the platform reports. Your break-even ROAS is the floor, and you can model it in the break-even calculator.
- Subscription. Forget first-order ROAS. Measure payback period and cohort LTV, and optimize to the subscription-start event. A subscription with a four-month payback looks like a disaster on day-one ROAS and a great business on a cohort curve. See the AG1 and Whoop plans.
- Lead gen. Cost per qualified lead, not cost per form fill. The real measurement job is connecting spend to lead quality and closed revenue in the CRM, or you will optimize toward cheap junk.
- CPG and retail. The pixel is blind, because the sale happens on a shelf or on Amazon. Measure on incrementality, retail velocity and marketplace rank, not last-click ROAS. See the Liquid Death plan, where the CMO says platform ROAS "doesn't tell me anything."
Connect your reporting to AI
A lot of measurement time dies in the reporting layer: exporting, pivoting, and eyeballing rows. That is exactly what a language model is good at, and almost nobody has wired it up well yet.
The simple version: export your Ads Manager or GA4 data, by CSV or API, and hand it to an AI model with a precise job. Not "analyze this," which returns mush, but a specific ask:
- Anomaly detection. "Here is thirty days of daily spend, CPA and CPM by campaign. Flag any day or campaign where CPA moved more than fifteen percent, and give the most likely cause from the other columns."
- What changed and why. "Compare this week to last. Was the CPA change driven by CPM, which is auction and reach, CTR, which is creative, or conversion rate, which is landing page and offer?" That CPM, CTR, conversion-rate split is the single most useful ad-diagnosis prompt there is.
- Creative fatigue. "Here is frequency, CPM-reach and outbound CTR by ad over time. Which ads are fatiguing, and in what order should I refresh them?"
- Reallocation. "Given these blended numbers and my target MER, where would you move budget, and what is the risk of each move?"
The guardrail matters. An AI model will confidently invent a story from noise, the same failure mode as an underpowered test. Give it the decomposition framework, ask it to show its reasoning, and treat it as a fast analyst you still supervise, not an oracle. The value is speed and a second read, not outsourced judgment. This is also where paid.social is heading: measurement you can talk to.
The whole thing in five lines
- The platform's ROAS is for steering the campaign, never for judging whether it made money.
- Blended MER is the honest number. Watch it against your P&L.
- Ask your customers how they heard about you before you spend on anything fancier.
- If you have never run a holdout or a geo test, you do not know what your ads are worth.
- Triangulate: platform for steering, experiments for truth, MMM for allocation, MER for the check. When they disagree, that is the lesson.
Get those right and you stop paying for sales you would have made anyway, and start seeing the ones you would not.
Want a second pair of eyes on your measurement setup before you scale spend? Book a free intake call.