Meta's A/B test measures the algorithm, not your creative.
In 181,890 Meta A/B tests, 22% show audience imbalance that compromises the creative comparison. Meta's own Lift tests: 0.16%. Here is the gap.

On May 22, 2026, two economists submitted a paper to arXiv.
Pallavi Pal at Stevens Institute of Technology. Anjana Susarla at Michigan State's Broad College of Business. The title of their paper: ‘Algorithm or Creative? A Three-Arm Experimental Design for Decomposing Algorithmic Bias in Platform A/B Tests.’[2]
Their live Meta field experiment measured a specific thing. When a creative ‘wins’ a Meta A/B test, how much of that win is the creative itself, and how much is Meta's delivery algorithm silently allocating impressions to whichever variant its early signal favors?
Their answer, from the field experiment: about 75% of the audience reallocation credited to the winner was the algorithm. Not the creative. Not what users saw. The delivery decision Meta's system made before either variant had enough impressions for statistical significance.[2]
And in a two-arm A/B test — the kind Ads Manager runs by default — the algorithmic channel is understated by a factor of two.
That paper was not the first to raise the alarm. Nine months earlier, in December 2025, a different group had already quantified the scale of the problem across 181,890 real Meta A/B tests and 27 billion user-test observations. Two of the five authors were Meta Research Scientists.[1]
Their finding: 22% of those 181,890 tests show audience imbalance large enough to compromise the creative-vs-creative comparison. Meta's own Lift tests, run with a true no-ad holdout, show 0.16%.[1]
That is a 137-fold gap between what Meta's standard A/B test tool measures and what Meta's own Lift tool measures.
If you have been A/B testing creatives on Meta, you have been A/B testing Meta's delivery algorithm.
Meta's standard A/B test tool does not cleanly measure creative quality. It measures a confound of creative quality plus Meta's delivery algorithm's targeting decisions. Across 181,890 real Meta A/B tests analyzed by researchers including two Meta Research Scientists, 22% show audience imbalance large enough to compromise the comparison at the standard academic threshold (SMD greater than 0.2).[1] Meta's own Lift tests, which use a randomized no-ad holdout, show 0.16% at the same threshold. In a live causal decomposition, about 75% of the audience reallocation credited to a ‘winning’ creative traces to the algorithm, not the creative.[2] The tool most media buyers use to pick between creatives is the tool the peer-reviewed and preprint literature now documents as delivery-confounded.
What ‘divergent delivery’ actually means
Divergent delivery is the technical name for what happens inside Meta's auction when you run an A/B test.
In a clean A/B test, both variants should reach interchangeable audiences so the only thing that differs between the two groups is the creative. That is the entire logic of the design. Randomize the assignment, hold everything else constant, measure the difference.
Meta's delivery algorithm does not do that. It uses early signal — predicted click-through rate, predicted conversion likelihood, engagement history — to decide which variant to show which impression, and it makes that decision before either variant has accumulated enough impressions for statistical significance. If the algorithm predicts Ad A will perform better in the 25-34 female segment, it shows Ad A more to that segment. If Ad B looks better for older users on desktop, Ad B goes there.
At the end of the test, Ad A ‘wins.’ The verdict looks clean. But what actually happened is that Ad A was shown to a systematically different, systematically higher-converting mix of users than Ad B. The creative did not win the comparison. The algorithm won it, by pre-selecting who saw what.
Platform A/B testing tools deliver ads to distinct and undetectably optimized mixes of users that vary across ads, even during the test.Braun & Schwartz — Journal of Marketing, January 2025 (peer-reviewed)
Braun and Schwartz named this in the Journal of Marketing in January 2025.[3] Their argument is structural: the confound is not a bug that better test design fixes. It is baked into how the auction serves impressions to test arms.
The 22% study, unpacked
Burtch, Moakler, Gordon, Zhang, and Hill submitted their paper to arXiv on December 16, 2025. It is Marketing Science Institute Working Paper 25-140.[1] The MSI page for the associated webinar lists Robert Moakler (Research Scientist, Meta) and Poppy Zhang (Senior Research Scientist, Meta) as co-presenters alongside the academic authors from Kellogg and Minnesota.[5]
Their sample: 181,890 real Meta A/B tests, plus 3,204 Lift tests. Roughly 27 billion user-test observations across the A/B set. This is not a lab study or a single-campaign field experiment. It is a census-scale audit of Meta's own experimentation tool.
Their headline findings.
| Metric | Meta A/B tests (n=181,890) | Meta Lift tests (n=3,204) |
|---|---|---|
| Share of tests with t-statistic significant at p ≤ 0.05 | 25% | Not reported |
| Share of tests with SMD greater than 0.2 (compromised comparison) | 22% | 0.16% |
| Restricted awareness-optimized A/B subsample (n=612) | 18% t-stat significant, 5.07% SMD > 0.2 | — |
The paper's own framing, notable because two Meta scientists co-authored it: divergent delivery in A/B tests is intentional. It reflects real-world delivery under business-as-usual deployment. Meta is not calling it a bug. Meta is calling it a feature.[1]
That framing is the honest reading. But for a media buyer, the practical consequence is the same: when the tool measures a confound of the creative and the delivery, the buyer does not know which one drove the verdict. Doubling the media budget on the ‘winner’ scales whatever Meta's early-delivery signal favored, which may or may not be the creative that will keep winning under future auction conditions.
The 75% study, unpacked
Pal and Susarla's May 2026 paper takes the confound one step further. Instead of measuring imbalance across many tests, they built an experimental design that causally separates the algorithmic channel from the creative channel.[2]
Three arms. All three hold campaign objective, targeting, budget, placement, frequency cap, schedule, ad format, and auction infrastructure constant. Only one thing varies: a ‘women-targeted text fragment.’
Arm 1 — Control. Generic creative. No women-targeted text in metadata or the rendered ad. Algorithm sees no demographic signal.
Arm 2 — Metadata-only. Identical user-facing creative to Arm 1. But the delivery algorithm receives the women-targeted metadata fragment. Users cannot tell the difference from control.
Arm 3 — Visible treatment.User-facing creative explicitly states ‘Women applicants encouraged.’ Algorithm also sees the metadata signal.
Arm 2 versus Arm 1 isolates the algorithmic channel. Arm 3 versus Arm 2 isolates the creative channel. The design point-identifies both without requiring the strong assumptions a standard A/B test rests on.
Their headline result, on female impression share for the high-bid, top-percentile bracket:
| Channel | Effect on female impression share | Interpretation |
|---|---|---|
| Algorithmic (Arm 2 vs Arm 1) | +2.07 pp | Metadata alone reallocates delivery |
| Creative (Arm 3 vs Arm 2) | −0.68 pp | Visible women-targeted text actually opposes the algorithmic move |
| Total effect | +1.39 pp | Net observed shift |
| Share of absolute reallocation that is algorithmic | ~75% | |2.07| / (|2.07| + |0.68|) |
Two things fall out of that table that are worth reading carefully.
First, the algorithm is doing most of the work. Even in a design where the ‘creative treatment’ is a literal ‘Women applicants encouraged’ text overlay, the algorithm's reallocation swamps what users see.
Second, the creative channel can oppose the algorithmic channel. The visible ‘women-targeted’ text actually reduced female impression share compared to the metadata-only condition. A two-arm A/B test that compares Arm 3 to Arm 1 measures the net effect (+1.39 pp) and cannot see that the two channels are pulling in opposite directions. The buyer sees a positive verdict and concludes the creative worked. What actually happened is the algorithm did the work and the creative undid a chunk of it.
The 7-day, single-campaign pilot is underpowered at p=0.05. The 2 pp algorithmic effect is real in the point estimate but the confidence intervals are wide. The mechanism and direction are the load-bearing contributions; the exact 75% share should be read as directional, not final. What is not in doubt is that the three-arm design point-identifies channels a standard two-arm test conflates.
The problem is not new
Reading the Pal and Susarla paper in isolation is misleading. It is the sharpest recent statement of a problem the field has been quantifying for at least seven years.
2019. Gordon, Zettelmeyer, Bhargava, and Chapsky publish ‘A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook’ in Marketing Science. Fifteen U.S. field experiments at Facebook. 500 million user-experiment observations. 1.6 billion ad impressions.[4] Their finding: observational advertising-measurement methods routinely fail to match the ground-truth causal estimate from randomized experiments. The tools most media buyers use overstate advertising effectiveness. The problem has a name (measurement bias). Two of the co-authors were Facebook Research Scientists at the time.
January 2025. Braun and Schwartz publish ‘Where A/B testing goes wrong’ in the Journal of Marketing.[3] They name the mechanism: divergent delivery. Platform A/B testing tools serve ads to different, algorithmically-optimized user mixes. Structural, not fixable by better test design. Peer-reviewed at the top marketing journal.
December 2025. Burtch, Moakler (Meta), Gordon (Kellogg, also on the 2019 paper), Zhang (Meta), and Hill publish the 181,890-A/B-test census. 22% confounded on A/B, 0.16% on Lift. Meta co-authors on the byline.[1]
May 2026. Pal and Susarla publish the three-arm causal decomposition. ~75% of the observed audience reallocation is algorithmic. Two-arm tests understate the algorithmic channel by 2x.[2]
Four papers, three of them peer-reviewed or under peer review, one preprint, all four converging on the same finding: the A/B test tool Meta ships to advertisers does not cleanly measure creative quality. It measures a confound. Meta's own scientists have co-authored two of the four papers.
The tool has not been fixed. Advertisers still open Ads Manager, launch an A/B test, and act on the verdict.
Why the pre-launch tools shipping this year do not fully fix this
The industry is starting to respond. Two moves in the last few months.
Brunner acquired AdSkate on July 16, 2026.[8]AdSkate is a pre-launch AI creative testing tool that uses synthetic-audience personas to estimate creative performance before campaigns go live. Their pitch: sidestep both the media cost of a live test and the data-privacy constraints of a real-user panel. Testing moves off Meta's auction and into a synthetic simulation.
Toluna (formerly MetrixLab) ships ACT Instant. AI-powered 24-hour ad-copy pre-testing tool that predicts ad performance using ML models coding ads across 130+ variables. No surveys, no respondents.[9]
Both are directionally correct. Both move the creative-choice question off Meta's delivery-confounded A/B tool and into a pre-launch simulation. That is the right move. It is the same category Splitroom operates in.
Neither, however, publishes a claim that they isolate creative from delivery in a causal sense. If the ML models that power AdSkate's or Toluna's predictions were trained on historical Meta A/B test winners, they inherit the same confound Burtch et al. quantified: the ‘winning’ creatives in the training set were often winners because Meta's algorithm reallocated to them, not because the creative itself was preferred by users. A pre-launch tool that treats a confounded post-launch signal as ground truth is a faster version of the same problem.
The methodological question the pre-launch category has to answer is whether the synthetic-audience prediction is trained against clean, delivery-uncontaminated signal or against Meta's confounded A/B verdicts. Neither vendor has publicly disclosed the answer.
What actually isolates creative from delivery
Splitroomjudges creative choice off-platform, before Meta's delivery algorithm gets involved. Two creatives go in. Up to 1,000 synthetic buyers evaluate both, in parallel, against a symmetric scorecard. Every panelist scores both ads against the same canonical dimensions in the same order. There is no algorithm making choices about which panelist sees which ad more.
The confound Braun, Burtch, and Pal all quantify cannot happen inside a Splitroom simulation, because the structural precondition (an algorithm allocating impressions) does not exist. Every panelist evaluates both creatives. Every scorecard uses the same dimensions in the same order. Nothing in the pipeline can systematically show one creative to a different mix than the other.
That is not the same as claiming synthetic panels perfectly predict real-buyer behavior. They do not, and we have written about the peer-reviewed limits in ‘What synthetic panels can and can not do.’ What is true is the specific structural claim: the delivery-vs-creative confound cannot enter a Splitroom judgment by construction.
For the on-platform question of whether Meta advertising works at all versus not advertising, use Meta's Lift tests. They show 0.16% audience imbalance in Burtch et al.'s data — that is the clean tool Meta itself points serious causal-measurement customers toward.[1] For the question of which creative to pick, run the pre-launch judgment off-platform. That is Splitroom.
What this means for a media buyer today
Four practical calls that fall directly out of the research.
Stop treating Meta A/B test verdicts as clean creative signal. They are not. The tool Meta's own research scientists co-authored papers about is delivery-confounded 22% of the time at the standard academic threshold, and structurally confounded 100% of the time by the mechanism Braun and Schwartz named in the Journal of Marketing.[3]Continuing to double-down media budget on Meta A/B winners without understanding the confound is funding whatever Meta's early-delivery signal favored.
Use Meta Lift tests for the incremental-effect question.For the question ‘is my Meta advertising working at all,’ Lift tests with a true holdout have 0.16% audience imbalance in Burtch's data. That is a valid measurement. Do not use Lift tests to compare creatives. That is not what they are built for.
Move the creative-vs-creative comparison off-platform. Pre-launch judgment is now a defensible category, backed by peer-reviewed literature naming the confound the on-platform tool has. Options exist. Splitroom (synthetic panel), PickFu (real-human panel), AdSkate, Toluna. Split your measurement problem into two questions each tool actually answers.
Ask any pre-launch vendor how they trained their models.If the answer is ‘on historical A/B test winners’ without explicit correction for divergent delivery, you are getting a faster version of the confound, not a fix. Vendors that isolate creative from delivery structurally (no auction inside the tool) have a stronger claim than vendors whose predictions are learned from delivery-contaminated ground truth.
The science is settled. The tool has not been.
Four papers, seven years, one converging finding. Gordon et al. 2019 warned that observational ad measurement is broken. Braun and Schwartz named the mechanism in a peer-reviewed 2025 Journal of Marketing article. Burtch et al., with two Meta Research Scientists on the byline, quantified the confound across 181,890 real Meta A/B tests. Pal and Susarla causally decomposed it into a ~75% algorithmic share.
Meta has not shipped a new A/B test tool that isolates the creative from the delivery. Meta's own scientists co-authored the paper documenting the confound and pointed customers toward Lift tests instead. Ads Manager still shows the same A/B verdict interface it did five years ago. Media buyers still act on the numbers it returns.
The gap between what the research knows and what the tool measures is where every Meta media buyer is currently making decisions.
You did not A/B test your creative. You A/B tested Meta's algorithm. Now you know.
Fair questions
What does 'divergent delivery' actually mean?
Divergent delivery is what happens when Meta's algorithm decides on its own to show two ads in the same A/B test to systematically different mixes of users. In a clean A/B test, the two variants should reach interchangeable audiences so the only thing that differs is the creative. In a Meta A/B test, the delivery algorithm uses early signal (predicted CTR, predicted conversion likelihood) to allocate impressions to whichever variant it thinks will perform better. That reallocation happens before either variant has enough impressions for statistical significance. The verdict looks like 'Ad A won,' but what actually happened is Ad A was shown to users predicted to convert more. The creative is confounded with the algorithm's targeting decision. Braun and Schwartz named this in the Journal of Marketing in January 2025. Burtch, Moakler, Gordon, Zhang, and Hill quantified it across 181,890 real Meta A/B tests in December 2025.
How much does Meta's delivery algorithm actually confound my A/B test results?
Two studies give two answers. Burtch et al. (2025), analyzing 181,890 real Meta A/B tests with two Meta Research Scientists among the authors, found that 22% of those tests show audience imbalance with a standardized mean difference greater than 0.2 — the conventional academic threshold for a compromised experimental comparison. In Meta's own Lift tests, which use a true no-ad holdout, that number is 0.16%. That is a 137-fold gap. Pal and Susarla (2026), using a three-arm experimental design on a live Meta field campaign, causally decomposed the effect: approximately 75% of the audience reallocation credited to a 'winning' creative was attributable to the delivery algorithm, not the creative itself. A conventional two-arm A/B test understates the algorithmic channel by roughly a factor of two.
Are Meta's own Lift tests reliable?
More reliable than A/B tests, yes. Meta Lift tests use a randomized no-ad holdout group and measure the incremental effect of advertising versus no advertising. Because both groups are randomly selected and neither is subject to algorithmic reallocation between test arms, the audience-imbalance rate is 0.16% at the SMD > 0.2 threshold, versus 22% for A/B tests, per Burtch et al.'s analysis of 181,890 A/B tests and 3,204 Lift tests. The catch: Lift tests measure the effect of your ads versus no ads, not the effect of one creative versus another. They tell you whether Meta advertising works. They do not tell you which of two creatives to pick. For creative choice, Meta's own tool architecture leaves the media buyer with an A/B test that Meta's own research scientists have co-authored papers documenting as confounded.
Why don't AdSkate and Toluna solve this?
They partially solve it by moving pre-launch. AdSkate (acquired by Brunner in July 2026) uses synthetic-audience personas to estimate creative performance before campaigns go live. Toluna's ACT Instant predicts ad performance in 24 hours without surveys or respondents. Both are legitimate improvements on running the test on Meta and getting a delivery-confounded verdict. But neither claims to isolate creative from delivery in a causal sense. They avoid the confound by testing off-platform against synthetic panels. Whether the synthetic panels themselves inherit delivery bias depends on how their models were trained: if the training data includes historical A/B test winners (which are themselves confounded per Burtch et al.), the confound gets baked into the pre-launch prediction. That is a real methodological caveat neither vendor has publicly addressed.
Is post-launch testing on Meta ever reliable?
For measuring incremental lift versus no ads, yes: use Meta Lift tests with a true holdout, which show 0.16% audience imbalance in the Burtch study versus 22% for standard A/B tests. For measuring one creative versus another, no. The tool architecture Meta ships to advertisers for creative-vs-creative comparison — the standard A/B test — is the one the peer-reviewed and preprint literature has now documented as delivery-confounded. Media buyers who want a clean creative-vs-creative signal have to test off-platform (real-human panels like PickFu, synthetic panels like Splitroom) and then use the on-platform test only for lift-vs-no-ad measurement. That splits the measurement problem into two questions each tool is actually built to answer.
What does Splitroom do differently?
Splitroom judges creative choice off-platform, before Meta's delivery algorithm gets involved. Two creatives go in. Up to 1,000 synthetic buyers evaluate both, in parallel, against a symmetric scorecard. No auction. No Learning Phase. No delivery decisions. Every panelist scores both ads against the same canonical dimensions in the same order, so nothing in the pipeline can systematically show one ad to a different user mix than the other. The confound that breaks Meta's A/B test cannot happen because there is no algorithm making delivery choices. That is not the same as claiming synthetic panels perfectly predict real-buyer behavior (they do not — see 'What synthetic panels can and can not do' for the peer-reviewed limits). It is the specific structural claim that the delivery-vs-creative confound cannot enter a Splitroom simulation by construction.
Sources
- Divergent Delivery in Digital Advertising: Evidence from 181,890 Meta A/B Tests (MSI Working Paper 25-140) · Burtch, Moakler (Meta), Gordon, Zhang (Meta), Hill — Marketing Science Institute / arXiv preprint, Dec 2025 · retrieved 2026-08-14
- Algorithm or Creative? A Three-Arm Experimental Design for Decomposing Algorithmic Bias in Platform A/B Tests · Pal (Stevens Institute), Susarla (Michigan State) — arXiv preprint, May 2026 · retrieved 2026-08-14
- Where A/B testing goes wrong: How divergent delivery affects experiments on digital-advertising platforms · Braun, Schwartz — Journal of Marketing, Jan 2025 · retrieved 2026-08-14
- A Comparison of Approaches to Advertising Measurement: Evidence from Big Field Experiments at Facebook · Gordon, Zettelmeyer, Bhargava, Chapsky — Marketing Science 38(2), 2019 (500M user obs, 1.6B ad impressions) · retrieved 2026-08-14
- Meta Ad Testing Demystified: Divergent Delivery and What It Means for Your Results · Marketing Science Institute (Moakler + Zhang from Meta, Gordon + Burtch) · retrieved 2026-08-14
- Creative Benchmarks 2026 (578,750 ads, 6,015 accounts, ~$1.29B Meta spend) · Motion · retrieved 2026-08-14
- Facebook Ads Benchmarks 2025 (US CPL $27.66) · WordStream · retrieved 2026-08-14
- Brunner buys AdSkate as independent agency bets on AI creative testing · PPC Land · retrieved 2026-08-14
- AI-powered 24-hour ad-copy pre-testing (ACT Instant) · Toluna (formerly MetrixLab) · retrieved 2026-08-14
- Meta Q4 2025 earnings ($58.14B ad revenue, +24.3% YoY) · Meta Investor Relations · retrieved 2026-08-14
Two creatives go in. Up to a thousand synthetic consumers argue it out, before a dollar of media moves.
Free in early access · No card required