Why Splitroom judges the scroll-stop first.
Splitroom's verdict engine gates every duel on one question before any other: did the creative earn a second look. The eye-tracking and ad-effectiveness research behind that order.

Splitroom now scores every duel in a fixed order. First: did the creative earn a second look. That's a gate, not a tiebreaker. Second: creative quality and messaging, scored together, because the research doesn't cleanly separate them. Third: persona-specific priorities, like how much a given segment weighs trust cues versus price.
That order isn't a preference. It's what the published research on attention and creative effectiveness actually supports. Below is the sourced case for it, complications included.
What made us build a gate
A synthetic panel duel returned a 93-7 landslide for one creative.
The losing creative was, separately, the one more likely to actually get looked at in a feed. Its headline read faster. Its focal point resolved sooner. A repeated first-glance check on both images called it the clear winner of the race to be noticed at all.
And in a live campaign running the same two creatives, the real click data was tracking the "loser."
That's not a contradiction the panel's reasoning could talk its way out of. Give a synthetic consumer ten seconds to study two ads side by side, and it will find real, defensible reasons to prefer the one with better trust badges, tighter copy, a more premium mood.
A real scroller never grants those ten seconds to the ad that didn't stop the thumb in the first place. The panel was answering "which ad wins an argument." The question that actually decides ad performance is "which ad gets heard at all."
What the research says the order should be
The foundational eye-tracking work here is Pieters and Wedel's 2004 study in the Journal of Marketing. They tracked gaze across 1,363 print ads and more than 3,600 consumers.[1]
Their finding: the pictorial element captures attention first, largely independent of its size. Text captures attention in proportion to how much space it takes up. The brand element's job is to transfer attention it already received to everything else on the page.
Translated to a gate: something has to win the first fraction of a second before copy, trust badges, or brand get read at all.
That study is 20 years old and about print. The closest thing to a direct feed-scrolling replication is Mayer, Ohme, Maslowska, and Segijn's 2024 eye-tracking study in Social Media + Society. They put 201 participants through a real, scrollable Facebook newsfeed and measured gaze on desktop versus mobile, and in a quiet lab versus a busy cafeteria.[2]
It complicates the simple "pictures win" story. On mobile, participants spent less dwell time and fewer fixations on the picture than on desktop (d = .44 to .59), and significantly more on text. In the public, distraction-heavy condition — the one that actually resembles scrolling a feed on the go — attention to text rose again, not fell.
We could cite only the convenient half of this paper: pictures grab attention first. We're not doing that.
What survives across both studies, print and mobile feed, isn't "images always win." It's narrower: something in the ad, whichever element it is for that specific execution, has to win a first, fast, low-effort pass before anything requiring deliberate reading gets attempted. That's the claim behind gating on scroll-stop. Not a claim about which element should carry it.
Kantar's own attention research draws the same two-stage line from a different angle. Their 2023 framework splits attention into passive attention (eyes on screen) and active attention (measurable emotional engagement), built on a database of over 50,000 ads.[3]
Their data on feed and short-form formats — Facebook and Instagram feed, Stories, TikTok — shows passive attention diminishing fastest of any format in the first seconds of exposure. Their own conclusion: these formats need clear communication within the initial seconds, because that window is shorter than in formats people can't skip past.
A synthetic panel scoring an ad has none of that time pressure. It reads calmly. A gate is how we put the time pressure back in.
Then creative quality and messaging, together
Once a creative clears the scroll-stop gate, the next and largest factor is creative quality and messaging. We score it as one bundle, not split into separate "visuals" and "copy" scores.
NCSolutions' 2023 update to their Five Keys to Advertising Effectiveness study, based on nearly 450 CPG campaigns, found creative responsible for 49% of incremental sales. The single largest of the five factors they measure, unchanged from their original 2017 study.[4] Restricted to social media specifically, creative's share was 46%.
| Factor | All media | Social media |
|---|---|---|
| Creative (quality + messaging) | 49% | 46% |
| Brand (loyalty, penetration, share) | 21% | 26% |
| Targeting + reach + recency | 30% | 28% |
Why not split it into a "visuals score" and "copy score"? Because NCSolutions' own methodology doesn't split it. Their creative-quality figure is one bundled measure. No published study we found decomposes it cleanly into a visual-only and copy-only share without conflating the two. A striking headline is often what makes a viewer look back at the image, and vice versa.
Scoring them independently would manufacture a precision the research doesn't have.
Then persona-specific weighting
The last tier is one Splitroom already had built. Each synthetic panelist carries its own priority weights across trust cues, price sensitivity, emotional tone, and more, generated from an archetype brief with enforced variance so the panel doesn't collapse into one homogenized voter.
This tier answers "which specific thing about this creative moved this specific kind of buyer." It's the tier most responsible for segment-level dissent in a report — the finding that a headline wins the aggregate vote but loses inside your highest-value segment.
The honest caveat: persona conditioning is real but bounded. Hu and Collier's 2024 ACL paper found persona conditioning explains less than 10% of variance in subjective LLM annotations, even with explicit prompting.[5]
That doesn't make persona-level dissent noise. It means a single persona's stated reasoning should carry less confidence than an aggregate finding across the whole panel. That's exactly why this tier sits third, and why segment dissent is reported as a lean, not a separate verdict.
Making the gate statistically real
A single vision-model call judging "which creative wins the scroll-stop" has no more statistical standing than asking one person.
Zheng et al.'s NeurIPS 2023 benchmark, the canonical paper on using LLMs as judges, found GPT-4 flips its verdict 35% of the time on the same pair of candidates when their A/B labels are swapped.[6] One call is genuinely unreliable at the single-instance level.
GPT-4 flips its verdict on identical candidates 35% of the time when the A/B labels are swapped.Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023
The same paper is also the strongest evidence that aggregating those noisy calls works. GPT-4's judgments agreed with human raters 85% of the time on MT-Bench, against an 81% human-human baseline.
Argyle et al. found the same pattern for full synthetic panels. A well-conditioned LLM sample tracked real ANES respondents' vote choices at tetrachoric correlations of 0.90 to 0.94 across three election waves, even though no single simulated respondent's answer was trustworthy on its own.[7]
Noisy at the single-call level, stable in aggregate. Same principle a polling sample runs on.
So the scroll-stop read is no longer a single call. It runs independently multiple times — the count scales with panel size, 5% of the panel, floored at 2, capped at 15 — and the gate fires off the majority vote. The vote split itself is surfaced on the report.
A near-even split, say 7 of 13, is now visible as exactly what it is: these two creatives don't have a clean scroll-stop winner. That's a real, honest finding. Not a confident-sounding single opinion papering over genuine ambiguity.
When the gate and the panel disagree
When the scroll-stop majority disagrees with the duel's actual winner, the report carries a named warning. Which creative won the scroll-stop race, and by what margin. Which creative won the head-to-head comparison. A plain-language note that a comparison win means little if the losing side is the one that actually gets looked at in a real feed.
It's not a silent adjustment to the score. The panel's verdict stands as reported. The mismatch sits next to it, in the open, on the report page — not buried in a log file.
That distinction matters. The gate is designed to make disagreement visible, not to override a verdict quietly on the customer's behalf.
What's still open
Two things we're not pretending to have solved.
First, exactly how hard the gate should bite. Right now a scroll-stop loss produces a warning next to the verdict, not an automatic score adjustment. Whether a weak or tied result should discount the verdict proportionally, or only a clear loss should trigger any adjustment at all, is a question we're leaving to calibrate against real outcome data over time. Not to guess at from research that wasn't run on ad creative specifically.
Second, none of the studies above measured actual downstream conversion. They measured attention and aggregate sales lift, not click-through on the exact ad formats Splitroom tests. The gate is built on the best published evidence for the underlying mechanism — attention has to happen before persuasion can — not on a study that ran this exact experiment.
The one-liner: a synthetic panel that's allowed to study an ad calmly for ten seconds will always find something to like about the ad that lost the real feed. The scroll-stop gate is the part of the pipeline that refuses to let it forget the ten seconds were never real.
Fair questions
What does 'scroll-stop gate' mean in a Splitroom verdict?
The first-glance engine reads which creative wins the race to be noticed in a feed. That read gets checked against the panel's head-to-head verdict. If the two disagree, the report carries a named warning right next to the verdict. It's a gate, not a silent score adjustment. The panel's vote still stands as reported. But a creative that lost the race to be looked at gets flagged when it also won the comparison, because a synthetic panel that studies both ads calmly for ten seconds can find real reasons to prefer an ad a real scroller never stayed to read.
Why weight scroll-stop above creative quality and messaging?
Because attention is a precondition for everything downstream of it, not a competing factor to average in. Pieters and Wedel's 2004 eye-tracking study of 1,363 print ads found the dominant element captures attention first, largely independent of its size, before text or brand elements get read at all. Mayer et al.'s 2024 mobile-feed eye-tracking study extends this to social feeds directly. If nothing in an ad earns that first fast pass, the messaging behind it never gets evaluated by a real scroller. Scoring message quality without checking whether the ad earns the chance to be read would overweight ads that never get that chance in market.
Why does Splitroom score creative quality and messaging together instead of separately?
Because the research doesn't cleanly separate them. NCSolutions' 2023 Five Keys to Advertising Effectiveness study, based on nearly 450 CPG campaigns, measures 'creative' as one bundled factor responsible for 49% of incremental sales, 46% for social media specifically. It doesn't decompose that into an independent visual-only and copy-only share. A striking headline is often what makes a viewer look back at an image. A strong image is often what makes a viewer read a headline. Treating them as cleanly separable would manufacture a precision the research doesn't support.
Isn't a single AI model's read of 'which ad stops the scroll' just a guess?
A single read is unreliable. Zheng et al.'s NeurIPS 2023 benchmark found GPT-4 flips its verdict 35% of the time on identical candidates when their labels are swapped. That's why Splitroom no longer asks once. The scroll-stop read now runs independently multiple times, roughly 5% of the panel size, floored at 2 and capped at 15, and the gate fires on the majority vote, with the vote split itself shown on the report. A near-even split is surfaced honestly as a real finding: these two creatives don't have a clean scroll-stop winner. The same paper found GPT-4 agrees with human raters 85% of the time in aggregate, against an 81% human-human baseline. Noisy at the single-call level, stable in aggregate. Same principle a polling sample runs on.
Does the scroll-stop gate override the synthetic panel's vote?
No. The panel's head-to-head verdict is reported exactly as the panel returned it. The gate adds a visible warning when the scroll-stop majority disagrees with that verdict, naming which creative won the scroll-stop race, by what margin, and what that means for trusting the headline number. It doesn't silently adjust the score. The goal is to put a real disagreement between two different instruments, a calm comparative read and a fast first-glance read, in front of the person making the decision, not to have one instrument quietly overrule the other.
What hasn't this methodology settled yet?
Two open questions, stated plainly rather than guessed at. First, how much a scroll-stop loss should discount a verdict. Right now it produces a warning, not an automatic score penalty. Whether a weak or tied result should discount proportionally, or only a clear loss should trigger any adjustment, is left to calibrate against real report-back data over time. Second, none of the cited research measured downstream click-through on the exact ad formats Splitroom tests. It measured attention allocation and aggregate sales lift. The gate is built on the strongest available evidence for the underlying mechanism, that attention has to happen before persuasion can, not on a study that ran this precise experiment on this precise product.
Sources
- Attention Capture and Transfer in Advertising: Brand, Pictorial, and Text-Size Effects (1,363 print ads, 3,600+ consumers eye-tracked) · Pieters, Wedel — Journal of Marketing, Vol. 68, No. 2, 2004, pp. 36-50 · retrieved 2026-08-23
- Headlines, Pictures, Likes: Attention to Social Media Newsfeed Post Elements on Smartphones and in Public (N=201 eye-tracking experiment, desktop vs. mobile, private vs. public) · Mayer, Ohme, Maslowska, Segijn — Social Media + Society, Vol. 10, Issue 2, April 2024 · retrieved 2026-08-23
- Beyond Viewability: The Role of Attention in Creative Effectiveness (passive vs. active attention framework, 50,000+ ad normative database) · Kantar — Ecem Erdem, November 2023 · retrieved 2026-08-23
- Five Keys to Advertising Effectiveness, 2023 update (~450 CPG campaigns; creative = 49% of incremental sales, 46% for social media specifically) · NCSolutions, 2023 e-book · retrieved 2026-08-23
- Quantifying the Persona Effect in LLM Simulations (persona conditioning explains less than 10% of variance in subjective annotations even with prompting) · Hu & Collier — ACL 2024, Long Papers · retrieved 2026-08-23
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (GPT-4 flips verdict 35% of the time on swapped A/B labels; 85% agreement with human raters vs. 81% human-human baseline) · Zheng et al., NeurIPS 2023 · retrieved 2026-08-23
- Out of One, Many: Using Language Models to Simulate Human Samples (GPT-3 matched real ANES vote choices at tetrachoric correlations of 0.90-0.94 across three election waves) · Argyle, Busby, Fulda, Gubler, Rytting, Wingate — Political Analysis 2023 · retrieved 2026-08-23
Your ad creatives go in, up to six at once. A thousand synthetic consumers argue it out, before a dollar of media moves.
Free in early access · No card required