← Blog11 min read

Why Splitroom judges the scroll-stop first.

Splitroom's verdict engine gates every duel on one question before any other: did the creative earn a second look. The eye-tracking and ad-effectiveness research behind that order.

SA
Syed Asif Sultan
Founder, Splitroom
The short answer

Splitroom's verdict engine now scores every duel in a fixed order: whether the creative earns a second look at feed size comes first, as a gate, not a tiebreaker. Creative quality and messaging, scored together because the research doesn't cleanly separate them, comes second. Persona-specific priorities — how much a given segment weighs trust cues versus price versus tone — come third. That order isn't a design preference. It's the order the published research on attention and creative effectiveness actually supports, and this piece is the sourced case for it, including the parts that complicate the story.

What made us build a gate instead of a footnote

A synthetic panel duel had just returned a 93-7 landslide for one creative. The losing creative was, separately, the one independently identified as more likely to actually get looked at in a feed — its headline read faster, its focal point resolved sooner, and a repeated first-glance check on both images called it the clear winner of the race to be noticed at all. And in a live campaign running the same two creatives, the real click data was tracking the "loser."

That's not a contradiction the panel's reasoning could talk its way out of. A synthetic consumer given ten seconds to study two ads side by side will find real, defensible reasons to prefer the one with better trust badges, tighter copy, or a more premium mood. A real scroller never grants those ten seconds to the ad that didn't stop the thumb in the first place. The panel was answering "which ad wins an argument," and the question that actually determines ad performance is "which ad gets heard at all."

The order isn't arbitrary: what the research says

The foundational eye-tracking work on this is Pieters and Wedel's 2004 study in the Journal of Marketing, which tracked gaze across 1,363 print ads and more than 3,600 consumers.[1] Their finding: the pictorial element captures attention first and largely independent of its size, text captures attention in proportion to how much space it takes up, and the brand element's job is to transfer attention it's already received to everything else on the page. Translated to a gate: something has to win the first fraction of a second before anything else — copy, trust badges, brand — gets read at all.

That finding is 20 years old and about print. The closest thing to a direct feed-scrolling replication is Mayer, Ohme, Maslowska, and Segijn's 2024 eye-tracking study in Social Media + Society, which put 201 participants through a real, scrollable Facebook newsfeed and measured gaze on desktop versus mobile, and in a quiet lab versus a busy cafeteria.[2] It complicates the simple "pictures win" story in a useful way: on mobile, participants spent significantly less dwell time and fewer fixations on the picture than on desktop (d = .44 to .59), and significantly more on text elements. In the public, distraction-heavy condition — the one that actually resembles someone scrolling a feed on the go — attention to text-carrying elements rose again, not fell.

What we take from a study that doesn't confirm the simple version

We could have cited only the part of this paper that's convenient — pictures grab attention first — and left out the part that complicates it. We're not doing that. What survives across both studies, print and mobile feed, isn't "images always win." It's narrower and more useful: something in the ad, whichever element it is for that specific execution, has to win a first, fast, low-effort pass before anything requiring deliberate reading gets attempted. That's the actual claim behind gating on scroll-stop — not a claim about which element should carry it.

Kantar's own attention research draws the same two-stage line from a different angle. Their 2023 framework splits attention into passive attention (eyes on screen) and active attention (measurable emotional engagement), built on a normative database of over 50,000 ads.[3] Their data on feed and short-form formats specifically (Facebook and Instagram feed, Stories, TikTok) shows passive attention diminishing fastest of any format in the first seconds of exposure — their stated implication is that these formats need clear communication "within the initial seconds," because that window is shorter than in formats people can't skip past. A synthetic panel scoring an ad has none of that time pressure. It reads calmly. A gate is how we put the time pressure back in.

Then creative quality and messaging, together

Once a creative clears the scroll-stop gate, the next and largest factor is creative quality and messaging — scored as one bundle, not split into separate "visuals" and "copy" scores. NCSolutions' 2023 update to their Five Keys to Advertising Effectiveness study, based on nearly 450 CPG campaigns, found creative responsible for 49% of incremental sales — the single largest of the five factors they measure, unchanged from their original 2017 study.[4] Restricted to social media specifically, creative's share was 46%.

What drives incremental sales — NCSolutions Five Keys, 2023 (n≈450 CPG campaigns)
FactorAll mediaSocial media
Creative (quality + messaging)49%46%
Brand (loyalty, penetration, share)21%26%
Targeting + reach + recency30%28%
Source: NCSolutions, Five Keys to Advertising Effectiveness, 2023 e-book. [4]

The reason we don't split this into a separate "visuals score" and "copy score" is that NCSolutions' own methodology doesn't split it — their creative-quality figure is a single bundled measure of "creative message," and no published study we found decomposes it cleanly into a visual-only and copy-only share without conflating the two (a striking headline is often the thing that makes a viewer look back at the image, and vice versa). Pretending to score them independently would manufacture a precision the underlying research doesn't have.

Then persona-specific weighting

The last tier is the one Splitroom already had built before this change: each synthetic panelist carries its own priority weights across trust cues, price sensitivity, emotional tone, and so on, generated from an archetype brief with enforced variance so the panel doesn't collapse into one homogenized voter. This tier answers "which specific thing about this creative moved this specific kind of buyer," and it's the one most responsible for segment-level dissent in a report — the finding that a headline wins the aggregate vote but loses inside your highest-value segment.

The honest caveat on this tier: persona conditioning is real but bounded. Hu and Collier's 2024 ACL paper found that persona conditioning explains less than 10% of variance in subjective LLM annotations even with explicit prompting.[5] That doesn't mean persona-level dissent in a Splitroom report is noise — it means the confidence we put on any single persona's stated reasoning should be lower than the confidence we put on an aggregate finding across the whole panel, which is exactly why this tier sits third and why segment dissent is reported as a lean, not a separate verdict.

Making the gate statistically real, not one model's opinion

A single vision-model call judging "which creative wins the scroll-stop" has no more statistical standing than asking one person. Zheng et al.'s NeurIPS 2023 benchmark, the canonical paper on using LLMs as judges, found GPT-4 flips its verdict 35% of the time on the same pair of candidates when their A/B labels are swapped.[6] One call is genuinely unreliable at the single-instance level.

GPT-4 flips its verdict on identical candidates 35% of the time when the A/B labels are swapped.
Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena, NeurIPS 2023

The same paper is also the strongest evidence that aggregating those noisy single calls works: GPT-4's judgments agreed with human raters 85% of the time on MT-Bench, against an 81% human-human baseline. That's the same logic Argyle et al. established for full synthetic panels — a well-conditioned LLM sample tracked real ANES respondents' vote choices at tetrachoric correlations of 0.90 to 0.94 across three election waves, even though no individual simulated respondent's answer was trustworthy on its own.[7] Noisy at the single-call level, stable in aggregate — that's the same principle a polling sample runs on.

So the scroll-stop read is no longer a single call. It now runs repeated, independent times — the count scales with panel size (5% of the panel, floored at 2, capped at 15) — and the gate fires off the majority vote, with the vote split itself surfaced on the report. A near-even split (say 7 of 13) is now visible as exactly what it is: these two creatives don't have a clean scroll-stop winner, which is a real, honest finding, not a confident-sounding single opinion papering over genuine ambiguity.

What happens when the gate and the panel disagree

When the scroll-stop majority disagrees with the duel's actual winner, the report now carries a named warning: which creative won the scroll-stop race and by what margin, which creative won the head-to-head comparison, and a plain-language note that a comparison win means little if the losing side of that comparison is the one that actually gets looked at in a real feed. It's not a silent adjustment to the score — the panel's verdict stands as reported, and the mismatch sits next to it in the open, on the report page, not buried in a log file. That distinction matters: the gate is designed to make disagreement visible, not to override a verdict quietly on the customer's behalf.

What's still open

Two things we're not pretending to have solved. First, exactly how hard the gate should bite — right now a scroll-stop loss produces a warning next to the verdict, not an automatic score adjustment. Whether a weak or tied scroll-stop result should discount the verdict proportionally, or whether only a clear loss should trigger any adjustment at all, is a question we're deliberately leaving to calibrate against real outcome data over time, not to guess at from research that wasn't run on ad creative specifically. Second, none of the studies above measured actual downstream conversion — they measured attention and aggregate sales lift, not click-through on the specific ad format Splitroom tests. The gate is built on the best published evidence that exists for the underlying mechanism (attention has to happen before persuasion can), not on a study that ran this exact experiment.

The one-liner: a synthetic panel that's allowed to study an ad calmly for ten seconds will always find something to like about the ad that lost the real feed. The scroll-stop gate is the part of the pipeline that refuses to let it forget the ten seconds were never real.

Fair questions

What does 'scroll-stop gate' mean in a Splitroom verdict?

It means the first-glance engine's read of which creative wins the race to be noticed in a feed — the scroll-stop race — is checked against the panel's head-to-head verdict, and if the two disagree, the report carries a named warning right next to the verdict. It is a gate, not a silent score adjustment: the panel's vote still stands as reported, but a creative that lost the race to be looked at in the first place gets flagged when it also won the comparison, because a synthetic panel that studies both ads calmly for ten seconds can find real reasons to prefer an ad a real scroller never stayed to read.

Why weight scroll-stop above creative quality and messaging?

Because attention is a precondition for everything downstream of it, not a competing factor to average in. Pieters and Wedel's 2004 eye-tracking study of 1,363 print ads found the pictorial or dominant element captures attention first, largely independent of its size, before text or brand elements get read at all. Mayer et al.'s 2024 mobile-feed eye-tracking study extends this to social feeds directly. If nothing in an ad earns that first fast pass, the quality of the messaging behind it never gets evaluated by a real scroller — so scoring message quality without first checking whether the ad earns the chance to be read would overweight ads that never get that chance in market.

Why does Splitroom score creative quality and messaging together instead of separately?

Because the research doesn't cleanly separate them. NCSolutions' 2023 Five Keys to Advertising Effectiveness study, based on nearly 450 CPG campaigns, measures 'creative' as a single bundled factor responsible for 49% of incremental sales (46% for social media specifically) and does not decompose it into an independent visual-only and copy-only share. A striking headline is often what makes a viewer look back at an image, and a strong image is often what makes a viewer read a headline — treating them as cleanly separable would manufacture a precision the underlying research doesn't support.

Isn't a single AI model's read of 'which ad stops the scroll' just a guess?

A single read is unreliable — Zheng et al.'s NeurIPS 2023 benchmark found GPT-4 flips its verdict 35% of the time on identical candidates when their labels are swapped. That's why Splitroom no longer asks once. The scroll-stop read now runs independently multiple times (roughly 5% of the panel size, floored at 2 and capped at 15), and the gate fires on the majority vote, with the vote split itself shown on the report. A near-even split is surfaced honestly as a real finding — these two creatives don't have a clean scroll-stop winner — instead of being hidden behind one confident-sounding call. The same paper found GPT-4 agrees with human raters 85% of the time in aggregate, against an 81% human-human baseline: noisy at the single-call level, stable in aggregate, which is the same principle a polling sample runs on.

Does the scroll-stop gate override the synthetic panel's vote?

No. The panel's head-to-head verdict is reported exactly as the panel returned it. The gate adds a visible warning when the scroll-stop majority disagrees with that verdict — naming which creative won the scroll-stop race, by what margin, and what that means for trusting the headline number — rather than silently adjusting the score. The goal is to put a real disagreement between two different instruments (a calm comparative read and a fast first-glance read) in front of the person making the decision, not to have one instrument quietly overrule the other.

What hasn't this methodology settled yet?

Two open questions, stated plainly rather than guessed at. First, exactly how much a scroll-stop loss should discount a verdict — right now it produces a warning, not an automatic score penalty, and whether a weak or tied result should discount proportionally versus only a clear loss triggering any adjustment is left to calibrate against real report-back data over time. Second, none of the cited research measured downstream click-through on the exact ad formats Splitroom tests; it measured attention allocation and aggregate sales lift. The gate is built on the strongest available evidence for the underlying mechanism — attention has to happen before persuasion can — not on a study that ran this precise experiment on this precise product.

Sources

  1. Attention Capture and Transfer in Advertising: Brand, Pictorial, and Text-Size Effects (1,363 print ads, 3,600+ consumers eye-tracked) · Pieters, Wedel — Journal of Marketing, Vol. 68, No. 2, 2004, pp. 36-50 · retrieved 2026-08-23
  2. Headlines, Pictures, Likes: Attention to Social Media Newsfeed Post Elements on Smartphones and in Public (N=201 eye-tracking experiment, desktop vs. mobile, private vs. public) · Mayer, Ohme, Maslowska, Segijn — Social Media + Society, Vol. 10, Issue 2, April 2024 · retrieved 2026-08-23
  3. Beyond Viewability: The Role of Attention in Creative Effectiveness (passive vs. active attention framework, 50,000+ ad normative database) · Kantar — Ecem Erdem, November 2023 · retrieved 2026-08-23
  4. Five Keys to Advertising Effectiveness, 2023 update (~450 CPG campaigns; creative = 49% of incremental sales, 46% for social media specifically) · NCSolutions, 2023 e-book · retrieved 2026-08-23
  5. Quantifying the Persona Effect in LLM Simulations (persona conditioning explains less than 10% of variance in subjective annotations even with prompting) · Hu & Collier — ACL 2024, Long Papers · retrieved 2026-08-23
  6. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (GPT-4 flips verdict 35% of the time on swapped A/B labels; 85% agreement with human raters vs. 81% human-human baseline) · Zheng et al., NeurIPS 2023 · retrieved 2026-08-23
  7. Out of One, Many: Using Language Models to Simulate Human Samples (GPT-3 matched real ANES vote choices at tetrachoric correlations of 0.90-0.94 across three election waves) · Argyle, Busby, Fulda, Gubler, Rytting, Wingate — Political Analysis 2023 · retrieved 2026-08-23
Stop guessing which creative wins.

Your ad creatives go in, up to six at once. A thousand synthetic consumers argue it out, before a dollar of media moves.

Free in early access · No card required