Stop A/B Testing Garbage
Generative AI broke a cost curve that held for fifty years. Producing a marketing variant used to require a designer, a copywriter, a round of internal review, and three days. Now it requires a prompt.
Twenty headlines before lunch, ten static ads by the afternoon, then six hooks, four landing pages, and three offers. A creative team of two can generate more concepts in a week than a 2019 agency shipped in a quarter.
Traffic didn't get cheaper.

Every concept that reaches a real experiment consumes impressions, clicks, conversion events, analyst time, and calendar. Making the thing got almost free. Finding out whether it works costs what it always did.
The gap between those two costs is where a business lives. A new bottleneck formed the moment creative supply outran the ability to evaluate it, and nobody owns the layer that decides what gets tested.
Somebody should own it. Here's the shape of that company:
The money: Twenty agencies at $1,500 a month is $30,000 MRR. Fifty at $2,000 is $100,000. The first version is a $2,500 sprint with no software.
Inside:
• The $2,500 sprint that proves the ranking
• MVP split: LLM features, stats layer
• Agency pricing from $750 to $3,000/mo
• The ledger moat competitors can't copy
The Math That Makes This a Business
Run the numbers on a single Meta creative test and the problem stops being abstract.
Performance marketers budget these tests off a rule of thumb. At a $40 target CPA, a concept needs roughly 50 purchase events before anyone can read it with confidence, which works out to about $2,000 of live spend per variant. Lighter reads still get $500 to $1,000 per creative for anything a team would call rigorous, plus 7 to 10 days of runtime before declaring a winner. No benchmark study blesses those figures. They're what practitioners actually put behind a test.

The hit rate is worse than the price tag suggests. Conversion Team audited 2,288 A/B tests that ran to a clean result across 71 client engagements. The raw win rate looked healthy at 50.5%, but only 19.1% of individual tests produced a statistically significant winner. That figure sits comfortably inside the published industry band, which runs from roughly 12% at Optimizely to around 20% at CXL and Convert, so the picture below is conservative rather than cherry-picked. Four out of five tests spend real money to teach the team nothing they can defend.
An agency pushing 20 concepts through live tests for one brand in a month spends somewhere between $10,000 and $40,000 in traffic to surface maybe four defensible winners. The other sixteen were tuition.
Cut that queue to the six concepts most likely to matter, keep three of the four winners, and the same brand learns nearly as much for a quarter of the spend. Somebody should be selling that difference.
What Amazon Actually Proved
On August 3, 2026, Stefan Hut and Lorenzo Masoero published Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation through Amazon Science. They ran an AI-agent simulation framework against 67 historical marketing A/B tests and reported the results honestly, which is rarer than it should be.
The naive version failed. An off-the-shelf foundation model hit 0.70 sign overlap, meaning it usually knew which direction a treatment would push the metric. But it systematically overshot effect sizes, and its launch alignment landed at 0.41 against a random floor of 0.33. A coin flip with a little taste. The authors were direct: the uncalibrated simulator has no business making launch decisions.

The laziest version of this startup dies right there. Nobody needs another dashboard announcing "our AI predicts a 17.4% conversion lift."
The second result is where the money is. When the researchers calibrated simulated behavior against observed pre-period behavior using Platt scaling, squared prediction error across those 16 experiments compressed by roughly 77 times. Exposing each synthetic agent to both alternatives instead of polling independent synthetic crowds cut standard errors by about 2.4 times.
The paper is equally clear about what it can't do. Personas miss plenty of what drives a purchase, the benchmark covers a single domain, and the evaluation was retrospective, with no ground truth available at prediction time. Magnitude overshoot survives calibration. The authors describe their goal as a faster, better-informed experimentation pipeline, and they say plainly that it doesn't replace one.
The product hiding in that paper is a filter. Screen for direction, kill the likely losers before they touch live traffic, and leave the oracle business to somebody else.
Everyone Is Building the Wrong Half
The market isn't waiting for someone to invent AI creative evaluation. It already exists in two mature forms, and both leave the interesting problem untouched.
On one side, the experimentation platforms prove merchants will pay to run tests. Shopify merchants moved $378.4 billion in GMV during 2025, up 29% year over year, and Q2 2026 alone hit $115.6 billion, up 32%. Five straight quarters of GMV growth above 30% have built a software budget around testing. Intelligems sells Shopify testing from $69 a month up to $349, and now markets an AI layer that helps merchants "know which content tests to run next." Shoplift runs $99 to $999 a month and ships Lift Assist, which auto-creates branded split tests.

The pre-flight creative graders prove the other half: marketers will pay to score work before launch. Kantar's Link AI is built on a database of 300,000-plus ad tests and 35 million human interactions, predicts a TV ad's in-market performance in about 15 minutes, and counts Coca-Cola, Google, and Unilever as clients. Neurons scores attention with neuroscience-derived models. Dragonfly AI does the same for visual saliency across Nestlé, PepsiCo, and L'Oréal. Test My Advertising has taken the low end, returning a verdict in 20 seconds starting at $99. Newer persona-based entrants keep arriving, with POPJAM and Deja Vu running creative variants past synthetic cohorts and forecasting engagement before spend, alongside a dozen more credible vendors selling some flavor of the same thing.
Neither side holds the useful record. The platforms only know what happened inside their own tool. The graders generalize across everyone's ads, and none of them keeps a per-client history of what was predicted against what actually happened in market. Nobody keeps the record of what this specific agency predicted, what it launched, what the live result was, and what that says about the batch sitting in the queue right now.
A crowded synthetic-scoring category is a signal, and the signal says build the other half. The white space is narrower than "AI tells you if your ad is good," and better for it.
The Heist
A performance agency with eight DTC clients runs on a Monday rhythm. The creative team dumps another batch into the pipeline: twelve Meta concepts for Brand A, eighteen hooks for Brand B, eight PDP variations for Brand C, six landing-page treatments for Brand D, plus twenty concepts recycled from portfolio winners.
Today the shortlist gets picked by some blend of strategist judgment, creative director taste, client politics, and which files happened to be finished on time.
Unlock the Vault.
Join founders who spot opportunities ahead of the crowd. Actionable insights. Zero fluff.
“Intelligent, bold, minus the pretense.”
“Like discovering the cheat codes of the startup world.”
“SH is off-Broadway for founders — weird, sharp, and ahead of the curve.”