Our email team ran 120 A/B subject-line tests last year across three client newsletters, tracking open rate, click-to-open, and unsubscribe rate for each one. I'm sharing the aggregate because the pattern across 120 tests is worth more than any single clever result, and because a few of these findings directly contradict advice I used to give clients with a completely straight face and a slide deck to match.
We'd been calling a popularity contest a test
The embarrassing confession first: before this year, our 'A/B tests' usually compared two totally different subject lines written by two different people who disagreed about everything from tone to comma placement. When one won, we learned nothing transferable, because forty things were different between the two, not one. We'd basically been running a popularity contest between coworkers, declaring a winner, and calling it data. We did this for years. We charged clients for it. I'd like that time back.
Designing tests that actually isolate a variable
We switched to the Subject Line A/B Set Generator — 10 Isolated-Variable Tests prompt to design pairs where one thing changes and everything else holds still — short vs long with the same words, question vs statement with the same claim, number vs no number with the same offer. GPT-4o is the right model here precisely because it's fast and cheap. We're generating dozens of these a week across three clients, and we do not need a flagship reasoner to write 'short vs long.' What we need is consistency, clean formatting, and volume, and GPT-4o delivers a fully structured set in seconds without choking the budget. Crucially, each pair ships with a one-sentence hypothesis, which forces us to commit to a belief before the data arrives to flatter or humiliate us.
The discipline of isolation changed what 'a result' even means to us. When the long subject beats the short one now, we know it's the length, because length was the only thing we changed. That sounds obvious written down. It is shockingly rare in practice, because real teams under deadline test whatever two lines they happened to write, and then build a strategy on noise.
- Emoji in the subject: won 7 of 12 tests for consumer brands, lost 9 of 12 for B2B — the audience interaction dwarfs the variable itself, so a blanket 'emojis work' rule is just wrong
- Negative framing ('What you're getting wrong about X') beat positive framing by 31% in informational sends, but lost in promotional ones — frame depends on whether you're teaching or selling
- First-name personalisation: no statistically significant difference in 8 tests — audiences have fully learned to ignore 'Hey [Name]' and so should you
- Specificity won almost everything: '3 ways to cut your SaaS spend this month' beat 'How to reduce software costs' in all 15 specificity tests we ran
Why the hypothesis is the most valuable line
The single best habit this prompt installed in our team is the hypothesis. Writing down 'B wins because curiosity beats benefit for cold readers' before you send means that when B loses, you feel something — surprise — and surprise is the only place the learning actually lives. Without a prediction, a result is just a number you'll have forgotten by Friday lunch. With a prediction on record, a wrong result rewires how you think for the next hundred sends. It turned out that half the value of these tests was never the open rate at all. It was the sentence I was forced to write before pressing send.
Occasionally I flat-out disagree with GPT-4o's hypothesis, and that's the best possible outcome, not a flaw. Now there are two predictions on the table — mine and the model's — and the test will settle the argument with money instead of opinion. That's a real disagreement with stakes, not a coin flip between two coworkers' egos and one loud manager. The model becomes a sparring partner that always shows up with a stated position, which is more than I can say for some humans I've tested subject lines with.
Grab the Subject Line A/B Set Generator — 10 Isolated-Variable Tests prompt on Prompt Dock, run it on your next campaign, and make one promise to yourself: write the hypothesis before you look at a single open rate. The discipline is the actual product here; the subject lines are just the delivery vehicle. When the issue underneath the winning subject still needs writing, I feed everything into my Newsletter From Bullet Points — A Finished Issue From a Scrappy List prompt and finish the whole send in one sitting.