Cold email A/B testing: a free prompt to test one variable at a time

WORKSHOP

Workshop 28: Cold email A/B testing: a free prompt to test one variable at a time

If you change the list, the copy, and the offer at once and the campaign works, you cannot say what worked. The next campaign starts from a guess, and a lucky result gets adopted as the new normal.

Cold email A/B testing has a second trap: replies trickle in over weeks. A test read after the first email has sent is biased toward that email, and a test with a handful of sends per arm cannot tell signal from noise.

What it does

The prompt works through the test with you one step at a time. It asks, waits for your answer, then moves on. Its job is to hold you to one variable per experiment.

It starts by checking you have a baseline: at least one campaign that ran for three weeks, so you know your normal. Then it names the kind of test, writes the hypothesis, sizes it, and sets the success line. At the end you get a plain-text experiment plan with a date to read the result. Once results are in, it adds the outcome per arm, a confidence rating, and the next test.

What it catches

  • Testing without a control. No campaign has run for three weeks yet, so there is no normal to compare against. Run one first.

  • A hypothesis you cannot state. If it does not fit in one sentence with a reason, the experiment is not understood yet.

  • A "constant" that is moving. Every fixed element is written down. If one is changing, the test is fixed or called combined.

  • Deciding success after the data. The target and the failure line are written before launch, so a different "learning" cannot be argued afterwards.

  • Staggered launches. Both arms go out at the same time, on the same mailbox split and schedule, or day of week and warmup state muddy the result.

  • Calling it early. Results are read after the last step plus a grace period for late replies, on positive reply rate, with the send count next to every rate.

  • Testing copy on a broken list. The usual order of impact is list, then offer, subject line, opener, call to action, and timing. A bad list beats any copy.

  • Adopting a combined win as the baseline. It gets split into single-variable follow-ups to find what drove it.

How to use it

Paste it into any chat. Copy the prompt below into ChatGPT, Claude, or any other assistant and describe the campaign you want to improve.

Add it to a Claude Project. Upload the prompt as a project file. If you connect Apollo, it pulls the numbers for each arm itself and confirms with you before it creates or changes anything live. Setup: the Claude setup guide.

The prompt

Copy the whole prompt below, from the first line to the last.

You are helping me design an experiment on a cold outbound campaign so I actually learn something from it. If I change the list, the copy, and the offer at once and it works, I cannot say what worked. Your job is to hold me to one variable per experiment. Work through the steps below with me one at a time. Ask, wait for my answer, then move on.

If my sending platform is connected as a tool, pull the numbers yourself and confirm with me before you create, activate, or change anything live. If nothing is connected, tell me exactly what to set up and what to pull, and I will bring the results back.

You advise and I decide. If I want to run a test you think is muddy, say why, then help me run it.

First, check there is a baseline. An experiment needs a control. If I have not had at least one campaign run for three weeks, so I know my normal, tell me to run one first.

Name the kind of experiment: - List only. Change the targeting, hold copy and offer fixed. Tells me whether a segment fits better. - Copy only. Change the copy or a variant, hold list and offer fixed. Tells me whether a message lands. - Combined. A whole new campaign for a new audience. Nothing is isolated, so the result is a hypothesis, not a conclusion.

Then the framework: 1. Write the hypothesis in one sentence, with the reason. "Targeting heads of operations instead of sales leaders will get a higher positive reply rate, because they feel this pain daily." If it does not fit in a sentence, I do not understand the experiment yet. 2. Name the single variable. Write down exactly what changes and everything that stays the same. If a "constant" is actually moving, fix it or call the test combined. 3. Size it honestly. Small samples lie. An arm with a handful of sends cannot separate signal from noise. Big effects need fewer sends, small effects need many. When unsure, run more before concluding. 4. Decide success before launch. Write the target and the failure line down first, so a different "learning" cannot be rationalised after the data comes in. 5. Launch both arms at the same time, on the same mailbox split and the same schedule. A staggered test is confounded by day of week and warmup state. 6. Measure after the full sequence has run, through the last step plus a grace period for late replies. Measuring early biases the result toward the first email. Judge on positive reply rate, not reply rate, and keep the send count next to every rate. For copy tests, group results by template. 7. Weight the result by confidence. A clean single-variable test with enough volume is high confidence. A combined test is low. Say which.

If I do not know what to test first, the usual order of impact is: list, then offer, then subject line, then opener, then call to action, then timing. A bad list beats any copy, so fix targeting before testing a subject line.

No borrowed benchmarks. Judge every result against my own past campaigns, never an industry figure. My baseline is the only honest yardstick.

Mistakes to stop me making:

- Changing three things and claiming a win. Nothing was learned.

- Calling it early. Cold replies trickle in over weeks.

- Testing copy on a broken list. Fix the list first.

- Adopting a combined win as the new baseline. Split it into single-variable follow-ups to find what drove it.


What to hand me at the end: a plain-text experiment plan with the kind, the one-sentence hypothesis, the variable and the full list of constants, the sample size per arm with your confidence in it, the success and failure lines, the launch plan, and the date to read the result. Once results are in, add the outcome per arm with send counts, the confidence rating, and the next test.

Method adapted from the experiment design skill in Apollo Operator, a free, open-source headless GTM toolkit by Creatop: github.com/creatop-gtm/apollo-operator

Part of the Apollo Operator prompt pack

This is one of 19 free prompts from Apollo Operator, the open-source headless GTM toolkit Creatop builds and runs on its own campaigns. Get the full pack here: the Apollo Operator prompt pack.

LATEST

FEATURED

Learn from our work

Logo

OUR NEWSLETTER

Notes from live campaigns. No theory, no filler.

PAGES

MORE

LEGAL

© 2026 Creatop. All rights reserved.

Learn from our work

Logo

OUR NEWSLETTER

Notes from live campaigns. No theory, no filler.

PAGES

MORE

LEGAL

© 2026 Creatop. All rights reserved.

Learn from our work

Logo

OUR NEWSLETTER

Notes from live campaigns. No theory, no filler.

PAGES

MORE

LEGAL

© 2026 Creatop. All rights reserved.