Most outbound copy tests finish the same way. Someone writes two versions, sends forty of each, one comes back with three replies and the other with one, and the team adopts the winner by Friday. Four weeks later the winner performs the way the loser did. We have run that loop on our own pipeline and on the pipelines we operate for clients, and this page is what the arithmetic allows at the volumes a small B2B company can actually send.
The window below is 28 August to 25 September 2026. Three first-touch variants ran side by side across six pipelines: one written in the founder's own speech, one in our previous house style, one built from a published outreach playbook. Assignment happens by lead id, stays fixed for the whole thread, and gets stamped on every outgoing message by the database rather than by whoever wrote the text.
Four weeks of three-way testing, with the error bars
871 first messages went out under the test. Split by variant: 279, 307, 285. Replies: 24, 28, 19. Reply rates: 8.6%, 9.1%, 6.7%.
Read that quickly and the second variant beats the third by 2.4 points, a 36% relative lift, on nearly nine hundred messages. That is the number a weekly outbound report prints. That is also the number a person acts on, and by the following Monday two of the three variants are switched off.
The same data with intervals attached looks different. The 95% confidence interval around 8.6% on 279 messages runs from 5.3% to 11.9%. Around 9.1% on 307 messages it runs 5.9% to 12.3%. Around 6.7% on 285 it runs 3.8% to 9.6%. The three intervals sit on top of each other for most of their length. Testing the widest gap, the 9.1% against the 6.7%, gives p = 0.27. The gap between the two leaders gives p = 0.83, which is another way of saying the coin landed the way coins land.
Acceptance on the invitation note repeats the pattern: 23.5% on 136 notes, 26.8% on 142, 21.7% on 152, intervals running roughly 15% to 34% in every case.
Four weeks, three voices, 871 messages, and the honest summary fits in one sentence. All three perform inside the same band, and the test keeps running.
The volume each question needs
The base reply rate across all three variants in this window is 8.2%. Our cold LinkedIn benchmarks for 2026 put that figure in context. With that base, standard sample size arithmetic at 80% power and a 5% two-sided threshold gives the following per variant:
- to catch a doubling, 8.2% to 16.3%: 252 first messages per variant
- to catch a half, 8.2% to 12.2%: 864 first messages per variant
- to catch a quarter, 8.2% to 10.2%: 3 147 first messages per variant
- to resolve the difference we are currently staring at, 6.7% against 8.6%: 2 954 per variant
Our own pace under this test is about 97 first messages per variant per week, roughly 290 first touches in total. So a doubling shows up in under three weeks. A half takes nine weeks. A quarter takes about eight months, and eight months from now the market, the segment and the product pitch have all moved, which makes the answer historical by the time it arrives.
The practical consequence shapes what goes into the test at all. Pick changes big enough that your volume can see them: the voice the message is written in, whether the first line names what you sell, the length, whether the message carries a finished piece of work. Word-level polish belongs to senders doing tens of thousands a month.
The assignment key is the experiment
Our first version of the key took the lead id modulo 100 and cut it into three bands of 33. It reads as fair. It was the single largest defect in the first two days of the test.
A batch in our system is a block of consecutive lead ids, median size 3 people, largest 78. A band of 33 consecutive ids is wider than nearly every batch we build, so the batch inherits one voice from end to end. Of the batches with 6 or more people built under that key, 9 of 17 went out in a single voice, and exactly 1 of 17 carried all three.
That matters because a batch is one country, one segment, one signal and one morning of sourcing. When the variant and the batch line up, the experiment measures the country and reports it as the copy.
We changed the key to lead id modulo 3 on 30 August. Since then, of the 96 batches with 6 or more people, 83 carry all three voices. The rule that generalises past our case: the period of the assignment key has to be shorter than the smallest unit you group by. Check it by running one query, counting distinct variants per batch, before you trust a single result.
The differences already in your data
Same window, same three variants, same machinery, reply rate by pipeline: 8.2% on 402 messages, 12.6% on 191, 5.4% on 74, 3.8% on 106, 3.7% on 82. The spread between the busiest audience and the quietest one is 3.4x, the same kind of spread our benchmarks page shows across audiences.
The copy difference under test is 1.9 points. The audience difference sitting in the same table is 8.9 points. Any imbalance in how the variants land across pipelines swallows the thing you are trying to measure. So assignment belongs at the level of the individual person and stays with them for the life of the thread. Choosing the variant per campaign or per day puts the audience and the copy on the same axis, and then the report describes the audience.
The same logic applies to which messages carry a label at all. For the first week of our test only the first touch was labelled, and 115 messages went out with no variant recorded, written by the follow-up loop, the reply loop and by hand in the console. We moved the stamp into the single database door that every writer passes through: variant from the thread when the thread already has one, variant from the key when the person is new. The stamp now covers 2 107 sent messages across 1 827 people, and of the labels applied in that window, 291 came from the thread and 167 from the key. Any message that a reader receives from you belongs in the count, and the only reliable way to guarantee that is to let the storage layer apply the label.
The tests a small sample does settle
Two examples from the same database, both decided in days.
Invitation notes by account type. Over 21 days, accounts on the standard plan that attached a note to a connection request hit 66% platform refusals. The same accounts sending a plain request sat at 0.3%. Premium accounts sent 309 requests with a note attached and collected zero refusals. An effect of that size is visible in the first week, and the interval around it stays clear of the alternative by a wide margin.
The farewell message. Our older sequence closed silent threads with a polite goodbye. It went out 133 times and produced 2 replies. That one is decided by absence: 133 attempts returning 1.5% retires a step without needing a control group at all, because the step has to beat zero and it barely does.
Both share a shape. The effect was large, the measurement was cheap, and the decision followed inside a week. Smaller volume raises the size of the effect you should be hunting, and the discipline is to spend your sending on questions that size can answer.
What we run while the numbers accumulate
Set the threshold before the data arrives. Ours is 30 first messages and 3 replies per variant before anyone is allowed to discuss the result. We passed both on day 6 of this test and still hold no verdict, because the threshold licenses a conversation and the confidence interval decides the outcome. The distinction saves a lot of Mondays.
Watch for changes in the leading variant. Since day 6, the leader here has changed zero times. Stable ranking is mild evidence in favour of the leader, and it stays mild while the intervals overlap.
Treat a fourth variant as a separate clock. We added one on 18 September, a first touch that carries a finished sample of the work rather than a description of it, running on a quarter of a single pipeline. It has 32 first messages and 3 replies so far, which technically clears our own threshold in its first week. That is precisely the trap this page is about, so it has its own, later reading date and no verdict until then.
Three things worth doing tomorrow with data you already have. Put a confidence interval around last month's winning variant and see whether the loser sits inside it. Run one query that counts distinct variants inside each batch, and read what your assignment key is actually doing. Then pick your next test by the size of effect you could plausibly produce, given the number of messages you send in a month.
If you would rather point your own model at this kind of arithmetic against your own pipeline, that is what our MCP access does: the sourcing and campaign tools behind one token, and the counting stays yours. The way we run the whole cycle for a company, from signals through to replies, is laid out on the service levels page.
One closing number for scale. Our current test needs 864 messages per variant to resolve a change of one half, and it produces 97 per variant per week. Everything above follows from those two figures. Work out yours before you write the second version of anything.
Questions buyers ask
How many messages do I need for a valid A/B test of outbound copy?
At an 8.2% base reply rate, 80% power and a 5% threshold, catching a doubling takes 252 first messages per variant. Catching a lift of one half takes 864 per variant, and a quarter takes 3,147. At about 97 first messages per variant per week, a half takes nine weeks.
Is a 9% versus 7% reply rate difference real?
Often it is noise. In our test, 9.1% on 307 messages against 6.7% on 285 gave p = 0.27, and the confidence intervals overlapped for most of their length. We set a minimum of 30 first messages and 3 replies per variant before discussing a result, and let the interval decide.
How should leads be split between copy variants?
Assign by person, keep the variant for the whole thread, and make the key's period shorter than your smallest batch. When we split by bands of 33 lead ids, 9 of 17 larger batches went out in a single voice. After switching to lead id modulo 3, 83 of 96 batches carried all three.
Which outbound tests can a small team settle quickly?
Tests with large effects. Invitation notes from standard accounts drew 66% platform refusals against 0.3% for plain requests, which was visible in the first week. A farewell message that went out 133 times and produced 2 replies was retired without a control group.