Buying signals beat better lists

Everyone tells you to tighten the list. We tested eleven list attributes across our own campaigns and none of them predicted a reply. Two things did, and neither is an attribute of the list.

11 / 0
list attributes tested, none predictive of replies
×3.1
French market vs English-speaking, on the same machine
×2.8
leads emailed within 7 days vs after 30
1
closed-won from outbound, on a campaign built around a dated deadline

What everyone does when a campaign underperforms

Rewrite the subject line. Tighten the ICP. Add a personalisation variable. Buy a better database. We did all of it, on our own campaigns and our clients', and we kept getting reply rates between 1 and 2%.

So we went back through every campaign that had contacted at least 100 people and asked a duller question: across everything we changed, what actually correlates with a human reply? The answer surprised us enough to change how we build campaigns.

Eleven list attributes, none predictive

Each attribute below was compared as an observed-to-expected ratio, stratified by campaign and by sequence step, so a single good campaign cannot carry the result. None of them separated from noise.

What we testedResult
Seniority and job function of the prospect×1.00, no effect
Role addresses (contact@, info@)0.20% of the base, nothing to filter
Size of the listno effect
Automated flow vs hand-built batchno effect
Personalised copy vs volume copyno effect
Prospect domain extensionno effect
Depth of the ICP definitionno effect
Which sending fleet was used×0.82
Domains listed on SURBL×0.78

The last two matter for a different reason: they hold the list constant and vary the pipe. Neither moved the result, which means our sending infrastructure was not the thing separating a good campaign from a bad one either.

Two things did pass

1

The market, not the list

French-speaking prospects replied at 1.52%, English-speaking ones at 0.49%. Same product, same sequences, same senders: a factor of 3.1, on the one variable nobody calls targeting.

2

Time spent waiting in the queue

A lead emailed within 7 days of being sourced replied 2.8 times more often than one that waited more than 30. At the time, only 30% of our leads went out inside that window and 14.5% waited over a month.

Both say the same thing in different words: what decides the reply is the context around the send, not the attributes of the person receiving it. Freshness is a signal. Market fit is a signal. Job title is not.

The uncomfortable part: most list comparisons are noise

Across 27 English-speaking campaigns, the spread between the best and the worst list was statistically indistinguishable from chance (dispersion 1.08 against an expected 1.00, permutation p = 0.34). The apparent range, from 1.69% down to 0.00%, is what random variation looks like at these volumes.

To tell a 0.5% list from a 1.5% list with any confidence, you need roughly 2,800 contacts on the same list. Almost nobody runs a single list that long before declaring it good or bad. We did the maths after months of confidently ranking lists we could not actually rank.

What actually moved deals: a dated trigger

Reply rate only tells you whether the copy got past the first second. So we looked at the CRM instead, at the deals that advanced past the first meeting.

Two campaigns stood out, and they share nothing in firmographics. One targets suppliers around NIS2, the European cybersecurity directive: three mature deals, all three advanced. The other targets CMMC 2.0 in US defense supply chains: our only closed-won from outbound. Different countries, different company sizes, different industries.

What they share is a public deadline and an audience that sells compliance rather than suffering it. Our campaigns defined by vertical instead, design agencies, founders, AI agencies, plateau at the first meeting: they book it, they rarely go further.

Honest limit: the NIS2 campaign has three mature deals. That is the minimum we allow ourselves to look at, and it is a signal to retest, not proof to wave around.

What makes a trigger usable

We now keep a trigger only when it ticks all four:

1

A public date, three to nine months out

Close enough to create urgency, far enough that budget still exists.

2

An audience that sells because of it

Whoever monetises the deadline moves faster than whoever merely complies with it.

3

Buyers you can actually list

A trigger you cannot turn into a list of named accounts is a blog post, not a campaign.

4

A window that closes

When the deadline moves, the urgency dies with it. We dropped a campaign built on the EU AI Act the day its high-risk obligations slipped to December 2027.

What we changed

1

We stopped qualifying on list attributes

No more gates on seniority, company size or ICP depth. They predicted nothing, and every gate costs volume.

2

We cap the queue at seven days

A campaign is never loaded with more stock than it can send in a week. Freshness is the second strongest effect we found, and it is free.

3

We steer on pipeline stage, not reply rate

Reply rate judges the copy. Whether a campaign deserves to live is decided by what happens after the meeting.

How to apply this on your own outbound

Start by asking what changed at the account this quarter: a funding round, a regulation with a date, a hire, a migration. Build the list from that event, not from a persona filter. Then check the two boring things: that your leads leave the queue within a week, and that your emails reach the inbox at all, which is a separate problem entirely, as our Microsoft 365 study showed.

If you want to pressure-test the rest of the machine: the cold email auditor scores the message itself, the follow-up simulator sizes the sequence, and our intent signal tools page reviews the software that detects these events for you.

Method and limits: this is our own outbound, measured on campaigns with at least 100 prospects contacted, each prospect counted once, replies counted as distinct people. The comparisons are observational, not randomised: we did not assign markets or delays at random, so they carry the biases of how we run the business. The trigger-level findings rest on a handful of deals. We publish them because measured numbers with small samples still beat confident advice with none.

Want outbound built on signals, not lists?

We build the system, run it, and measure what it actually produces.

Book a GTM audit