A/B testing WhatsApp broadcasts: what actually works
AndySendy academy
← All posts

🔗 «This is better» is a feeling. A/B test is a number

Marketers argue for years - question or direct offer, name or no name, short or long. Settled by split test on your own list, not intuition or someone else's blog case. How to split fairly, sample size to avoid mistaking noise for pattern, and what to test first.


A/B in WhatsApp isn't a feature - it's a process

No built-in «run A/B test» in WhatsApp Cloud API. Marketers register separate templates and manually split sends via CRM or mailing platform - see WABA breakdown. Discipline: split, send variants, record, decide.


Before / After: sample size

Before: Minimum 200 contacts per variant for reliable result.

After: 200–500 per variant - common industry floor for somewhat stable results. No official Meta standard. Some use 500+ for more confidence; 50-per-variant groups are mostly noise. 200 is speed vs stability compromise - 5–10 point gaps on small groups may be random.


What to compare - why read isn't king

Metric Shows Limit
Delivered Reached device Not interest
Read (blue ticks) Opened Up to ~25% disable read receipts - systematically low
Response Rate User replied Best success marker
Complaints/blocks Negative reaction Can kill variant despite good conversion

Winner = delivered + inbound text reply, not read alone - track segments via database segmentation.


What to test first

Priority order - elements visible before opening chat - within first cold message structure:

  1. First line - lock-screen preview. Open or not.
  2. CTA/structure - question vs statement, explicit reply ask.
  3. Length - short trigger vs long text.
  4. Name personalization - wrong name worse than none.

Rule: one parameter at a time. Change length + name + CTA together = guessing, not testing.


Mini-case: question vs statement in B2B

Distributor A/B tested cold supplier base, 300 numbers per variant. A (statement + attachment): ban on test numbers at message 40 - stop words, reports. B (question): 32% response, 0 blocks - see mass mailing ban mechanics.

Shows: at 300, difference was survival, not just rate. Doesn't prove: questions always hit 32%.

Second case: online school 2000 leads, 1000 per variant. Long course pitch vs short «[Name], closed session starts 7pm - send link here?». Short won +180% webinar click - single case, direction only.


Pitfalls breaking test purity

Link tests in first cold message are pointless. Links unclickable until save or text reply - comparing site vs channel links in touch one is methodological error.

Many simultaneous variants on gray numbers risky. Rapid unique text bursts from one number = script pattern - behavioral antispam - regardless of test outcome.

Small-sample winner doesn't auto-scale. 200–300 tests linguistic viability; 50k blast adds throughput limits, send speed, cumulative complaints over hours.


Common misconceptions

«10 variants at once like email» - False. Messenger punishes mass content experiments from one number.

«Won at 200 → blast 50k immediately» - False. Small test ≠ scale load profile.

«Blue ticks are exact» - False. Use reply conversion or target action.

«Change whole text and see» - False. Can't isolate what moved the needle.


Practical takeaways


🎯 Next step

Take current first-message template, one hypothesis - e.g. question vs statement. Split next batch 50/50, compare reply share not open rate.

Conclusion

Practical rule:

Test what recipients answer to, not what you like - the only metric that doesn't lie.