WhatsApp A/B testing: how to run broadcasts that don't lie
AndySendy academy
← All posts

🧪 WhatsApp A/B tests that don't lie

Split your file top vs bottom - the test is broken before it starts. CRMs usually sort contacts by date added, so the "top" is often fresher and warmer than the "bottom." When variant A beats B tenfold, that's not copy winning - that's audience age.

Most "A/B tests" in WhatsApp broadcasts aren't tests: two texts on two different audiences, sent at different times, with multiple variables changed at once. Results look convincing but don't answer "which text works" - confounders snuck in.

In 2026 this costs more: WABA rates each template separately; complaints on a bad variant can push the account to Low Quality; hypotheses without randomization burn numbers - see WABA and templates.

After this article: how to split the list, which variable to change, how many contacts per variant, and which metric wins - without fooling yourself.


What's an A/B test vs chaos

A/B test = two message variants on two comparable audience segments, exactly one element different. Change offer, length, CTA, and send time together - that's guessing, not testing.

No built-in split-test in WhatsApp Business Suite - no random split, no significance calculator. It all rides on the marketer and external CRM.

Third pattern: blast 5–10 different texts from one number "to find a winner faster." Doesn't speed search - many different wordings from one account reads as spam script to Meta, not marketer - can trigger preemptive session ban - see ban mechanics.

A test is a controlled experiment, not shots in the dark.


Mistake #1: splitting the list in half

Grossest error - first half of rows gets A, second half B. Sounds like randomization; in practice it's chronological split.

Mini case. E-commerce marketer split 2000 clients: first 1000 rows (this month) got A, second 1000 (two years of base) got B. A got 35% response, B got 2% and cascading number ban. "Text A is perfect" was wrong - test compared warm vs cold, not texts - see warm vs cold base.

Fix - shuffle before splitting (random function in spreadsheet or CRM sample), don't cut by position - see list segmentation.


One-variable rule

One element changes per test: first line only, question at end only, length only, name personalization via {{1}} only. Rest of body identical.

Working pairs:

Change offer and tone together - you get a result, not a cause.


How many contacts per variant: 200 or 500?

Original thesis: "Minimum 200 per variant = significant."

Updated: 200 per variant is operator practice guide, not statistical standard. Split-test playbooks often cite 500+ as more reliable. No universal number - depends on baseline conversion and effect size you need to detect.

Source Sample Status
Broadcast operators 200–500 per variant practice guide
Split-test playbooks 500+ per variant stricter standard
Groups of 30–50 - random noise, test useless

Only 200 per variant? Run anyway - as hypothesis, not final verdict. Re-check on larger base next iteration.


Which metric wins

"Clicks vs replies" needs nuance.

First touch: clicks aren't really a metric - WhatsApp makes links from unknown numbers non-clickable until contact saved or reply sent. Not that clicks matter less - they don't exist as workable first-touch metric.

Response Rate and dialogs started - primary first-touch metric - see metrics dashboard. After dialog starts - catalog, deck, site link - clicks and conversions matter again.

Read Rate - least reliable dashboard number - see read no reply. 15–25% disable read receipts; many read via push shade without opening chat - stays delivered in CRM.

Metric Shows Test reliability
Response Rate Replied to first message High - primary first-touch KPI
Read Rate "Read" status Low - distorted for 15–25%
Link CTR Link click Meaningless first touch; useful step 2+
Reports "Report" taps Critical for WABA template quality

Complaints and template quality: invisible part of test

WABA: each variant = separate template; Meta scores complaints per template. Above 0.1–0.2% reports vs deliveries → Low Quality, template blocked mid-test.

Higher Response Rate + rising complaints ≠ winner - account risk. Watch reply and complaint trend together.


How to launch correctly

Send A and B same day, same hour - else you measure weekday/time activity, not text. A Monday morning vs B Friday evening kills the test before analysis.

Keep send speed and delays identical. One account, equal groups, same delays but wildly different semantics can flag as formulation probing, not manual outreach. Closer tone, structure, send length = fewer security flags.


Mini case: question vs pitch

Commercial real estate company: 1000 cold retail contacts → two randomized groups of 500.

A - pitch with proposal link. B - short situational question: "expanding in [city] or locations frozen?" A: 1.2% Response Rate, two numbers lost to reports. B: 29%, no bans; full pitch went step two only to repliers.

Doesn't mean questions always win 10× - scale specific to this cold B2B case. Direction matches first-touch logic: dialog invite beats instant pitch on average.

Forum case - no public methodology verification.


Winning a test ≠ scale immediately

Win on 200–500 sample = primary audience reaction only. Roll to tens of thousands adds send limits, Meta filters, cumulative complaints over time.

Small-test winner = scaling hypothesis, not guarantee. Roll out in stages, rising volume, watch complaints each step - see text vs account diagnosis.


Pre-launch checklist

  1. Random split, not list position (chronology, alphabet, ID).
  2. Exactly one element differs between variants.
  3. 200–500+ contacts per variant - by how critical result is and whether you can re-check later.
  4. Both variants start same day and hour.
  5. Same send speed and delays for both.
  6. Default first-touch metric - Response Rate and dialogs started, not read rate or clicks.
  7. Track complaints per variant; alert ~0.1–0.2% of deliveries.
  8. Don't stop early - premature stop risks random, not real winner.

🎯 Next step

Before next broadcast, check current split: how was the list divided, how many variables differ? If either answer is "not sure" - recalculate past results, don't scale.

Conclusion

Practical rule:

A test without randomization isn't A/B - it's comparing two audiences that accidentally looks like comparing two texts.