When you compare two offers, keep the customer, channel, budget rule, follow-up, measurement, and test period as similar as the platform allows. Change the offer deliberately and record every other change. Otherwise, a better result may belong to a different audience, page, week, salesperson, or tracking setup rather than the offer itself.
An offer is the complete deal a customer is asked to accept. It can include the promised outcome, scope, price, payment terms, timing, guarantee, eligibility, and next step. If version A and version B change several of those parts together, you are comparing two offer packages. You may learn which package performed better under the test conditions, but not which individual ingredient caused the difference.
Make both offers real before testing them
Do not spend money comparing an offer the business cannot deliver. Complete the one-page offer worksheet for each version, then check four conditions:
- The same type of customer could reasonably choose either offer.
- The business can honor the price, scope, timing, and capacity promised in both versions.
- The economics of a completed sale are known well enough to compare customers, not only clicks or leads.
- The landing page, form, phone path, lead delivery, and outcome tracking work for both versions.
Use the landing-page preflight to test the complete response path. A broken form in one version is not evidence that its offer is weaker.
Write an offer test card before launch
The card prevents the test from becoming a story assembled after the numbers arrive. Keep one copy where the person changing ads, pages, and follow-up can see it.
Business question: Customer and eligibility: Offer A: Offer B: What is intentionally different: What must remain fixed: Primary business outcome: Secondary diagnostic measures: Allocated budget and spend rule: Start and end rule: Capacity limit: Tracking and lead-owner check: Decision rule: Changes or incidents during the test: Final decision and remaining uncertainty:
Write the business question in plain English. For example: "Which of these two deliverable service packages produces more completed, contribution-positive jobs from the same type of customer?" That is more useful than asking which ad gets the cheapest click.
Freeze the conditions that could explain the result
Use this register before the test starts. Mark each row fixed, intentionally different, or not controllable.
| Condition | What to record | Why it matters |
|---|---|---|
| Customer | Location, need, eligibility, exclusions | A different audience may have different demand or ability to buy |
| Channel and placement | Platform, campaign type, placement, device rules | Traffic sources can behave differently |
| Budget | Allocation, daily limits, bid strategy, spend changes | Unequal delivery can change both volume and cost |
| Timing | Start, end, weekdays, promotions, disruptions | Seasonality and unusual days can distort a sequential comparison |
| Message | Ad promise and creative that introduce the offer | A different message can attract a different mix of people |
| Destination | Page, form, phone number, booking path | Page friction can hide or exaggerate offer response |
| Follow-up | Owner, response standard, sales process | One version should not receive better handling |
| Measurement | Event definitions and revenue source | A changed conversion event makes the two results incomparable |
| Capacity | Appointments, inventory, service area | A full calendar can suppress one version independently of demand |
Advertising platforms provide structured ways to hold some conditions steady. Google's custom experiment guidance says an experiment shares traffic and budget with the original campaign and warns that changing either campaign while it runs can make results harder to interpret. Google's separate experiment sync documentation explains what base-campaign changes are copied to a trial and what still requires care.
Microsoft Advertising's experiment documentation describes a duplicate campaign receiving a split of the base campaign's budget and traffic. It also recommends avoiding base-campaign setting changes during the experiment for a fair comparison. LinkedIn defines its A/B Testing as two ad sets with separate budgets that differ by one variable. TikTok warns in its split-test editing guidance that if a control variable must change, the same edit should be applied to both ad groups.
Those tools can improve the comparison inside one platform. They do not make a Google campaign and a Meta campaign a controlled offer test, and they do not repair inconsistent sales follow-up or missing customer outcomes.
Choose the outcome before you see the results
Use a primary measure tied as closely as practical to the business result. Keep earlier measures as diagnostics.
| Stage | Useful measure | What it cannot prove by itself |
|---|---|---|
| Attention | Impressions, views, clicks | That the offer attracted a qualified buyer |
| Response | Calls, forms, bookings | That the inquiry was eligible or serious |
| Qualification | Qualified opportunities | That the sales process converted them fairly |
| Customer | Completed first sales | That revenue was collected or profitable |
| Economics | Collected revenue and contribution after the defined marketing cost | That the same result will repeat at a larger budget |
If the business has enough downstream records, completed contribution-positive customers make a stronger primary outcome than clicks. If volume is limited, keep the test framed as a learning record rather than declaring a universal winner. Do not invent a sample threshold. Use the platform's experiment readout where applicable, keep raw counts beside percentages, and record the practical amount of evidence the business required before launch.
The lead-quality scorecard provides the stages needed to compare inquiries fairly. The collected-revenue guide keeps estimated pipeline from being mistaken for money received.
Work a hypothetical service-business example
A heating company compares two fall offers to the same eligible service area. The platform allocates $600 of media spend to each version during the same test period. Both versions use the same campaign type, follow-up owner, booking process, and completed-job record.
- Offer A: A $129 system tune-up with a defined inspection checklist.
- Offer B: A $39 diagnostic visit credited toward an approved repair.
The two versions are complete packages, not a one-variable price test. The business can compare the packages, but it should not claim that the price alone caused the outcome.
| Outcome | Offer A | Offer B |
|---|---|---|
| Leads | 30 | 42 |
| Qualified leads | 18 | 15 |
| Completed customers | 6 | 4 |
| Collected revenue | $2,700 | $2,000 |
| Direct delivery cost | $1,080 | $900 |
| Contribution before media | $1,620 | $1,100 |
| Media spend | $600 | $600 |
| Contribution after media | $1,020 | $500 |
Offer B produced more leads and a lower cost per lead: $600 divided by 42 is $14.29, compared with $20 for Offer A. But Offer A produced more qualified leads and completed customers. Its media cost per completed customer was $100, compared with $150 for Offer B. In this simplified record, Offer A also left $520 more contribution after media.
That does not prove Offer A will always win. It gives the company a better business reason to keep or retest A than the lead count alone would provide. The company should also review recorded loss reasons, cancellations, capacity constraints, and any incident that affected one version. The loss-reason guide shows how to separate customer evidence from staff guesses.
End with one of four honest decisions
| Decision | Use it when |
|---|---|
| Keep A or keep B | The planned business outcome favors one offer, delivery was valid, and no major imbalance explains the result |
| Retest | The result is useful but too limited, close, or context-specific for the next decision |
| Inconclusive | The records do not distinguish the offers well enough to choose |
| Invalid | Tracking, delivery, eligibility, follow-up, capacity, or an unplanned change made the comparison unfair |
An invalid test is not a failed business. Record what broke and repair it before buying more evidence. An inconclusive test is also a result when it prevents a confident decision from being manufactured.
After choosing, keep the winning version stable long enough to establish a new baseline. Change the next important question separately. If you need to compare the full acquisition economics of the two offers, run each version through the marketing profitability calculator. If the setup requires campaign experiments, measurement, and controlled rollout support, review Ocean Media's paid advertising service.
Your next step
Run your numbersSources checked
- Google Ads Help: Set up a custom experimentChecked 2026-09-30
- Google Ads Help: About experiment syncChecked 2026-09-30
- Microsoft Learn: Experiment data objectChecked 2026-09-30
- LinkedIn Marketing Solutions Help: A/B TestingChecked 2026-09-30
- TikTok Ads Manager: How to edit a split testChecked 2026-09-30
About this resource
Created with AI assistance for Ocean Media Marketing. Examples are illustrative unless explicitly identified otherwise. Platform claims are checked against the listed sources. We do not claim that a quality score proves accuracy or guarantees results.
Original contribution: An original offer test card, fixed-condition register, downstream outcome ladder, four-decision validity rule, and independently checked heating-service example.
Suggest a correction