Evidence status: Agency-published case study with repeated 50/50 tests; detailed statistical appendix not public
The strongest part of the RAC case is the skepticism
Many automation case studies begin with an assumption that the newer system is better.
RAC’s paid-search story begins differently.
According to Merkle, the UK roadside-assistance organization had previously tested auction-time bidding in 2021 and 2022 and found that its existing intraday bidding approach with manual CPC performed better.
That history matters.
RAC did not have a theoretical objection to automation. It had internal evidence supporting the incumbent method.
The challenge was therefore not “How do we implement value-based bidding?”
It was:
Has the technology improved enough that the old conclusion should be tested again?
That is a much stronger operating question.
The incumbent had earned the right to be the baseline
Teams often treat legacy campaign logic as technical debt by default.
Sometimes it is.
Sometimes it reflects a system that was tested and worked.
Merkle says RAC used intraday bidding alongside manual CPC to drive online sales revenue and had prior testing that supported the approach.
Replacing that system without a controlled comparison would have mixed technology adoption with strategy.
Instead, the team used the incumbent as the baseline.
This is a useful principle:
Automation should compete against the best current process, not against a straw-man manual account.
That makes the result more meaningful.
The treatment combined auction-time and value signals
The test did not evaluate “automation” in the abstract.
RAC combined:
- auction-time bidding;
- value-based bidding;
- tROAS;
- Floodlight conversion values;
- structured testing through Search Ads 360.
The primary success measures included average revenue per user and ROAS, with conversion rate and high-intent traffic also monitored.
Merkle says the team deployed multiple 50/50 tests across high-volume campaigns.
That phrase is important.
Repeated testing across multiple campaigns is stronger operational evidence than a single before-and-after comparison.
A before-and-after test can be contaminated by seasonality, promotions, competitor changes or changes in demand.
A simultaneous 50/50 design reduces some of those risks.
The reported outcomes
Merkle reports that campaigns using tROAS with auction-time bidding and Floodlight values outperformed the alternatives.
The agency lists the following results:
- 27% increase in revenue
- 30% increase in conversion rate
- 4% increase in ARPU
- 14% increase in ROAS
The case also says the advantage held during critical sales periods and ordinary business periods.
Based on those results, RAC adopted auction-time value-based bidding as the default for performance-focused paid-search campaigns.
These figures are significant.
The public case does not disclose the exact campaign count, test duration for each split, media spend, statistical significance or raw treatment/control tables.
They should therefore be treated as agency-reported outcomes from a structured test program, not as a fully reproducible academic experiment.
Revenue and ROAS moving together matters
A common automation failure is improving efficiency by reducing scale.
If a bidding system cuts spend aggressively, ROAS can increase while revenue falls.
That can be rational when the prior spend was unprofitable.
It is not automatically a performance improvement.
The RAC result is more interesting because Merkle reports higher revenue and higher ROAS at the same time.
If accurate, that suggests the treatment found a better allocation of spend rather than simply shrinking exposure.
The 30% conversion-rate increase and 4% ARPU increase also point toward a mix effect: the system may have found users or auctions with both higher conversion probability and higher expected value.
That is exactly what value-based bidding is designed to do.
Floodlight values were not a reporting detail
Value-based bidding requires a value signal.
If every conversion is treated as equal, the system can optimize conversion volume but cannot distinguish a high-value outcome from a low-value one.
Merkle says Floodlight conversion values were aligned with RAC’s business goals.
That is the hidden dependency in the case.
The bidding algorithm may be sophisticated.
Its objective is only as useful as the value model supplied to it.
This is why teams should not copy the RAC bidding setup before asking:
- What does one conversion mean economically?
- Are values static or dynamic?
- Do values reflect revenue, margin or customer quality?
- Are important downstream outcomes missing?
- How quickly do those values become available?
- Can they be trusted enough to guide auction-time decisions?
The most transferable part of the case may be the signal design, not the bidding label.
Why prior failure should increase test quality, not block retesting
Organizations often develop institutional memory around automation.
“We tested broad match and it was bad.”
“Smart Bidding did not work for us.”
“PMax failed last year.”
Those statements can remain true for the original test and become outdated.
Platforms change.
Conversion architecture changes.
Creative improves.
The business changes.
A failed test should create a retest threshold, not a permanent doctrine.
For RAC, Merkle says previous internal tests had supported intraday bidding. The company later chose to reevaluate auction-time bidding because the technology had advanced.
That is a mature response.
The alternative extremes are both weaker:
- adopting every new feature because the platform promotes it;
- refusing to retest because an earlier version failed.
Failure mode: switching the whole account before proving the treatment
Automation changes can have large portfolio effects.
If RAC had moved every performance campaign to the new method at once, the team would have lost its baseline.
Any subsequent improvement or deterioration would be difficult to attribute.
The 50/50 structure preserved a counterfactual.
For other teams, the practical lesson is to create a migration lane.
Select campaigns that are:
- large enough to generate evidence;
- representative enough to matter;
- operationally safe enough to test;
- measurable with a stable business KPI.
Then expand only if the treatment earns the rollout.
Failure mode: measuring only CPA
Suppose the new system had increased conversion volume by 30% while pushing users toward lower-value products.
CPA might look excellent.
The business could still lose value.
RAC’s use of ARPU, revenue and ROAS alongside conversion rate is therefore important.
Value-based systems should be evaluated with value-based outcomes.
The exact metric depends on the business:
- revenue per conversion;
- contribution margin;
- expected LTV;
- policy value;
- qualified opportunity value;
- profit per booking.
If the evaluation metric does not reflect the value signal the system is optimizing, the team cannot tell whether the automation improved the intended objective.
Failure mode: letting automation erase promotion strategy
RAC operates in a business where promotional periods matter.
That creates a classic concern: can an automated system react appropriately during unusual demand?
Merkle says the treatment continued to perform during critical sales events as well as ordinary periods.
That is useful evidence because many teams distrust automated bidding precisely when commercial conditions become abnormal.
The correct operating response is not to assume the algorithm will handle every promotion.
It is to include promotional periods in the test program.
If the business cares about Black Friday, renewal season, product launches or holiday demand, the bidding system should earn trust in those conditions before becoming the default.
What the case does not prove
RAC’s result does not prove:
- auction-time bidding always beats manual bidding;
- value-based bidding works without accurate conversion values;
- tROAS is the correct strategy for every campaign;
- 14% ROAS growth is a reasonable benchmark;
- prior failed automation tests should be ignored;
- Search Ads 360 is required to apply the general method.
The strongest conclusion is narrower:
RAC used repeated controlled comparisons to determine that a newer automated bidding approach had become more effective than its established process under the tested conditions.
That is a test-and-rollout discipline.
A transferable rollout framework
Teams can copy the process without copying the settings.
1. Define the incumbent. Document the current method and why it exists.
2. Improve the value signal. Make sure the treatment optimizes something economically meaningful.
3. Choose representative campaigns. Avoid tiny campaigns that cannot produce useful evidence.
4. Run simultaneous tests. Preserve a valid baseline.
5. Measure value, not only conversion count.
6. Include abnormal periods where relevant.
7. Roll out only after the treatment beats the incumbent repeatedly.
8. Keep monitoring after migration.
The result is not blind faith in automation.
It is an evidence-based reason to trust it more.
Evidence note
This article relies primarily on a case study published by Merkle, the agency involved in the implementation. The page describes multiple 50/50 tests and reports the headline outcomes, but it does not publish the full underlying dataset or statistical appendix. Radar therefore treats the 27% revenue, 30% conversion-rate, 4% ARPU and 14% ROAS figures as agency-reported results from a structured test program.



