Creative volume is not the same as creative learning

A team can launch fifty ads and learn almost nothing.

That happens when assets are produced without a hypothesis, naming system or feedback loop.

The media team sees ad-level performance.

The creative team sees files.

The strategy team remembers that “UGC worked.”

Nobody can explain which concept, angle, hook or proof created the useful difference.

This tracker is designed to connect those layers.

It contains:

  • Creative Registry;
  • Weekly Performance;
  • Fatigue Tracker;
  • Test Scorecard;
  • Dashboard.

The purpose is not to create another reporting spreadsheet.

It is to preserve the chain:

why we made it → what we made → how the auction treated it → what happened → what we should make next

Start with taxonomy before metrics

The Creative Registry requires each asset to have a Creative ID and structured fields such as:

  • concept;
  • angle;
  • hook;
  • format;
  • creator;
  • product;
  • launch date;
  • primary hypothesis;
  • status;
  • owner.

This helps solve a basic analytical problem.

If one ad changes only the opening three seconds and another changes the entire customer argument, they are not equivalent variations.

A taxonomy lets the team separate:

Conceptual variation — a different idea.

Execution variation — a different expression of the same idea.

That distinction becomes critical when production automation makes cosmetic variation cheap.

The organization needs to know whether it is increasing idea diversity or merely asset count.

The hypothesis field is the most important column

“Test new video” is not a hypothesis.

A useful hypothesis predicts a relationship.

Examples:

Demonstrating the product in the first five seconds will reduce skepticism and improve qualified CVR.
Customer proof will improve new-customer ROAS relative to offer-first creative.
A price-forward hook will increase CTR but may reduce downstream conversion quality.

A hypothesis gives the result somewhere to go.

If the test fails, the team can revise a belief.

If the asset simply “underperforms,” the team has no structured learning.

Weekly Performance separates allocation from efficiency

The performance sheet tracks:

  • spend;
  • impressions;
  • clicks;
  • conversions;
  • revenue;
  • CTR;
  • CVR;
  • CPA;
  • ROAS;
  • frequency;
  • spend share.

Spend share is intentionally included.

Motion’s 2026 Creative Benchmarks analyzed $1.29 billion in realized Meta spend across 578,750 creatives and 6,015 advertiser accounts. Motion uses spend as its cross-account performance signal because ROAS and CPA are not comparable across heterogeneous businesses. The company explicitly warns that spend concentration does not prove downstream business value. citeturn862450search0turn862450search1

That is a useful methodological lesson.

An asset that absorbs significant spend has demonstrated that the platform is willing to allocate budget to it.

The business still needs to determine whether that allocation creates acceptable economics.

Do not rank creative by ROAS alone

Imagine two ads.

Creative A has 2.7x ROAS at $80,000 spend.

Creative B has 4.1x ROAS at $3,000 spend.

Which is stronger?

The answer is not obvious.

Creative B may be excellent and under-distributed.

It may also be living in a small high-intent pocket that cannot absorb scale.

Creative A may be the more valuable portfolio asset because it continues to perform while receiving much more budget.

This is why the tracker keeps spend, spend share and efficiency together.

The correct creative decision often depends on capacity, not the highest ratio.

The Fatigue Tracker is a diagnostic, not a kill switch

Creative fatigue is difficult because several different problems look similar.

CTR falls.

CPA rises.

Frequency increases.

Spend shifts.

The temptation is to label the creative “fatigued” and replace it.

The workbook creates a heuristic Fatigue Score using:

  • CTR deterioration;
  • CPA deterioration;
  • frequency;
  • creative age.

That score is intentionally labeled as a diagnostic.

It is not a causal model.

A rising fatigue score means “investigate.”

The underlying problem could be:

  • audience saturation;
  • auction inflation;
  • landing-page decline;
  • offer fatigue;
  • product availability;
  • tracking changes;
  • competitor activity.

Do not pause a profitable creative because a spreadsheet score crossed an arbitrary line.

Why the score combines several signals

Frequency alone is weak.

A high-frequency ad can remain productive.

CTR decline alone is weak.

A product-price increase can lower CTR without creative fatigue.

CPA decline can be caused by conversion tracking.

The tracker combines several indicators to create a triage signal.

The weights are examples.

They should be adapted to the account.

A long-consideration financial product and a fast-moving fashion brand do not have the same fatigue behavior.

Motion’s 2026 research reinforces that context: creative testing volume and format performance vary substantially by vertical and spend tier, and the report explicitly rejects universal production targets. citeturn862450search2turn862450search10

The tracker should therefore be calibrated from your own history.

The Test Scorecard protects interpretation

The Test Scorecard contains:

  • Test ID;
  • hypothesis;
  • control;
  • treatment;
  • primary KPI;
  • minimum spend;
  • start and end;
  • result;
  • confidence;
  • decision;
  • learning.

This prevents a common creative-testing failure:

The ad launches first.

The team decides later what it was testing.

That creates retrospective storytelling.

If CTR improved, it was a hook test.

If CVR improved, it was a proof test.

If neither improved, “the audience was wrong.”

The scorecard forces the question to exist before the result.

Minimum spend is not statistical significance

The workbook contains a Minimum Spend field.

That is an operating guardrail, not a mathematical proof threshold.

A creative needs enough opportunity to produce a meaningful signal.

The correct threshold depends on:

  • CPM;
  • conversion rate;
  • target CPA;
  • account scale;
  • test design.

For rigorous A/B tests, use proper power and sample-size calculations.

For portfolio creative testing inside an automated campaign, the minimum-spend field is simply a rule preventing the team from declaring winners or losers after trivial exposure.

Winners are rare; do not overreact

Motion’s benchmark classifies “winners” as creatives spending at least ten times the account median and at least $500. Under that definition, roughly 5% of creatives qualify. Motion’s own methodology stresses that this is a reporting convention, not a universal quality law. citeturn862450search0turn862450search7

The practical implication is useful:

A creative system should expect many assets not to become major spend absorbers.

That is not permission to produce junk.

It means production should be designed as a portfolio with enough shots on goal to discover outliers.

Track learning velocity

The Dashboard summarizes portfolio state, but the strongest metric is not included as one formula because it requires judgment:

How much of the next creative batch is informed by the last batch?

A team can measure this manually:

  • percentage of briefs citing prior learning;
  • number of hypotheses resolved;
  • time from signal to iteration;
  • number of winning concept families extended;
  • number of repeated failed ideas retired.

That is learning velocity.

Creative velocity without learning velocity creates content debt.

Failure mode: comparing different jobs with one KPI

A creator testimonial designed to introduce a category and a retargeting offer ad may not have the same job.

If every asset is ranked by immediate CPA, the portfolio will drift toward lower-funnel creative.

Tag creative intent.

Compare assets against appropriate peers.

The tracker’s taxonomy is there to prevent one universal leaderboard from flattening strategic differences.

Failure mode: retiring the asset instead of the concept

One execution can fatigue while the underlying concept remains powerful.

If a product demo worked for six weeks and deteriorates, the team should ask:

  • Is the concept exhausted?
  • Or is this execution exhausted?

A new creator, proof structure, setting or hook can extend a concept.

This is why concept and variant need separate fields.

Failure mode: benchmark chasing

Do not turn Motion’s 5% winner rate or weekly testing volumes into production quotas.

The report itself warns against this.

A team with $20,000 monthly spend does not need the same creative throughput as an enterprise advertiser spending millions.

Use benchmarks to ask whether the organization’s production capacity is obviously out of line with its scale.

Then use account evidence to set the cadence.

Download the tracker

Use the workbook as a shared operating layer between media and creative.

The media team should be able to explain what the auction did.

The creative team should be able to explain what was different about the assets.

The next production brief should connect both.

The final question is not:

Which ad won?

It is:

What did we learn that changes what we produce next?

Download the Creative Testing Scorecard & Fatigue Tracker (XLSX)

More from Radar