Guide
Write the creative test before you read the result
A useful test names the decision, the meaningful difference, and the evidence that would change your mind.
The useful part
Keep these three things in mind.
- Define the decision and the specific difference before launch.
- Do not call ordinary optimized delivery a controlled experiment.
- Keep outcome definitions, counts, disturbances, and uncertainty with the result.
How this piece was prepared. Original evaluation framework informed by primary documentation; all examples are hypothetical.
Name the decision the test will inform
Start with a choice you actually need to make. An imagined home-goods brand might be deciding whether a product demonstration explains an unfamiliar mechanism better than a styled room photograph. That is more precise than asking which ad is best. It also tells the creative team what each version must communicate.
Write the hypothesis in one sentence and identify the business outcome it is intended to affect. Google’s experiment guidance likewise emphasizes a clear hypothesis tied to a goal. Our framework adds a practical boundary: specify what the test will not resolve, such as the best offer or the right audience, when those are being held constant.
Sources: Google
Make the comparison match the question
If you want to understand the effect of a demonstration, keep the offer and central claim consistent where practical. Changing the image, price, audience, and destination together may help compare two complete campaigns, but it will not isolate the image as the cause of a difference. Name the comparison honestly.
For a causal question, use an appropriate experiment design and confirm how people or traffic are assigned. Two ads launched at the same time do not automatically receive comparable exposure. A platform may optimize delivery toward different people or moments. If the available setup is only an informal comparison, describe it that way rather than borrowing the language of a controlled test.
Choose the outcome before looking
Pick a primary outcome that matches the decision and define its denominator. Clicks per impression, purchases per landing-page visit, and revenue per order answer different questions. Record the date range, conversion definition, attribution window, and source you will use. A memorable percentage without those details is difficult to interpret.
Use secondary observations to diagnose the result, not to keep searching for a favorable metric. A comprehension check can explain what people think the creative shows; it does not establish commercial impact. A purchase result can matter commercially without telling you exactly which visual detail changed behavior. Keep the claims proportional to the measurement.
Set the review conditions in advance
Agree on the planned duration, the evidence needed for a useful decision, and conditions that would invalidate or interrupt the comparison. There is no universal number of days or conversions that makes every creative test conclusive. The required sensitivity depends on the decision, expected variation, and experiment design.
Record operational disturbances such as an unavailable product, a broken destination, or a changed offer. Do not quietly restart the clock after seeing an inconvenient result. If the data are sparse or the setup was compromised, the responsible finding may be that the test did not answer the question.
Keep a decision record, including uncertainty
At review time, write the original hypothesis, actual exposure, outcome counts and rates, and material limitations. State what you will keep, change, or investigate next. Google’s guidance recommends retaining experiment records; the value is being able to see why an earlier decision made sense with the evidence available then.
This process fits teams with a real decision and enough traffic or research access to evaluate it. When commercial evidence will be too thin, first test comprehension or usability with an appropriate method and label that narrower learning. The goal is a better next decision, not an endless sequence of declarations that every new asset is a winner.
Sources: Google
Sources & method
Original evaluation framework informed by primary documentation; all examples are hypothetical.
Sources checked Sep 6, 2026. Product capabilities can change; verify the current documentation before making a commitment.
Published by Pixel & Shelf. Prepared with AI assistance, with claims checked against the linked sources.
Our editorial approach Suggest a correction ↗