Measuring Incrementality on a Channel With No Search Term Report

By 10 min read

Summarise with

Prompt copied

Screen-print illustration of a precision balance whose two pans each hold so few grains that the beam has barely moved off level, the pointer sitting almost dead centre.

In brief

Evaluate ChatGPT Ads incrementality without user-level data using geo tests, holdouts and explicit statistical-power checks.

Last verified: 12 September 2026 | Version: 1.0 | Next scheduled review: 12 October 2026

There are three incrementality designs available on ChatGPT Ads, and for most B2B advertisers all three are underpowered at the budgets the channel currently supports. That is the finding. The rest of this piece is how each design works, what it can detect, and how to decide whether you are in a position to run one at all.

Start with what is removed. Incrementality testing normally relies on splitting users into exposed and unexposed groups and comparing outcomes. Every version of that design requires user-level data: a user ID list, a hashed identifier, a ghost-ad log, or an impression record you can join to a customer record. OpenAI provides none of these. There is no user-level export, no impression log, no PSA placeholder, and custom audiences are upload-only, meaning you can send a list in but you cannot get a treatment assignment back out.

So user-split designs are not hard on this channel. They are unavailable. What remains operates on geography and on time.

Design one: geo holdout

Split your serving markets into a treatment set and a control set, run ChatGPT Ads only in treatment, and compare outcomes.

This is possible because location targeting goes down to country, region, DMA and postal code. DMA is the working unit for the US, because it is large enough to carry volume and small enough that you can hold out ten or fifteen of them without conceding a meaningful share of the market.

Mechanically: pick paired DMAs matched on historical conversion volume, category seasonality and existing paid spend. Serve to one of each pair. Keep every other channel identical across both sets, which in practice means your Google, LinkedIn and Meta geo settings must not differ, and your sales team must not treat the two sets differently.

What it cannot do. Nothing prevents a person in a control DMA from seeing your ad. ChatGPT location targeting operates on the signals OpenAI has for the session, and a user on a VPN, travelling, or on a corporate network egressing elsewhere lands wherever their signals say. Contamination pushes the measured lift down, so a geo holdout produces a conservative estimate rather than an unbiased one. That is the safer direction to be wrong in, and it matters when you are near a decision threshold.

Where it breaks for B2B specifically. Business buyers do not respect DMA boundaries. A head office in one DMA buying for offices in six others turns a clean geographic split into a contaminated one, and there is no way to detect it from the campaign side. If your accounts are multi-site, geo splitting is measuring something other than what you intended.

Design two: matched-market testing

The same idea with the control constructed rather than randomised. Instead of holding out DMAs and comparing raw outcomes, you build a synthetic control: a weighted combination of non-treated markets whose historical outcome series tracks the treated markets closely in the pre-period, then measure the gap that opens after launch.

This is the better design when you have few markets, uneven sizes, or a strong pre-existing trend, which describes most B2B advertisers. It is also the design with the highest analytical requirement, because a synthetic control fitted to a short or noisy pre-period will fit noise and produce a confident, wrong answer.

The practical requirement is a pre-period long enough to establish fit and to validate it. The usual discipline is to fit on one span and check the fit on a held-out span before the test starts. Below roughly a year of stable weekly conversion data, the fit is not testable and you are producing a chart rather than an estimate.

Design three: spend pulsing, or switchback in time

Run the channel on and off in alternating intervals nationally, or step spend up and down, and compare outcome rates across the intervals.

The attraction is that it needs no geographic split and no control markets, so it works for advertisers too small or too concentrated for a geo design. Measurement vendors specifically recommend it alongside geo testing for privacy-constrained channels.

The problem is carryover, and on this channel carryover is the entire subject. If the dominant response to a ChatGPT ad is a remembered brand name acted on weeks later, then an off interval still contains the effects of the previous on interval, and the difference between intervals understates the effect by an unknown amount. Switchback designs assume effects decay within the interval. For a B2B cycle measured in weeks, they do not.

You can lengthen intervals to accommodate carryover, but each lengthening costs you observations, and observations are exactly what this channel cannot spare. Which brings us to the real constraint.

The power floor

Call it the power floor: the smallest true lift the channel's serving footprint can distinguish from noise at the budget you can actually spend.

Every argument above is a design argument. The binding constraint is arithmetic, and almost nobody publishing on this topic does it.

Take a simple comparison of conversion counts between two arms with equal exposure. A standard approximation for the events required per arm to detect a relative lift r at 5 percent significance and 80 percent power is 2 × (1.96 + 0.84)² ÷ (ln r)².

Relative lift you want to detect Conversions needed per arm At 40 conversions per arm per month
10 percent about 1,730 about 43 months
20 percent about 470 about 12 months
30 percent about 230 about 6 months
50 percent about 96 about 2.4 months

Read the right-hand column against a real B2B account. A company generating 200 leads a month across the US, holding out half its DMAs, is looking at roughly 100 conversions per arm per month if the split is even, and considerably fewer if the held-out DMAs are the smaller ones. At 40 per arm, a test capable of detecting anything under a 30 percent lift takes longer than the channel has existed.

Now add the budget side. The minimum daily budget is 25 USD, and circulating bid guidance sits at 3 to 5 USD per click. Spread across ten treatment DMAs, that is a handful of clicks per DMA per day. To reach the conversion counts above you either run for many months or you spend well above the minimum, and the second option is the one people take without noticing that they have converted a measurement exercise into a large budget commitment justified by a test that has not concluded.

The honest conclusion: for most B2B advertisers at this channel's current scale, a properly powered incrementality test is not affordable, and running an underpowered one is worse than running none, because an underpowered test that reads positive will be believed.

What to do instead, if you are below the floor

Three things, none of which are incrementality tests, and all of which should be labelled as what they are.

Run a geo holdout for the directional read and pre-register the threshold. Before launch, write down the lift you would need to see to continue, and the confidence interval you will accept. A wide interval that excludes zero is still information. A point estimate quoted without an interval is not.

Use branded search lift as the population-level proxy. It does not isolate causation, but it moves on the same clock as remembered demand and it is measurable with data you already have. Our post on the conversions ChatGPT Ads will never show you sets out the procedure.

Treat the channel's payback on observed conversions as the decision rule. If the reported numbers alone clear your threshold, the unobserved portion is upside. If they do not, an unverifiable uplift multiplier should not be what keeps the budget alive.

The reckoning: this advice argues against the advice

There is a tension in this article worth naming. Everything above says measure causally, and then says most readers cannot afford to. The resolution is not to run the test anyway.

It is to notice that the channel's minimum budget of 25 USD a day is small enough that the cost of being wrong about ChatGPT Ads for six months is, for most B2B advertisers, smaller than the cost of a properly powered test. When the price of certainty exceeds the value of the decision, the rational move is to spend at a level where you do not need certainty, and revisit when the serving footprint is large enough for the arithmetic to work.

That is an unsatisfying answer for anyone who reports to a board. It is the correct one.

What we cannot tell you

  • The actual serving footprint of any ad group by geography. No impression share, no reach, no frequency reporting. Absent.
  • How OpenAI resolves location for a session. Undocumented, so geo contamination cannot be estimated, only assumed to exist.
  • Whether custom audience uploads can be used to construct a holdout. Uploads are one-directional and OpenAI returns no treatment assignment. Absent by design.
  • The true decay curve of a ChatGPT ad impression. Nobody has published one, which is why switchback interval lengths are a guess.
  • Any incrementality result of our own. InPromptAds runs no campaigns and has never run one of these tests.

Quick answers

Can you run a proper incrementality test on ChatGPT Ads? Geo and time-based designs are possible. User-split designs are not, because OpenAI exposes no user-level data, no impression log and no way to read back a treatment assignment from a custom audience upload.

Why does no user-level data rule out matchback? Matchback joins a customer record to an impression or click record for the same person. ChatGPT Ads produces no exportable user-level record to join against, so there is nothing on the ad side of the join.

What geographic unit should a geo test use? DMA in the US, because location targeting supports it and it is large enough to carry volume. Region or country elsewhere, accepting that fewer, larger units make a matched design harder.

How many conversions do I need for a geo test? About 230 per arm to detect a 30 percent lift at 80 percent power and 5 percent significance, and about 1,730 per arm for a 10 percent lift. Most B2B accounts on this channel will not reach the second figure inside a year.

Is spend pulsing a valid alternative? It avoids the geographic split but assumes effects decay inside the interval. On a channel where response is often a remembered brand name acted on weeks later, that assumption fails, and the design understates the effect.

What if I cannot power a test? Say so, pre-register your decision threshold, run a directional geo read with a stated confidence interval, and set the budget at a level where being wrong for six months is cheaper than the test would have been.

Sources

Claim Source Tier
Location targeting supports country, region, DMA and postal code OpenAI Help Center, campaign setup documentation, 2026 Confirmed, primary
Custom audiences are first-party upload only OpenAI Help Center, audiences documentation, 2026 Confirmed, primary
No user-level data, no impression log, no search terms report, no impression share OpenAI Ads Manager documentation, by absence Absent
Minimum daily budget 25 USD OpenAI Help Center, billing and budgets, 2026 Confirmed, primary
Bid guidance of 3 to 5 USD CPC circulating in 2026, not published by OpenAI as a floor Multiple agency write-ups, 2026 Reported
Geo holdouts and spend pulsing recommended as the standard approaches for privacy-constrained channels; controlled experiments described as the only causal source in year one Measured, ChatGPT Ads incrementality guidance, 2026. Vendor sells incrementality measurement Vendor
Ads serve to Free and Go plan users in the reported serving markets OpenAI Help Center, 2026 Confirmed, primary
Power calculation, the 2(z+z)²/(ln r)² approximation and the resulting sample requirements Inference, ours. Standard two-sample Poisson approximation, equal exposure assumed Inference, ours
  • The Conversions ChatGPT Ads Will Never Show You
  • How to Measure ChatGPT Ads, and What You Cannot Measure Yet
  • What OpenAI's Attribution Windows Actually Count
  • Location Targeting in ChatGPT Ads
  • Do ChatGPT Ads Work?

Changelog

12 September 2026, v1.0. First publication. Establishes the power floor as the binding constraint on causal measurement for this channel, with the sample requirements worked out, and states the case for not running an underpowered test.

Field kit

Tools

Site

OpenAI (primary)

Share this

AK

Ansh works across GEO strategy, B2B research, and execution. At InPromptAds, he translates new AI advertising products into clear operating advice, tests, and measurement questions for marketing teams.

Get placed before the self-serve rush.

We're onboarding a founding cohort of B2B brands for managed ChatGPT ad placements. Early access, limited seats.

Now capturing early access
Contact us Contact us