Two or Three Ads Per Ad Group, Because Ad Count Is Your Only Experiment

By 8 min read

Summarise with

Prompt copied

Screen-print illustration of six test tubes in a rack, two of them filled to a clearly marked line and the other four holding only a shallow film at the bottom.

In brief

Choose how many ChatGPT ads to run per ad group using a two-arm test design that preserves volume and interpretable results.

Last verified: 12 September 2026 | Version: 1.0 | Next scheduled review: 12 October 2026

Start from what gets reported. Ads Manager gives impressions, clicks, spend, CTR, average CPC, average CPM and conversions at campaign level, ad group level and ad level. Nothing below that. No hint-level breakdown, no conversation data, no adjacency, no search terms.

Ad level is therefore the finest grain of truth available anywhere in this platform. Which means the ads inside an ad group are the only place you can run a controlled comparison, and how many you put in is a question about statistical power, not about coverage.

Most advertisers answer it the way they answer it on Meta: upload six, let the system sort it out. That answer imports an assumption this platform does not support.

What an ad group can and cannot teach you

Everything above the ad is confounded. If ad group A outperforms ad group B, the cause could be the hints, the creative, the bid, or the interaction between them, and you cannot separate those without running a clean test.

Inside an ad group, the hints, bid and destination default are held constant by the structure itself. So a difference between two ads in the same group is attributable to the thing you changed in the ad. That is the only clean experiment the platform gives you for free.

This is why ad count is design rather than housekeeping. Each additional ad is another arm, and every arm divides the same impressions.

Vary one thing

An ad has four changeable parts: headline, description, image, destination URL. A readable test changes exactly one of them across variants.

Headline test. Same image, same description, same URL. Two headlines carrying different claims, not different phrasings of the same claim. "RTI filed automatically" against "£29 a month payroll" tests mechanism against price, which is a decision you actually have to make. "RTI filed automatically" against "Automatic RTI filing" tests word order, which is not.

Image test. Same headline, same description, same URL. Single object against type card is the test worth running first, because the two formats are genuinely different bets and neither has published evidence behind it.

Description test. Same headline, same image. Useful mainly for testing whether a disqualifying line, a price or a size qualifier, changes cost per qualified lead. Expect CTR to fall and judge it downstream.

Destination URL test. Ad-level URLs override the ad group default. This is also the mechanism you need for separating arms in your own analytics, since the platform will not pass you anything a CRM can join on. Parameter every variant differently even when the URL test is not the test.

Changing two things at once produces a result you cannot act on. With four variants covering two headlines crossed with two images, you would need roughly four times the spend to read the interaction, and you will not have that spend.

What a readable comparison costs

Here is the arithmetic that decides the number, and it is less forgiving than the advice usually admits.

To distinguish a 1.0 percent CTR from a 1.5 percent CTR with conventional confidence, you need on the order of 8,000 impressions per variant. That is a standard sample size approximation, not a platform figure.

Turn that into money using reported bid guidance. Circulating guidance sits around 3 to 5 dollars CPC and a 60 dollar default maximum CPM, both reported rather than published by OpenAI as floors.

Variants Impressions needed Approx spend at 60 USD CPM Days at 25 USD/day minimum
2 16,000 960 USD 38
3 24,000 1,440 USD 58
4 32,000 1,920 USD 77
6 48,000 2,880 USD 115

Read the right-hand column before the middle one. At the 25 dollar daily minimum, a six-variant ad group takes about four months to produce a readable comparison, by which time the creative, the market and quite possibly the platform have all changed. The test finishes after the question stopped mattering.

That is the case for two or three ads per ad group. Not because more ads are harmful, but because arms you cannot fund are arms you cannot read, and an unreadable test is worse than no test because it produces a number that feels like evidence.

If your budget is genuinely large, say 300 dollars a day, four arms become readable inside a fortnight and the calculus changes. Scale the number of variants to the spend, not to the number of ideas you have.

Detecting a difference in conversions costs far more

The table above is for CTR, which is the cheapest metric to move and the least important one.

Conversion rate differences need much larger samples, because conversions are rarer. If your click to conversion rate is 5 percent, distinguishing 5 percent from 7 percent needs several thousand clicks per arm, which at a 4 dollar CPC is well past 10,000 dollars per variant.

The practical consequence: ad-level tests on this platform are CTR tests, and CTR tests tell you which ad gets clicked, not which ad gets you customers. Use conversion data at ad level as a directional check for a variant that is clearly failing, and do not treat a conversion difference between two ads as a result unless the click volume behind it is in the thousands.

The reckoning: rotation is not in your control

All of the above assumes the platform divides impressions roughly evenly between variants in an ad group. Nothing in OpenAI's documentation says it does.

Most auction-based platforms optimise delivery toward the better-performing creative, which is good for performance and destructive for testing, because the winner accumulates impressions once it starts winning and the loser never gets the sample that would have proved it. If ChatGPT Ads does this, and there is no published statement either way, your two-arm test is not two arms of 8,000. It is one arm of 14,000 and one of 2,000, and the second one never reaches readability.

There is no rotation setting documented, so there is no way to force even delivery. The partial defence is to check the impression split in ad-level reporting early and often. If one variant has taken more than about 70 percent of impressions inside the first few days, stop treating the comparison as a test and treat it as a delivery signal instead: the platform has told you something, just not the thing you asked.

What we cannot tell you

  • Whether impressions are split evenly across ads in an ad group. Undocumented. No rotation control is exposed.
  • Whether a maximum number of ads per ad group exists. Not published, and nobody has reported hitting a cap.
  • Whether more variants improve delivery. Undocumented. The claim circulates by analogy with other platforms, which is not evidence.
  • How long the platform takes to favour a winning variant. Undocumented.
  • Our own test results. InPromptAds runs no campaigns and has no first-party data on variant counts or rotation.

Quick answers

How many ads should I put in a ChatGPT ad group? Two or three for most budgets. The limit is statistical, not structural: each variant needs roughly 8,000 impressions to produce a readable CTR comparison, which at reported CPMs is around 480 dollars per arm.

Is there a maximum number of ads per ad group? None is documented and no advertiser has reported a cap. The constraint that binds is spend, not the interface.

Why not upload six variants like on Meta? Because six arms need roughly 48,000 impressions to read, which is about four months at the 25 dollar daily minimum. The test finishes long after the answer would have been useful.

What should I vary between ads? One element: headline, image, description or destination URL. And vary claims rather than phrasings. Two ways of saying the same thing produce a result you cannot act on.

Can I control ad rotation? No rotation setting is documented. Check the impression split in ad-level reporting within the first few days, and treat a heavily skewed split as a delivery finding rather than a failed test.

Can I test conversion rate between two ads? Rarely. Conversion differences need thousands of clicks per arm, which is five figures of spend per variant at reported CPCs. Ad-level tests on this platform are realistically CTR tests.

Sources

Claim Source Tier
Reporting covers impressions, clicks, spend, CTR, avg CPC, avg CPM and conversions at campaign, ad group and ad level OpenAI Ads Manager documentation, 2026 Confirmed, primary
Ad consists of headline, description, image and destination URL, with ad-level URL overriding the ad group default OpenAI Ads Manager documentation, 2026 Confirmed, primary
No hint-level reporting, no conversation data, no adjacency reporting OpenAI Ads Manager documentation, by absence Absent
Minimum daily budget 25 USD, no minimum lifetime commitment OpenAI Ads Manager documentation, 2026 Confirmed, primary
Bid guidance of 3 to 5 USD CPC and 60 USD default max CPM Multiple agency write-ups citing OpenAI guidance, 2026 Reported
No ad rotation control is exposed or documented OpenAI documentation, by absence Absent
Sample size arithmetic and the variants-to-spend table InPromptAds calculation from standard sample size approximation and reported CPM guidance Inference, ours
  • How Many Context Hints Should One Ad Group Have?
  • How to Measure ChatGPT Ads, and What You Cannot Measure Yet
  • Keyword List, Full Question or Narrative: The Only Public Test Went the Wrong Way
  • Budget Calculator
  • ChatGPT Ads Bidding

Changelog

12 September 2026, v1.0. First publication. Reframes ad count as experimental design, publishes the variants-to-spend arithmetic, and records undocumented ad rotation as the assumption the whole method rests on.

Field kit

Tools

Site

OpenAI (primary)

Share this

KP

Keshav studies how AI systems retrieve, verify, and cite brand information. At InPromptAds, he leads source research and turns platform documentation into practical guidance for advertisers.

Get placed before the self-serve rush.

We're onboarding a founding cohort of B2B brands for managed ChatGPT ad placements. Early access, limited seats.

Now capturing early access
Contact us Contact us