When to Kill a ChatGPT Ads Test
By Keshav Parsai 9 min read
In brief
Set a ChatGPT Ads stopping rule before launch using minimum spend, conversion evidence and operational failure conditions.
Last verified: 12 September 2026 | Version: 1.0 | Next scheduled review: 12 October 2026
Write the stopping rule before you launch. Not because pre-commitment is a virtue, but because ChatGPT Ads gives you no benchmark, no query data and small samples, which means the decision to stop will otherwise be made on whichever number happens to look worst on the day somebody asks.
What a readable test actually costs
The minimum daily budget is 25 USD. Work forward from there.
Bid guidance circulating in 2026 puts CPC at three to five dollars, and the one named test in the trade press sits above that range: MediaPost reported on 31 August 2026 that Peter Jaffray of Choice OMG recorded roughly 7 USD per click across 415 USD of June 2026 spend. Take four to seven dollars as a working band.
At 25 USD a day that is roughly four to six clicks daily. A fortnight buys somewhere near 50 to 90 clicks.
Now attach a conversion rate. A B2B landing page converting cold paid traffic at two percent turns 90 clicks into fewer than two conversions. At one percent, fewer than one. That is not a conversion rate you have measured. It is a number you have failed to measure, and the two look identical in a report.
To accumulate ten conversions, which is roughly the floor at which a conversion rate stops being a coin flip, the same arithmetic gives you 500 clicks. At the 25 USD minimum that is somewhere between two and three months and 1,500 to 2,500 USD of spend, per readable unit. If you have split into three ad groups, multiply.
That number is the single most useful thing on this page. A decision-grade ChatGPT Ads test is a four-figure commitment over months, or it is a delivery test wearing a performance test's clothes.
Two failures that look the same
A channel that does not work and a test that was never readable produce identical reports: low or zero conversions, spend gone, nothing conclusive.
The MediaPost round-up in August 2026 is a worked example of the ambiguity. Agencies reported zero conversions across their ChatGPT Ads tests, on spends in the hundreds of dollars, with one test at 415 USD generating 60 clicks. Sixty clicks producing zero conversions is entirely consistent with a two percent conversion rate. It is also entirely consistent with a zero percent conversion rate. The data does not separate them and no amount of staring at it will.
Most advertisers reading that report concluded the channel does not convert. The defensible conclusion is narrower: those tests could not have detected conversion at a rate any B2B advertiser would consider normal.
This is not a defence of the channel. It might well not convert. It is a statement about what the published evidence can carry, and it applies with the same force to your own account.
The three rules, written before launch
Write these into a document before the campaign goes live, with numbers filled in, and date it.
Rule one: the delivery kill. If the campaign has not delivered a stated impression volume by a stated day, stop and treat the result as a delivery finding rather than a performance one. This one is cheap and fires early. Something like: fewer than 2,000 impressions by day seven with all eligibility checks passing, stop, and record that inventory for this theme was insufficient at this bid.
A delivery kill protects you from the most common waste on this channel, which is running a campaign for six weeks that was never going to accumulate a readable sample and then debating what its CTR means.
Rule two: the spend ceiling. State the total you are willing to spend before any conclusion is drawn, and state it as a number, not as a monthly rate that quietly renews. If your readable-sample arithmetic says 2,000 USD, the ceiling is 2,000 USD. If you are not willing to commit 2,000 USD, you are not running a conversion test, and the honest thing is to define the test as a delivery-and-cost-per-click test and cap it at a few hundred dollars with that stated up front.
Rule three: the qualified-lead threshold. Not conversions as the platform counts them. Qualified leads as your own CRM counts them, at a stated cost. If cost per qualified lead exceeds a number you set in advance, you stop.
Set that number against your other channels, because there is nothing else to set it against. OpenAI publishes no cross-advertiser benchmarks. If Google non-brand delivers a qualified lead at 180 USD, a ceiling of two or three times that is a defensible starting position for a new channel, and the important property of the number is that you chose it before you had a reason to want it to be higher.
Why judgement fails here specifically
On a mature channel, an experienced buyer's judgement is a compressed form of prior data. They have seen a thousand accounts, so a 0.4 percent CTR triggers an accurate instinct.
That mechanism does not work on ChatGPT Ads and will not for a while. The priors do not exist. OpenAI publishes no benchmarks. The published third-party figures are a handful of named tests with spends in the hundreds, plus a layer of vendor benchmark tables with unstated samples. Delivery itself changed materially during 2026, with Digiday reporting fill rates climbing 30 to 50 percent from launch levels by late April, so even the existing figures describe a moving target.
An experienced buyer applying Google-calibrated instincts to this channel is not exercising judgement. They are applying the wrong prior with high confidence, which is worse than applying no prior at all. Pre-committed rules are not a substitute for expertise here. They are what expertise looks like when the data to be expert about does not exist yet.
The reckoning: pre-commitment can be wrong too
A rule written on day zero was written by someone who knew less than you do on day thirty. That is a genuine cost and pretending otherwise would be dishonest.
Sometimes a campaign should be stopped early for a reason nobody anticipated, and sometimes it should be extended because you learned something mid-flight that changes the calculation. The rule should not override that.
What the rule does is change the burden of proof. Overriding a written rule requires you to state what new information justifies it, in writing, before you act. "It feels like it is starting to work" is not new information. "The pixel was misconfigured for the first nine days and the conversion data is invalid" is. Most overrides fail that test, which is the point.
What we cannot tell you
- Whether ChatGPT Ads converts for your category. No advertiser has published a conversion-rate study, and the named tests in the trade press were too small to detect normal B2B conversion rates.
- How long a decision-grade sample takes. Depends on your conversion rate, which you do not know before you test. The arithmetic here is inference from the 25 USD floor and reported CPCs.
- What cost per qualified lead is achievable. No cross-advertiser benchmarks exist, so the threshold has to come from your other channels.
- Whether current delivery matches the reported 2026 figures. Fill rates were reported improving through the year with no current published figure.
- What InPromptAds has seen. Nothing. No client campaigns have run, so there is no first-party stopping-rule data behind this article.
Quick answers
How long should I run a ChatGPT Ads test before stopping? Long enough to reach the sample your stopping rule requires, decided before launch. At the 25 USD minimum and reported four to seven dollar CPCs, a fortnight buys roughly 50 to 90 clicks, which measures cost per click and cannot measure conversion rate.
How much does a real ChatGPT Ads test cost? To accumulate roughly ten conversions at a two percent landing conversion rate, around 500 clicks, which is 1,500 to 2,500 USD over two to three months per readable unit at the minimum budget. Splitting into several ad groups multiplies it.
Does zero conversions mean the channel does not work? Not at small spends. Agencies reported zero conversions on ChatGPT Ads tests in the hundreds of dollars in August 2026, but 60 clicks producing zero conversions is consistent with both a two percent and a zero percent conversion rate.
What should the stopping rule measure? Three things: an impression volume by a stated day, a total spend ceiling, and a cost per qualified lead from your own CRM rather than the platform's conversion count.
Can I just use my judgement instead? Judgement on a new channel is a Google-calibrated prior applied where it does not hold. With no published benchmarks and delivery still shifting through 2026, a written rule is more defensible than an instinct.
When is it right to override the rule? When you can state, in writing and before acting, what new information invalidates the assumption the rule was built on. A misconfigured pixel qualifies. A sense that it is starting to work does not.
Sources
| Claim | Source | Tier |
|---|---|---|
| 25 USD minimum daily budget, no minimum lifetime commitment | OpenAI Ads Manager, 2026 | Confirmed, primary |
| Three to five dollar CPC starting bid guidance | Agency write-ups citing OpenAI guidance, 2026 | Reported |
| 415 USD across three campaigns, 9,000 impressions, 60 clicks, roughly 7 USD per click, June 2026 | MediaPost, 31 August 2026, citing Peter Jaffray of Choice OMG | Reported, third party, single test |
| Zero conversions reported across agency test campaigns | MediaPost, 31 August 2026 | Reported, third party |
| Fill rates up 30 to 50 percent from launch levels by late April 2026 | Digiday, 2026, anonymous sources | Reported, third party |
| No cross-advertiser benchmarks published | OpenAI Ads Manager, by absence | Absent |
| Clicks and spend required for a ten-conversion sample | Arithmetic from the 25 USD floor and reported CPC band | Inference, ours |
| InPromptAds has no first-party test data | InPromptAds | Confirmed, primary |
Related reading
- Reading CTR and CPC When No Benchmarks Exist
- The First Fourteen Days of a ChatGPT Ads Campaign, Decision by Decision
- Who Should Not Buy ChatGPT Ads Yet
- Why Your ChatGPT Ads Campaign Is Not Spending
- Budget Calculator
Changelog
12 September 2026, v1.0. First publication. Prices a decision-grade test from the 25 USD floor, separates a channel that does not work from a test that was never readable, and sets a written-override standard for pre-committed stopping rules.
Field kit
Tools
Site
OpenAI (primary)
Keshav studies how AI systems retrieve, verify, and cite brand information. At InPromptAds, he leads source research and turns platform documentation into practical guidance for advertisers.