Keyword List, Full Question or Narrative: The Only Public Test Went the Wrong Way

By 9 min read

Summarise with

Prompt copied

Screen-print illustration of a balance scale tipped decisively down on the side holding a plain scatter of loose beads, while the raised pan holds one elaborately carved ornament.

In brief

The only published context-hint test found keyword-style hints outperformed conversational questions. Here is what one test can prove.

Last verified: 12 September 2026 | Version: 1.0 | Next scheduled review: 12 October 2026

There are three ways people write context hints, and one published test comparing any of them. The test found the style everybody calls wrong is the one that won.

This article does three things: sets out the three styles precisely enough to test, reports the Search Engine Land result without softening it, and gives you a protocol to run the comparison yourself, because one test is not a finding.

The three styles, defined

Keyword list. Disconnected terms, comma separated, lifted from a Google Ads account or a keyword tool. payroll software, small business payroll, RTI submission, payroll for 5 employees, switch payroll provider.

Full question. A complete natural-language query of the kind a user would type. What payroll software should a five person UK company use if we keep missing RTI deadlines?

Narrative. A described situation carrying an audience, an intent and a constraint, which is the form OpenAI's own documentation illustrates. A UK limited company running payroll on spreadsheets who has just taken on their third employee and needs RTI submissions filed automatically rather than reminded about.

Published advice in this category, including OpenAI's own worked example of "cushioned everyday running shoes for beginners training for their first 5K" against the bare term "running shoes", points at narrative. The category has treated that as settled.

What the test found

In 2026 Search Engine Land ran one campaign with two ad groups and identical creative in both. One group held a Google-style keyword list. The other held full natural-language questions.

The keyword group served far more often, at a slightly higher click-through rate and a slightly lower average cost per click.

That is the whole public evidence base on hint style. One advertiser, one category, one moment in time, reported by a third party, not replicated by us or, as far as we can find, by anyone.

Treat that sentence as a limitation on the result, not as a way of dismissing it. It is still the only measurement anyone has published, and it points the opposite way to every piece of guidance sitting above it in the search results, including parts of our own.

Four things that could produce that result

Taking the result seriously means proposing mechanisms that would explain it, then noticing which ones are testable.

One. The matcher is less style-sensitive than the advice assumes. If OpenAI is embedding the hint and comparing meaning, a dense keyword list carries a lot of semantic surface per character. "payroll, RTI, small business, switch provider" names four concepts in nine words. The narrative version names the same four across thirty-five words with connective tissue that adds no matching signal. On that reading, style does not matter and density does.

Two. Breadth is being scored as performance. A keyword list contains no constraints, so it matches a wider set of conversations. Wider matching means more impressions. More impressions at the top of a funnel usually means a mix shifted toward cheaper, easier clicks, which can raise CTR and lower CPC while lowering the quality of everything downstream. The test reported serving, CTR and CPC. It did not report conversion rate or cost per qualified lead, and those are where breadth normally shows up as a cost.

Three. Relevance weighting rewarded the denser input. The auction is relevance weighted. If the keyword group matched more tightly to a larger set of conversations, it would both serve more and pay less, which is exactly the shape of the reported result.

Four. The questions were written badly. A full question is the weakest of the three styles on its face, because it describes one phrasing of one need. If the conversational group was built from single questions rather than described situations, the test may have compared a keyword list against the worst version of the alternative, rather than against narrative hints.

Mechanism two is the one that should worry you, and mechanism four is the one that most limits what the test proves. Neither can be settled from the published write-up.

A protocol you can actually run

This is designed so the result is readable given what ChatGPT Ads reports, which is nothing below ad level.

Hold constant. One campaign. One objective, Clicks or Conversions, not Reach. One bid. One set of locations and platforms. Identical creative in every group: same headline, same description, same image, same destination URL. Identical everything is not a nicety here, because with no hint-level reporting the only way to attribute a difference to hint style is to have changed nothing else.

Vary one thing. Three ad groups, one per style, describing the same single buying situation. Not three situations. If the keyword group covers four themes and the narrative group covers one, you have tested breadth, which is what makes the original result ambiguous.

Match semantic content across arms. Write the narrative hint first. List the concepts it contains. Build the keyword arm from those concepts and no others. Build the question arm from the same concepts phrased as queries. Otherwise you are testing what you happened to put in each box.

Give it enough spend. At reported CPC guidance of three to five dollars, a group needs roughly 100 clicks before its CTR and CPC are worth comparing, which is 300 to 500 dollars per arm, so 900 to 1,500 dollars across three arms. Minimum daily budget is 25 dollars per campaign. Below about 50 clicks per arm you are reading noise.

Run it for a fixed number of days, not until it looks decided. Two weeks minimum. Stopping when one arm looks good is how you generate a result that does not replicate.

Record the metrics that matter, not the ones that move first. Impressions and CTR move first and are the least informative. The comparison that decides anything is cost per qualified lead, which means the arms have to be separable in your CRM. Different destination URL parameters per ad group, set on the ad, are the only way to do that, because the platform will not tell you.

Expect the arms to serve unevenly. If one arm takes 70 percent of impressions, that is itself the finding, and you should report it as a delivery result rather than treating the other arms as underperforming.

What we would predict, and why we could be wrong

Our position, and it is a prediction rather than a finding: keyword lists will win on impressions and CTR, narrative hints will win on cost per qualified lead, and full questions will lose on both because they carry the narrowness of a narrative without its coverage.

That prediction is consistent with the Search Engine Land result, since that test measured only the metrics where we would expect keywords to win. It is also unfalsified, which is a polite way of saying unsupported. InPromptAds has no first-party data on hint style. We run no campaigns and hold no account data, so everything above is reasoning from documented auction mechanics and one third-party test.

If you run the protocol and get the opposite, the opposite is better evidence than this article.

What we cannot tell you

  • Whether the Search Engine Land result replicates. One test, one advertiser, one category, never repeated publicly.
  • What the test's conversion rates were. Not reported, and conversion is where the breadth explanation would show itself.
  • How OpenAI processes hint text. Whether hints are embedded, parsed, or used as retrieval input is entirely undocumented.
  • Whether hint style affects approval. No published data on rejection rates by hint style, or on rejection rates at all.
  • Our own results. InPromptAds has no first-party campaign data on this or any other ChatGPT Ads question.

Quick answers

Do keyword lists work as ChatGPT Ads context hints? One published test in 2026 found a keyword-style ad group served far more, with slightly higher CTR and slightly lower CPC, than a conversational one with identical creative. That is a single third-party test, not a settled answer.

Which hint style should I use? Narrative, describing a situation with an audience, intent and constraint, unless and until you have run your own comparison. The one contrary data point measured metrics that favour breadth and did not report downstream quality.

Why does the public test contradict the published advice? Four plausible reasons: the matcher may be indifferent to style, keyword lists buy breadth which flatters CTR and CPC, relevance weighting may reward semantic density, or the conversational arm may have been built from single questions rather than described situations.

How much should a hint style test cost? At reported three to five dollar CPCs, budget 300 to 500 dollars per arm to reach roughly 100 clicks, so 900 to 1,500 dollars for a three-arm test. Below 50 clicks per arm the numbers are noise.

Can I tell which hint inside a group did the work? No. There is no hint-level reporting, which is why a style test has to put each style in its own ad group with identical creative.

Does InPromptAds have data on this? No. We have no first-party campaign data. Everything in this article is either documented platform behaviour, the Search Engine Land test, or our reasoning labelled as such.

Sources

Claim Source Tier
Keyword-style ad group outserved a conversational one, with higher CTR and lower avg CPC, identical creative Search Engine Land, how to run ChatGPT ads, 2026 Reported, third party, single test
Hints should use clear natural phrases describing genuine use cases; "cushioned everyday running shoes for beginners training for their first 5K" OpenAI Help Center, Create Ad Groups for ChatGPT Ads, August 2026 Confirmed, primary
Auction is relevance weighted OpenAI documentation, 2026 Confirmed, primary
No hint-level reporting exists OpenAI Ads Manager documentation, by absence Absent
Minimum daily budget of 25 USD OpenAI Ads Manager documentation, 2026 Confirmed, primary
Three to five dollar CPC bid guidance Multiple agency write-ups citing OpenAI guidance, 2026 Reported
Four candidate mechanisms and the three-arm protocol InPromptAds reasoning Inference, ours
  • Context Hints Are Not Keywords: The Translation Table
  • How to Write a ChatGPT Ads Context Hint
  • How Many Context Hints Should One Ad Group Have?
  • How to Measure ChatGPT Ads, and What You Cannot Measure Yet
  • Budget Calculator

Changelog

12 September 2026, v1.0. First publication. Records the Search Engine Land hint style result in full, proposes four mechanisms that could produce it, and publishes a three-arm replication protocol.

Field kit

Tools

Site

OpenAI (primary)

Share this

AK

Ansh works across GEO strategy, B2B research, and execution. At InPromptAds, he translates new AI advertising products into clear operating advice, tests, and measurement questions for marketing teams.

Get placed before the self-serve rush.

We're onboarding a founding cohort of B2B brands for managed ChatGPT ad placements. Early access, limited seats.

Now capturing early access
Contact us Contact us