Skip to main content
AI Development

ChatGPT Ads: attributed conversions are not the same as extra sales

OpenAI’s October 5 measurement update adds partners and explores geo-based experiments. Separate conversion credit from causal lift before increasing spend.

Sean McLellan profile photo

Sean McLellan

Lead Architect & Founder

5 min read
Constructed comparison separates attribution, which assigns conversion credit, from incrementality, which asks what changed because of advertising.
Constructed diagramBaristaLabs conceptual measurement guide, not OpenAI product UI or a reported campaign result.

OpenAI’s October 5 ChatGPT Ads announcement combines a new visual format with more measurement partners. The useful distinction for an advertiser is not between an old dashboard and a new one. It is between attributed conversions and additional business caused by advertising.

Those are different questions. A report can connect a purchase to an ad interaction without establishing whether the customer would have purchased anyway. OpenAI now explicitly describes early work on incrementality, including partnerships to explore geo-based experiments. That is a reason to plan a better comparison, not evidence that your next campaign will create extra sales.

What changed, and what is still a test

OpenAI says the new visual ad format will initially be tested during image generation in ChatGPT. Testing is scheduled to begin later this month in the United States with an initial group of advertisers. Ads will be labeled and separate from the image being created; OpenAI says advertising does not influence ChatGPT’s answers. These are the company’s stated format boundaries, not results from an independent placement audit.

The measurement expansion includes integrations with Hightouch, Tealium, and LiveRamp for sending conversion data from existing systems. OpenAI also names attribution partners across web and app measurement, and full-funnel or advanced measurement partners including Fospha, Measured, and INCRMNTAL.

For incrementality, the announcement names Haus, Measured, and WorkMagic and says the work is still in its early stages. The stated goal is to explore geo-based experiments that help advertisers understand causal impact. The announcement does not say every advertiser can immediately run the same experiment through a self-service control.

Keep the visual-format rollout separate from measurement availability. A new creative opportunity, a supported reporting integration, and an experimental causal study are not interchangeable product promises.

What attribution can tell you

Attribution applies a reporting rule to eligible interactions and outcomes. It helps answer which channel gets credit for a purchase or lead under the chosen windows and matching rules. It is useful for tracing journeys and comparing reports, but the credited conversion is not automatically an incremental conversion.

OpenAI’s accompanying measurement post says advertisers can select a click-through reporting window and include a one-day view-through window. A view-through conversion follows an impression without a qualifying click. OpenAI says these settings control attribution in reports; they do not change how campaigns optimize.

That last distinction matters when reading a before-and-after dashboard. Changing the reporting window can change the credited count without showing that the business gained additional customers. Record the window, eligible event definition, and reporting period alongside the result. Otherwise two seemingly comparable acquisition-cost figures may use different rules.

Before asking a causal question, verify that the underlying conversion signal is trustworthy. Our conversion deduplication guide covers that earlier job: one business outcome should not become multiple events merely because browser and server reporting both observed it. Accurate event identity is necessary measurement plumbing. It still does not supply the missing counterfactual.

Why the early partner findings need separate labels

OpenAI reports that DV Rockerbox measured WeightWatchers’ attributed cost per acquisition on ChatGPT Ads at 15.3% below its blended paid-search benchmark. This is an attributed-cost comparison for a named advertiser, not a general causal lift estimate or a price promise for other businesses.

The announcement also says WorkMagic reported statistically significant lift for Dose, with 67% of incremental purchases coming from net-new customers. It separately reports Triple Whale’s finding that 93% of Portland Leather’s visitors from ChatGPT Ads were new. New visitors, incremental purchases, and attributed acquisition cost describe different outcomes. They should not be presented as three versions of the same success rate.

The inspected announcement does not disclose enough experiment detail to reproduce these findings: it does not provide the full assignment method, sample size, uncertainty interval, or campaign conditions. The results are publisher-reported partner findings. They may justify asking the partners for a study design; they do not justify borrowing the reported percentages for a forecast of your own campaign.

What an incrementality question requires

The causal question is: what would have happened without these ads? A geo-based experiment can try to estimate that by comparing areas with different exposure to the campaign. Simply comparing sales before and after launch is weaker: promotions, seasonality, other channels, and changes in demand can move the same outcome.

The following is BaristaLabs planning guidance, not OpenAI’s disclosed study protocol. Start with one business outcome that your systems already record consistently, such as completed orders or qualified leads. Specify the definition, geography, and observation period before results arrive. Do not switch from purchases to visits midway because the latter looks better.

Ask the measurement partner how areas will be assigned, how comparable they are before launch, and what could contaminate the comparison. Overlapping campaigns, customers moving between areas, and uneven promotions deserve explicit treatment. A holdout is useful only if it supports a credible comparison; the word alone does not make a study causal.

Proposed causal comparison defines an outcome, plans a credible holdout, and records uncertainty and stop rules before advertising spend.
Constructed diagramBaristaLabs proposed test guide, not a disclosed OpenAI experiment or observed lift result.

Agree how uncertainty will be reported and what would make the test inconclusive. A small campaign may not provide enough observations for a useful estimate. Ask the partner whether the proposed design can answer the question at your scale rather than assuming that a larger advertiser’s method transfers unchanged.

Set a spending limit and decision rule before the test. A result consistent with no detectable lift should remain a legitimate outcome. So should an inconclusive result. Neither should be rewritten as a successful experiment solely because the attribution report showed conversions.

The next step is a measurement plan, not a borrowed forecast

If you are considering the new format, first confirm account access and the US test’s actual requirements. Then keep two records: the attribution report with its event definitions and windows, and the causal comparison with its assumptions and uncertainty. They can inform the same spending decision without being collapsed into one number.

For teams already advertising, the immediate task is narrower: document which result currently justifies spend, whether it is attributed or incremental, and what evidence would change that decision. A supported integration improves the available measurement path. A credible comparison is what helps determine whether advertising added business.

Sources

Evidence-led campaign planning

Define the measurement question before increasing spend

BaristaLabs can help map one conversion path, reconcile the reporting definitions, and document what a campaign test can and cannot establish.

Bring a sanitized measurement outline, not customer records.

Turn this idea into a pilot

Which workflow should go first?

Use the readiness check to compare impact, effort, risk, owner, and next step before requesting a review.

  • 3-5 minutes
  • Deterministic score
  • No sensitive data
Check workflow readiness

Practical AI Workflow Notes

Want more practical AI operations ideas?

Get short notes on applying AI inside real small-business workflows — from document handling and customer follow-up to internal reporting, compliance, and automation guardrails.

A useful next step if you’re still exploring and not ready to request a 20-minute workflow assessment.

Occasional emails. Practical workflow guidance only. Unsubscribe anytime.