New IAB Framework Aims to Standardize Measurement of AI-Powered Visibility

Marketers have long lived by the axiom “You can’t improve what you can’t measure.” But although 50% of consumers today use AI-powered search, according to McKinsey, and such search will influence an estimated $750 billion in consumer spend by 2028, only 16% of brands systematically measure AI-powered discoverability.

The lack of a standardized measurement framework certainly makes tracking AI search performance challenging. Every provider of AI-visibility measurement tools seemingly has its own definitions, metrics and methods. Not only is knowing what to measure and how to do so difficult, but trying to compare vendors can become an exercise in frustration.

That’s why trade group IAB published “Measuring Visibility in the AI Era,” which sets out to codify the definitions, quality criteria and standards relevant to measuring GEO, AIO and other aspects of non-paid AI-powered discoverability. Marketers can use the report to determine the relevant questions about models, data sources, attribution logic and other mechanisms and metrics to include in RFPs, says Caroline Giegerich, VP, AI at IAB. Then, if providers report different AI visibility scores for a brand during the process, the brand can knowledgeably follow up to pinpoint the reasons for the discrepancies as well as gauge the transparency of each vendor.

Just as important, the report demystifies which types of metrics will help improve which sorts of decisions, ultimately encouraging more marketers to measure AI-powered visibility. Failure to measure is among the gravest mistakes a brand can make in this new world of AI discoverability, Giegerich says, and much of it stems from “not knowing what quality measurement looks like.”

The 4 P’s of AI Visibility

Giegerich suggests that brands focus on the 4 P’s of AI visibility:

  • Presence, or whether the brand is appearing in searches. Mention rate, citation rate (how often the brand is linked to), share of voice and visibility momentum (how presence is changing over time compared with past performance and that of competitors) are the key metrics here.
  • Prominence, or where the brand is appearing. Is it the only brand featured in searches? If it’s part of a list, is it cited first, fifth or last? Does it appear in the first few sentences or several paragraphs deep? Does it show up just once in a response or multiple times?
  • Portrayal, or the context in which the brand appears. This includes not just how positively or negatively the brand is described but also how factually correct the mentions are. “If I were a brand, this is what would keep me up at night,” Giegerich Imagine the damage to a brand if it is referenced as selling high-end menswear when it actually specializes in affordable womenswear. Given that AI models feed on data from other AI sources, a single inaccurate citation can spread as quickly and pervasively as crabgrass, and like crabgrass it can be difficult to eliminate. For that reason, finding the source of any errors is imperative, as is knowing measurement providers’ hallucination and factual inaccuracy rates per platform as well as how the vendors detect and monitor such errors.
  • Persuasion, or whether the brand’s visibility is driving actions. This is the most difficult to measure, Giegerich says, though post-citation clickthrough rates are a reasonable metric. Not all platforms, however, provide CTR data.

Directional vs. Decision-Grade

When assessing the value of visibility metrics, brands should distinguish between directional and decision-grade measurements. While the former are valid for detecting emerging signals and trends, they lack the sample size, data validation, methodology documentation and other quantification and precision elements to support decisions requiring a significant investment of resources.

Although the value of decision-grade measurements is obvious, a vendor should provide both types. Because directional measurements require less data, fewer queries and less validation, they cost less to run. “The better the measurements, the higher the computing costs,” Giegerich notes.

Running more frequent, less expensive directional queries can enable you to fine-tune the decision-grade queries you opt to run, saving you time and money over the long term. Understanding the measurement methodologies and inputs helps to ensure you’re making decisions based on valid, reliable outputs.

Suggestions as Opposed to Mandates

For all of the report’s definitions and standards, the IAB still lacks a means of making measurement providers comply with the suggested framework. The group is discussing a certification process as the next step, says Giegerich.

However, she believes vendors that voluntarily comply are giving themselves an advantage over competitors. A provider that fails to specify its query sets, prompt types, data validation methods and other significant information in an RFP will come across as far less credible than one that does.