Brands entering AI Search are discovering that the familiar ambition to “rank number one” does not translate cleanly into conversational discovery. A new company benchmark from SEOPulse puts an unusually sharp number on the problem: across 1,500 commercial prompts, the company says major AI engines selected the exact same top vendor in fewer than 1.5% of queries.
The finding was highlighted in a NetContentSEO analysis published September 3 and appeared alongside SEOPulse's launch announcement for an enterprise AI Visibility & Intelligence Platform. The reported benchmark covered commercial prompts and is being used by the company to argue for monitoring brand recommendations across systems including ChatGPT, Gemini, Perplexity, DeepSeek and ByteDance's Doubao.
The headline number deserves an important qualification. It is company-reported research released as part of a product announcement, and the public release does not provide enough methodological detail to reproduce the 1.5% result independently. It should therefore be treated as a benchmark from SEOPulse's test set rather than a universal law of AI Search. Even with that caveat, the underlying issue is increasingly difficult to ignore: different answer engines can construct materially different competitive realities for the same market.
There is no universal number-one position in AI Search
Traditional SEO has never produced identical results for every user. Location, device, language, personalization and query interpretation all introduce variation. Yet the industry still operates around a relatively stable abstraction: a ranked search-results page where positions can be observed and tracked.
Generative systems break that abstraction. ChatGPT may recommend one vendor first, Gemini another and Perplexity a third. The answer can change again when the wording of the prompt changes, when a model receives fresher web data, when the user's location changes or when the platform updates its retrieval and generation systems.
Even “first place” is difficult to define consistently. One assistant may explicitly rank three companies. Another may describe several alternatives without declaring a winner. A third may recommend one brand while citing a publisher that discusses a different set of vendors. Brand mention, recommendation and citation are related forms of visibility, but they are not the same event.
This is the practical extension of the shift Galloni.net described in the transition from SEO to GEO: being technically discoverable is no longer the end of the visibility problem. A brand must also be selected inside the generated answer, and that selection can vary by engine.
Different AI engines operate in different information environments
Cross-engine disagreement should not be surprising once the systems are treated as separate products rather than interchangeable interfaces to one intelligence layer. They can use different foundation models, retrieval indexes, ranking components, system instructions, freshness policies and source-selection logic.
Some systems may lean heavily on live web retrieval. Others can combine web results with prior model knowledge. Commercial and product queries may activate specialized data sources or shopping components. Regional AI products can also operate inside different publishing and language ecosystems.
The result is a fragmented discovery market. A brand that appears highly visible in ChatGPT can be weak in Gemini for the same topic. A company prominent in English-language answers may perform differently when the equivalent question is asked in Chinese, Spanish or another language. Measuring one engine and labeling the result “AI visibility” can therefore conceal significant blind spots.
Independent research supports the direction, but not one universal percentage
Other AI visibility studies reinforce the broader idea that engine agreement varies substantially, while also showing why marketers should resist turning any single benchmark into a fixed industry constant. Mapou's 2026 State of AI Search research, for example, tracks hundreds of brands across multiple engines and reports considerable divergence in a number of commercial segments. Its results also show that some categories produce much stronger consensus than others.
A separate AIVO footwear study illustrates the opposite extreme. Across its sneaker and footwear benchmark, all five tested engines selected Crocs as the leading brand overall. The point is not that one study disproves another. Their prompt sets, categories, models, methodology and definition of leadership differ. The contrast demonstrates exactly why methodology disclosure is essential in AI visibility research.
“How often do AI engines agree?” cannot have one useful answer without specifying what was asked, which engines were tested, when the tests ran, which markets were represented and how a top recommendation was defined.
Prompt tracking is becoming the AI equivalent of rank tracking
SEOPulse's product positioning reflects an important shift in measurement. Instead of recording a keyword and a numerical search position, AI visibility tools increasingly record a prompt, market, language, engine and response. The system can then extract which brands appeared, which were recommended, which sources were cited and how those outcomes changed over time.
That makes the measurement unit much richer than a conventional rank. “Best CRM” is not equivalent to “Which CRM is best for a 200-person European SaaS company that needs Salesforce integration and EU data residency?” Both concern the same product category, but the second prompt expresses constraints that can completely change the recommendation set.
For enterprise GEO, this means a useful monitoring panel needs to represent actual customer decisions rather than merely translating a keyword list into questions. Discovery prompts, comparison prompts, pricing questions, use-case constraints and high-intent vendor-selection prompts can each reveal different competitive positions.
AI share of voice is more useful than a single position
If the answer varies across systems and sessions, a single ranking number becomes fragile. Share of voice provides a more durable model. A brand can measure how often it appears across a controlled set of relevant prompts, how often it is recommended, which competitors appear instead and which engines are strongest or weakest.
The distinction between appearance and recommendation remains essential. A brand may be mentioned because it is the market incumbent but not selected as the best choice. It may be cited as the source of a technical fact without being recommended commercially. It may be recommended while a third-party review site receives the citation.
SEOPulse reports that more than half of the AI answers in its benchmark mentioned brands, while only 10% both mentioned and cited a brand. The public release does not provide the underlying dataset needed to independently validate that figure, but the conceptual distinction is important. Referral traffic captures only one part of AI influence. A recommendation can shape a buyer's consideration without sending a measurable click to the company's website.
The source behind the answer may be more actionable than the answer itself
For SEO, digital PR and content teams, one of the most useful questions is not simply “Did the AI mention us?” It is “What evidence appears to be helping the AI choose the companies it mentions?”
If competitors repeatedly win recommendations when an engine cites specialist review sites, industry rankings or authoritative comparison pages, those sources become part of the optimization landscape. The opportunity may exist outside the brand's own domain.
This extends familiar off-page SEO into generative discovery. First-party pages still need accurate, accessible and well-structured information, but a company's claims about itself may carry less weight than independent corroboration. Reviews, reputable media coverage, analyst research, public datasets and consistent entity information can help create the evidence environment from which an answer engine constructs recommendations.
Real consumer interfaces and APIs can produce different evidence
SEOPulse says its monitoring approach simulates real-user experiences rather than relying exclusively on model APIs. That distinction is methodologically important. A consumer-facing AI assistant is a product layer, not simply a raw model endpoint.
The interface may invoke web search, apply system instructions, use personalization, connect to specialized indexes or expose features that are not present in the developer API. Testing an API can be valuable for model research, but it does not necessarily reproduce the answer a prospective customer sees when asking the same question inside a commercial AI product.
Real-session monitoring introduces other complications. Answers are variable, interfaces change and platforms can run experiments. But if the business question is whether a buyer sees a brand inside ChatGPT, Gemini or Perplexity, the consumer experience is ultimately the surface that matters.
International brands face multiple AI visibility markets
The inclusion of DeepSeek and Doubao in SEOPulse's platform also highlights a weakness in Western-centric GEO strategies. Global AI discovery does not stop at ChatGPT and Google. Different markets have different dominant products, content ecosystems and language patterns.
Translation alone cannot solve that problem. An English prompt translated literally into another language may not represent how local buyers naturally frame the decision. Local sources may have different authority. Brands themselves may have different recognition and reputation. The information available to an engine can also vary by geography.
International AI visibility therefore requires market-specific prompt panels and source analysis. A global average can hide the fact that a company is highly visible in one region and almost absent in another.
A controlled prompt set is the foundation of credible measurement
AI share of voice can look precise while measuring an arbitrary universe. If a vendor's test set overrepresents prompts that closely match one company's strongest product, that company can appear dominant. If the panel focuses on categories the company barely serves, the opposite happens.
Enterprise teams should define a prompt taxonomy tied to real customer journeys. That means identifying the questions prospects ask while discovering a problem, evaluating approaches, comparing products, checking constraints and preparing to buy. Markets and languages should be specified, and a stable core of prompts should remain unchanged long enough to observe trends.
At the same time, the panel cannot be completely frozen. Customer vocabulary evolves, new products emerge and AI interfaces change how people formulate questions. The best monitoring design therefore combines a stable benchmark with a smaller exploratory layer that captures emerging demand.
Repeated runs matter because generative answers are variable
A single AI answer is weak evidence of a durable competitive position. Generative systems can change wording, ordering and citations across repeated sessions. Retrieval results change, models are updated and sampling itself can introduce variation.
If a brand appears first once and disappears during the next several runs, describing it as “number one in ChatGPT” would be misleading. Monitoring should distinguish persistent patterns from stochastic outcomes through repeated observations and clearly defined aggregation rules.
This is also why screenshots make poor long-term AI visibility metrics. They can document an interesting example, but they do not reveal frequency, stability or competitive share. GEO measurement increasingly needs panel data rather than anecdotes.
Different engines may require different optimization priorities
Low cross-engine agreement also challenges the idea that there is one universal set of GEO ranking factors. If different platforms consistently surface different vendors, at least part of the explanation may lie in engine-specific retrieval, source access and synthesis behavior.
That does not mean brands should create five completely different websites. Core principles remain portable: make factual information accessible, structure entities clearly, maintain current product data, publish evidence that answers real customer questions and build credible external corroboration. But diagnosis should happen engine by engine.
If a company performs strongly in ChatGPT and poorly in Perplexity, the useful question is what differs between those information environments. Which sources does the weaker engine cite? Does it retrieve the company's pages? Are competitors better represented in the publications it favors? Does the weakness occur mainly for one stage of the buying journey?
That approach is more actionable than searching for a generic “AI ranking factor” that supposedly applies identically everywhere.
AI Search is becoming a portfolio of visibility markets
The most durable lesson from SEOPulse's benchmark is not the exact 1.5% figure. That number needs more methodological disclosure before it can be generalized. The larger lesson is that conversational search increasingly behaves like a portfolio of distinct discovery markets rather than one replacement for Google.
Brands will need to decide which engines matter to their customers, build representative prompt panels, separate mentions from citations and recommendations, and study the sources associated with persistent winners. Performance should be measured across the customer journey and across time rather than reduced to a single screenshot or universal AI rank.
Traditional SEO taught marketers to ask where a page ranks. GEO is forcing a more complicated question: across the answer engines that influence our buyers, how often are we selected, in which contexts, and on the basis of which evidence? If major AI systems rarely agree on the same top brand, there may be no single number-one position to win—only a changing share of the answers.