Ronn Torossian is Founder & Chairman of 5W AI Communications. 5W is one of the largest private communications firms in the world.
If you want to see how fragile “AI visibility” really is as a metric, look at Vanguard. ChatGPT named it in 97.8% of the answers it returned in 5WPR’s new Robo-Advisors & Retail Investing AI Visibility Index 2026. Claude, asked the identical set of questions in the identical week, named Vanguard in just 30.4% of its own answers. Same brand. Same prompts. A 67-point gap that has nothing to do with Vanguard and everything to do with which engine happened to be open on the screen.
One Engine Is Not “AI.” It’s One Vendor’s Opinion
The study ran 24 prompts three times each across Perplexity, ChatGPT, Gemini and Claude, 288 attempted answers in total. Charles Schwab appeared in 91.3% of ChatGPT’s returned answers and 21.7% of Gemini’s. Fidelity showed up in 95.8% of Perplexity’s answers and 39.1% of Gemini’s. A brand team that checks one platform and reports a single “AI visibility score” to its board is reporting one vendor’s retrieval behavior dressed up as a category verdict.
I see this mistake constantly. A junior staffer runs three prompts through ChatGPT, screenshots a favorable answer, and the deck goes out calling it proof of market leadership. This data shows exactly how wrong that habit can be, and how differently the story reads twelve hours later on a different engine.
Why The Engines Disagree
Each of these four systems pairs a different underlying model with a different retrieval tool. Perplexity’s native web search behaves differently than Google’s search grounding inside Gemini, which behaves differently than OpenAI’s web search tool inside ChatGPT. A brand’s appearance rate on any single engine reflects that engine’s specific crawling and ranking choices as much as it reflects the brand’s actual market position. That is not a defect in the study. It is the finding.
Acorns is a clean illustration. Perplexity named it in 66.7% of returned answers. Claude named it in 6.5%. Nothing about Acorns changed between those two numbers, only the system doing the asking.
What A Real Measurement Looks Like
The fix isn’t complicated, but it is more work than most teams currently do. Measure all four engines, not the one someone happens to have a subscription to. Run the same prompt more than once, since the study’s own repeat trials show individual answers can vary even from the same engine on the same day. Report a range, not a single flattering number, and expect competitors’ names to show up in your own results, because they will.
Cross-engine consistency, a brand appearing at a comparable rate across all four systems rather than dominating one and vanishing from another, is the closest thing this discipline has to a durable asset. A brand propped up by a single engine’s quirk can lose that position the next time that engine’s retrieval tool gets updated, with no warning and no press release.
The Boardroom Version Of This Story
Every category I have measured shows some version of the same spread this study found in robo-advisors. Marketers are used to channel fragmentation, television, search, social, each with its own audience and its own measurement. AI answer engines are now a channel with the same property, except the fragmentation is invisible unless you go looking for it, because the chatbot never tells you it’s giving you a different answer than the one next to it would.
The firms that internalize this will stop reporting single-engine screenshots as strategy and start asking the harder, more useful question: does my brand hold up across all four systems, or am I one platform update away from finding out it never did?