AI Search Works. How to Prove It with Real Tests.

Adding FAQ sections to a set of test pages raised AI quotes. Removing them lowered the citations.
That conversion is the difference between correlation and causation, and almost no team measuring AI search today can produce it.
That level of evidence reinforced a recent SEJ webinar with SeoClarity’s Mark Traphagen, VP of Marketing and Product Training, Mihir Naik, Senior Product Manager, AI, and Suraj Lalchandani, IT Project Manager. Their main argument: “Visibility scores tell you what you did. Page-level performance and cross-checking will tell you if what you did was really important.”
The session went through seoClarity’s segmentation testing of business customers working across ChatGPT, Claude, Perplexity, Gemini, and Google’s AI areas: how to build a gold set that expands, how to build a control group where LLMs won’t allow A/B testing, and where new Google AI data enters Search for the first console.
They also share the results of three real client tests, including one change that passed quotes and two outcomes that no one in the room could have predicted.
Watch the full webinar on demand for the full testing method.
Can You Finally See AI Search Visibility in Google Search Console?
For a subset of sites, yes. On June 3, Google launched dedicated Search Console reports for AI Summary and AI Mode, which show page by page how often each URL appears within Google’s AI search features.
Lalchandani called it the biggest AI search experiment ever. “This has been the hardest thing to measure in AI search. Everyone was sampling. Everyone was thinking. But now Google is giving it to you.”
First-party data from the source carries a different level of trust than any third-party tool. But the team was specific about the limitations: the new reports cover only part of what an AI search testing system needs, and ChatGPT, Claude, and Perplexity still require third-party tracking.
In time, the team maps out exactly which posts block new reports, leave them open, and reference each platform for what the AI engine can crawl and provide.
Action item: Check the search console for new AI reports, and see where first-person data fits your test plan before you build around it.
What Alerts Should You Check Out First in an AI Search?
The ones where you almost won. The team created a set of golden alerts that included an AI search funnel, retention awareness, with all information tagged by stage, and then sorted each information according to where the brand currently stands in the AI response.
Step 1 information is an easy win. As Lalchandani puts it, “You’re important, but AI hasn’t been given a URL to link to.”
Tier 2 is the heaviest lift, and one notification bucket is dropped from testing altogether, a move that surprised many in attendance.
Succession is deliberate: early winners buy political capital for tougher trials later. The session covers how to create and tag a gold content set, how tiers are defined, and a tracking unit that pairs each piece of content with the specific page you want it to be cited.
How Do You Do a Split Exam in LLM?
You can’t split live traffic 50-50, so you build a control group instead: a set of related pages that act as your noise filter against model updates and algorithmic shifts.
“Without a control group, all results can be guesswork,” said Lalchandani. “With one, you can tell a real win in the background noise.”
Time is a discipline that many teams skip. The approach is to set an initial time before any change goes live and a small testing window after that, because AI search doesn’t respond instantly the way traditional SEO sometimes does. Cut the window short and, in Lalchandani’s words, “you might read the noise.”
All tests arrive at one of three results, and each one tells you something about your hypothesis. The full session goes over how to create a relative control group, the exact baseline and test windows, and how to read all three results.
Watch the full webinar if needed for a complete test setup.
FAQ Tests That Proved Reason, and Two Tests That Didn’t
seoClarity applied the same method to three clients and got three very different results, which is exactly the point.
The FAQ test was a clear win. With approximately 1,000 information under measurement, adding FAQ sections to the test pages pushed quotes up compared to the control, and they remained high as long as the change was live. The team then returned the change. “The citations went back down. That’s the second piece of evidence. It’s not that the citations just went up when we added the FAQs, but that they went back down when we removed them. That’s causation, not correlation.”
Two other experiments, one on meta descriptions and listicle formatting, ended very differently, and the reasons why hold lessons for anyone who will invest in any strategy. See how both tests played out at full time.
Naik’s framework: every result is a win-win, because you have evidence instead of guesswork. That’s more than most AI search teams have today.
The session also covers schema and markdown test plans, two of the most controversial questions in AEO right now, and a set of quick layout tests for high-value templates that you can use in a few weeks.
Q&A: Most Helpful Questions from the Webinar
Q: How do you measure AI authority if there is no pure authority metric?
“The authority of AI is really how much the model trusts you as a source of the subject. I don’t think there’s a clean number for it or a single number for it, but there are a few signals that you can put together to get a kind of picture that works.”
Lalchandani named four signals that can be stacked, starting with sharing a quote from your top commands and cross-engine consistency, because “consistency across engines means you become an authoritative source in your category for certain types of questions.” He goes through all four, and how to follow them, in a full session.
Question: Can AI bots read FAQ answers hidden behind collapsible toggles?
“Folding can mean many different things. It’s how you make it fold.”
It completely depends on the application: one standard setting keeps the folded FAQs fully readable by AI and Google search engines, and the other makes the content invisible to both, because “even Google will not click on your site.” Lalchandani explains what the recording is, with his stand-up advice attached: “If you’re not sure about something, just check it out. It takes effort, but it will give you a definite answer.”
Q: What is the ROI of an AI quote that doesn’t drive referral traffic?
“You want to be quoted because you control the response that will actually come out.”
Even without a click, Naik explained, your citation page shapes the narrative within the answer, especially compared to questions where citations do the hard work of placing both types. The question goes from traffic to representation: are your USPs well highlighted, the comparison set right, are there any obvious inaccuracies. Lalchandani added an example of an alert from a real restaurant client that shows exactly what happens when AI can’t access your content, which is explained in full in the recording.
Q: Is traditional SEO still a factor in moving the needle for AI discovery?
“Sure. It’s basic. It’s basic.”
Traphagan noted that seoClarity’s longest-standing clients, those with well-optimized content and technically sound sites, also do well in AI search, with AI optimization as an extra layer on top. Lalchandani added: “When we do tests with our clients, we rarely, if ever, find a situation where something works for SEO and doesn’t work for AI search.”
Watch the Full Webinar
The required recording contains everything that the recap reserves: a set of gold sets, category definitions, the creation of a control group with an accurate base and test windows, a reference to the search platform by platform, a meta description and list results, and a schema and markdown test plan. Register to watch the full session if needed.



