AI SEARCH TESTING

How to Run a Cross-Platform AI Source Test Without Misreading the Results

A defensible cross-platform AI source test uses a defined buyer-question set, documents the conditions, repeats important prompts and separates citations, mentions, recommendations, URLs and domains. It reports observations without inventing a universal ranking system.

One controlled buyer question branching into distinct AI source results for comparison.

Cross-platform AI visibility reports can look precise while mixing different events. A domain appears in one citation panel. A brand is mentioned in another answer. A competitor is recommended in a third. The results are then compressed into one score.

That shortcut creates false confidence. Current platform documentation describes different search and citation behaviors, while 2026 studies show that results can vary across model versions, interfaces, prompts and runs.

The useful response is not to stop testing. It is to design the test so leadership can distinguish an observable pattern from a one-off result.

This methodology compares AI source behavior without claiming access to hidden ranking formulas or complete retrieval traces.

If you need the strategic premise before the testing method, start with AI Search Is Not One Channel.

Define the decision before selecting platforms

A test should begin with the business decision it needs to support. A company investigating inaccurate brand descriptions needs a different design from one studying competitor recommendations, category sources or local visibility.

Choose platforms because buyers plausibly use them, not simply because a dashboard supports them. Google AI Overviews, Google AI Mode, ChatGPT, Claude, Gemini and Perplexity can all matter, but their interfaces and source behavior should remain separate in the results.

Document whether the test uses a consumer interface or an API. An API can support repeatable research, but it should not be presented as a perfect reproduction of the consumer product.

Record why the products are not interchangeable

Google says AI Overviews and AI Mode may use query fan-out and may use different models and techniques. ChatGPT can decide when to search and may show citations or a Sources panel. Claude’s search tooling can run multiple searches. Gemini grounding can generate searches and return source mappings. Perplexity documents real-time web-grounded responses with inline citations.

These descriptions do not reveal complete ranking formulas. They establish that the products can use different retrieval, interpretation and citation processes. Model versions, product modes, location, personalization and source availability may affect the answer.

The test log should preserve those differences rather than hiding them inside one composite score.

Choose the measurement unit before collecting results

A report cannot calculate meaningful overlap until it defines what is being compared. Domain overlap asks whether the same publishers appear. URL overlap asks whether the exact pages match. Brand overlap, citation overlap and recommendation overlap answer different questions.

A 2026 workshop study ran 1,000 identical ranking-oriented queries across Google Search, web-enabled GPT-4o, Claude 4.5 Sonnet, Gemini 2.5 Flash and Perplexity Sonar Pro. An Aug. 7 preprint audited 2,208 grounded responses across local-discovery questions. Both found material variation, but each tested specified prompts, models and conditions.

Those limitations show why every percentage needs a denominator, prompt set, date, model configuration and measurement level.

Build the question set around buyer decisions

A query for the best provider is not equivalent to one for the most experienced provider, the best fit for a small organization or the safest option in a regulated category. Each wording introduces different decision criteria.

Create groups for branded research, problem discovery, category fit, comparison, trust and action. Keep the core meaning stable across platforms, then use a limited set of meaning-preserving variations to test sensitivity to wording.

Record the intended buyer stage beside every question. That prevents an informational citation from being treated as a recommendation outcome.

Control the conditions that can change an answer

Record the date, platform, model or mode when visible, account state, location, conversation history, personalization and whether web search was active. A fresh session and an established conversation are not identical conditions.

Perfect control is rarely possible in consumer products. The purpose of a test log is to make material differences visible so the analyst does not attribute them to the company by mistake.

Repeat high-value questions across several runs or dates. A single result can illustrate an experience, but it is weak evidence of a stable pattern.

Log citations, mentions and recommendations separately

A citations panel gives the reader a path to verify parts of an answer. It does not reveal every item retrieved, every source considered or the model’s training knowledge.

A company can be cited without being recommended, mentioned without a link or absent even when one of its pages supports the category explanation. For each visible citation, record the exact URL, domain, source type and the claim it appears to support.

This creates an auditable dataset instead of a collection of screenshots.

  1. Capture the answer and visible source links.
  2. Record brands, descriptions, recommendations and material inaccuracies.
  3. Classify first-party, review, directory, publication and other sources.
  4. Repeat the question under the documented conditions.
  5. Compare patterns only within the same defined measurement unit.

Interpret a source by the job it performed

A source-frequency list can encourage crude advice such as publishing on the domain an engine appears to prefer. That ignores why the source was useful for the question.

A company website is usually the best place for current services, policies and accountable facts. Reviews can describe customer experience. Associations can confirm credentials. Independent publications can add analysis. Community discussions may reveal buyer language but can also contain error and bias.

Ask what information the source supplied, whether it was accurate and whether a stronger owned page or external reference is needed.

Report what the test can and cannot prove

A useful report describes the tested conditions, shows the observations and separates findings by platform and question type. It identifies repeated inaccuracies, source gaps and meaningful patterns without inventing a universal AI rank.

The test can show which brands, descriptions and visible sources appeared. It can compare repeatability. It usually cannot prove every source retrieved, the internal weight given to a page or the cause of a final recommendation.

End with decisions: what needs correction, which high-value page needs stronger evidence, which external fact is inconsistent and what should be monitored before further investment.

Test with discipline

Turn AI source observations into a decision-ready baseline.

The Buyer Discovery Audit uses controlled questions, platform-separated evidence and buyer-research analysis to identify which gaps deserve attention.

Explore the Buyer Discovery Audit

Sources

About the author

Giselle Banlat

Founder & Principal Consultant, OutsourceSy

Giselle Banlat is the founder and principal consultant of OutsourceSy, where she helps organizations improve how customers find, research and choose them across search, AI-driven discovery and the wider digital customer journey.

Meet Giselle

A practical starting point

Make the discovery system easier to understand and improve.

Start with a focused review of the buyer-research problem and the most useful next step.