GEO audit · Worksheet included

How to run a GEO audit: a worked example and checklist

A practical GEO audit for product teams: choose buying questions, inspect AI responses, identify evidence gaps, and write a revision you can retest.

Free files · No sign-up required

1. Define the decision you want to understand

A generative engine optimization (GEO) audit investigates how a product appears in AI-assisted research and what happens when an assistant evaluates it for a buyer. This guide focuses on product recommendations. It does not measure your share of all AI conversations or replace a technical SEO audit.

Choose one product, one audience, and a small set of real buying situations. Write down the canonical URL, market, language, customer constraints, and the question you need to answer. “Why are small teams choosing the other tool?” is a more useful scope than “Improve our AI score.”

Download the worksheet without signing in. It includes a filled Duolingo example, a blank revision brief, and a checklist. The separate CSV is a blank log: add one row for each response, not one row for an entire brand.

2. Check that the relevant pages can be reached

Open the product, pricing, and supporting policy pages while signed out. Check for errors, login gates, contradictory prices, and essential facts that are only visible inside images. Inspect the rendered text and the links a reader would follow to verify the claim.

For Google AI Overviews and AI Mode, Google says the existing SEO requirements apply; it does not prescribe special AI markup. Check indexing eligibility, crawl access, internal links, and visible content. These are access checks, not evidence that an assistant will recommend the product. Other assistants have their own retrieval systems.

3. Write questions before looking at the results

Use a mix of questions that withhold your brand and questions that introduce it. Keep the exact wording, including constraints. The examples below are templates, not observed search-volume data.

For a small initial audit, you could choose four questions and collect three fresh responses per question on each configuration you want to investigate. That is a practical starting protocol, not a statistically representative sample. Decide the coverage in advance and record incomplete attempts.

Category discovery

“What should a two-person design team use to share interactive prototypes with clients?”

Problem-led discovery

“Our clients struggle to review designs over email. What would help us gather feedback?”

Named-product evaluation

“Would [product URL] work for a two-person design team that needs client feedback? Explain any limitations.”

Competitive choice

“Compare [our product] and [alternative] for this team. Which would you choose, and what information is missing?”

4. Preserve the answer before summarizing it

Record the date, exact prompt, assistant or model identifier, interface, browsing mode, fresh-conversation status, response, and cited URLs. For API-based research, record the configured model and tools; do not label it an exact replay of a consumer app. If a setting is unavailable, write “unknown.”

Code discovery only for questions that withheld the brand. For named-product questions, record discovery as not applicable. Separately record whether the answer mentions the product, recommends trying it, chooses it for the stated job, and identifies a deciding concern.

A citation can establish what a response pointed to. It does not by itself prove the source was read accurately or that the product was recommended. Check the cited passage and keep the original response available for review.

5. Turn one finding into an edit

In the September 3 Duolingo study, a ChatGPT participant distinguished beginner habit-building from spontaneous conversation. The report proposed making learning-outcome studies easier to assess by language and skill.

The filled worksheet translates that into a revision brief: place a “Find evidence for your course” section on the studies page; identify language, level, measured skill, population, findings, and limitations; link the original publication. Do not use a reading result as evidence of speaking proficiency.

The acceptance check is whether a learner can find evidence for their course and intended skill. The retest asks whether that evidence is found and whether the concern remains. This proposed edit has not been implemented or retested in the Duolingo example.

6. Prioritize findings you can act on

Choose a concern that affects an actual buying decision and recurs in the responses you collected. Attach the evidence and distinguish an existing fact that needs explaining from a capability or independent proof you do not yet have.

Assign an owner, page location, proposed wording or structure, and acceptance check. Avoid collapsing repeated mentions of the same concern into several unrelated tasks. Keep genuine product limitations visible in the audit.

7. Retest without hiding the denominator

After publishing a revision, check the live page and preserve its version. Repeat the planned question set with compatible configuration. Report completed and expected observations, and compare like-for-like questions. Label changed prompts, models, or routes as a break in comparability.

Look at whether the original concern disappeared, whether the new evidence was cited, and whether the decision changed. Repeated observations help reveal variability; a single before-and-after pair cannot isolate causation.

Keep unchanged outcomes in the report. In an early Agentzia comparison, recorded understanding and trust counts rose while selection stayed at one of six scenarios. That is useful information for deciding what to investigate next.

See the evidence before choosing what to change.

Agentzia organizes buying research into decision records and proposed page improvements. Your first test is free.

Audit your product ↗