# GEO audit worksheet

By Agentzia · https://www.agentzia.dev/resources/geo-audit-checklist
Free to copy and adapt for your own audits. No account required.

## Start here

Choose one product and audience. Save exact questions before testing. Use a fresh conversation for each observation. Record every response, including failures. Use the companion CSV for one row per observation; it contains no sample measurements.

## Scope

- Product and canonical URL:
- Buyer, market, language, and constraints:
- Decision to investigate:
- Pages in scope:
- Exact discovery questions (no brand or target URL):
- Exact named-product questions:
- Planned configurations and repetitions per question:
- Planned total observations:
- Outcome definitions and rules for ambiguous answers:
- Audit owner and dates:

## Access checks

- [ ] Relevant pages load while signed out.
- [ ] Important product facts appear in readable text.
- [ ] Product, pricing, and policies agree.
- [ ] Evidence links resolve to the claimed information.
- [ ] Crawl and indexing eligibility checked separately from recommendations.

## Observation log field guide

Use geo-observation-log.csv in Excel, Sheets, or a text editor. Keep private research files private.

- observation_id: stable label, such as Q1-R1.
- question_type: discovery or named_product.
- exact_prompt: full request, including constraints.
- observed_at: ISO date/time with timezone.
- surface_or_api: consumer interface or API workflow actually used.
- model_identifier: exact available identifier, or unknown.
- browsing_mode: observed/configured setting, or unknown.
- fresh_conversation: yes/no/unknown.
- page_version: saved page snapshot or deployed version, if available.
- completion: complete/failed/incomplete.
- discovered: yes/no/unclear for discovery; not_applicable for named_product.
- mentioned, trial_recommended, selected_for_job: yes/no/unclear; leave unmeasured outcomes blank on failed attempts.
- deciding_concern: short description in your own words.
- cited_urls: URLs actually cited, separated by spaces.
- response_reference: saved original answer or report record.
- notes: missing information and configuration changes.

Code trial recommendations separately from selecting a product for the stated job. Do not count failed attempts as rejections. Report completed/expected coverage. Reused cached responses are not independent repeats.

## Filled example: Duolingo (a proposed edit, not a measured improvement)

- Study date: September 3, 2026.
- Evidence: https://www.agentzia.dev/sample-reports/duolingo#decision-chatgpt-known-product
- Research scope: representative AI profiles; not an exact consumer-app replay.
- Finding: the ChatGPT participant supported beginner practice but questioned whether the core exercises were enough for spontaneous conversation.
- Supporting passage: “These are useful, but they are easier than producing language spontaneously.”
- Source page to revise: https://www.duolingo.com/efficacy/studies
- Proposed location: before the research listings.
- Proposed section: “Find evidence for your course.” Identify course language, level, measured skill, learner population, findings, and limitations. Link each original study.
- Evidence required: actual study details; a reading/listening result cannot establish speaking proficiency.
- Acceptance check: a learner can locate evidence for the intended course and skill and identify untested outcomes.
- Retest question: does the matched response find the relevant evidence, describe its scope correctly, and still raise the original concern?
- Implementation/retest status: proposed only; no post-edit measurement.

## Your revision brief

- Finding and exact supporting response:
- Classification: discovery / misunderstanding / missing evidence / product limitation.
- Why it affects this buying decision:
- Existing source that resolves it, or evidence still needed:
- Page URL and location:
- Proposed change:
- Owner:
- Acceptance check:
- Published version and time:
- Expected effect (record before testing):
- Matched retest questions and configuration:

## Compare results

- Baseline completed / expected:
- Follow-up completed / expected:
- Independent repetitions per question and configuration:
- Prompt/model/tool/page changes that limit comparison:
- Was the updated evidence encountered?
- Did the original concern remain?
- Trial recommendation counts, with eligible denominator:
- Product selection counts, with eligible denominator:
- Discovery counts, using only brand-free questions:
- Unchanged or negative observations:
- Limits and next action:

One before/after pair does not isolate causation. Avoid equating research outcomes with human conversions, search rankings, or market share.
