Start with the decision you want to measure
If your goal is to learn whether an assistant recommends your product, counting its name in the answer is not enough. A response can mention you as a poor fit, cite your pricing page to explain a limitation, or recommend a free trial while preferring a competitor for the actual job.
Answer engine optimization (AEO) and generative engine optimization (GEO) overlap in current usage. You do not need two sets of measurements just because the work has two names. You do need to distinguish the answer surface, the buying question, and the outcome you care about.
Keep these four outcomes separate
Use explicit definitions before reading the results. The same response can satisfy several of these outcomes. They are not successive steps that every buyer or answer must follow.
| Outcome | Count it when… | It does not establish… |
|---|---|---|
| Brand mention | The answer explicitly names the product. | Positive sentiment or a recommendation. |
| Source citation | The answer includes an identifiable source reference. Record the URL and claim it accompanies. | That the claim is accurate or the cited product is preferred. |
| Trial recommendation | The answer explicitly suggests trying the product. | That it is chosen for ongoing use or paid purchase. |
| Product choice | The answer chooses it for the stated job. Preserve conditions and qualifications. | A human purchase, conversion, or market-wide preference. |
Duolingo: found in research, not always chosen
Agentzia’s September 3, 2026 Duolingo study used four representative AI research profiles across four buying situations, producing 16 responses. Eight situations withheld the brand for category or problem-led discovery. Eight introduced Duolingo for evaluation or competitive choice.
All eight discovery scenarios found Duolingo. Among the eight decisions that named the product, seven recommended a trial and three selected it beyond a trial. These counts use different eligible groups; they should not be drawn as a funnel from eight discoveries to three purchases.
The ChatGPT known-product response illustrates why: it supported beginner practice and habit-building, while questioning whether the exercises were enough for spontaneous conversation. In the participant’s words: “These are useful, but they are easier than producing language spontaneously.”
Read the entire response and its available sources before turning that concern into a task. The report proposed organizing learning-outcome evidence by language and skill. That proposal has not been implemented and retested in this study, and no resulting sales improvement was measured.
Audit what a citation actually supports
Keep the cited URL, the sentence or claim it accompanies, and whether the source belongs to the product, a competitor, or an independent publisher. Open the source and check the relevant passage. A citation attached to a price does not also verify a learning-outcome claim elsewhere in the answer.
Count responses with a relevant citation separately from the total number of links. One answer with five links is still one observation. If you report citation coverage, define the eligible prompt set and show the numerator and denominator. Do not infer missing URLs or treat a citation count as a recommendation score.
An answer may recommend a brand without linking to it. It may also link to a brand while recommending against it. Preserve both the evidence and the decision instead of using either as a substitute for the other.
Make the comparison interpretable
Choose your questions in advance from real buying situations. Keep brand-free discovery separate from named-product evaluation. A prompt containing your product’s name cannot establish whether the assistant would have found it unprompted.
Record the exact prompt, date, interface or API workflow, available model identifier, browsing configuration, original response, and cited URLs. An API-based research profile is not an exact replay of a consumer app. If a configuration detail is unavailable, label it unknown.
Use fresh conversations and repeat observations with compatible configuration. Report expected and completed coverage. A failed request is missing evidence, not a rejection. Cached responses reused during recovery are not independent repeats.
Break results down by question and research configuration before aggregating them. Explain any change in prompts, models, tools, or coverage. A headline percentage without those details can hide a different experiment.
Measure commercial outcomes separately
Research outcomes help you decide what to investigate or revise. Use web and product analytics to measure visits, completed studies, leads, and paying customers. A trial recommendation recorded in an AI response is not a visitor signing up on your site.
In an early Agentzia self-study, the later report recorded higher understanding and trust counts, while selection stayed at one of six scenarios. Several changes happened together and there was only one run before and one after, so the comparison cannot isolate their effect.
For your own revision, save the page version and repeat the planned questions. Check whether the updated evidence was encountered, whether the original concern remains, and whether the decision changed. Report unchanged outcomes alongside improvements.
Choose tools for the measurement you need
Use monitoring when you need recurring observations across a stable prompt set. Use recommendation research when you need to inspect why a product was chosen or rejected and what page change might address the concern.
Agentzia provides on-demand research through representative ChatGPT, Claude, Gemini, and Grok workflows. It supports the recommendation-testing part of AEO and GEO. It does not measure every answer engine, Google AI Overview placement, or voice-assistant response, and it does not replace a continuous monitoring program.