Historical case study · Agentzia testing Agentzia

GEO before and after: clearer pricing, unchanged selection

We added pricing, research evidence, and setup docs to Agentzia. In two early six-scenario runs, selection stayed at 1 of 6. Here is what the comparison can tell us.

Free files · No sign-up required

The question: would clearer buying information change the answer?

On August 28, 2026, we used an early version of Agentzia to research our own website. The first report described a product that was often understood but difficult to verify: pricing, proof, setup detail, and operator information were incomplete.

We then added explicit free and paid usage terms, an Agentzia-on-Agentzia research write-up, and canonical MCP setup documentation. We also narrowed unsupported market claims. A second run evaluated the changed site on the same day.

This is our own historical product research, not an independent customer study. Several things changed together. The record contains one run before and one after, with six scenarios in each. It does not establish that a particular page change caused an outcome.

What stayed comparable—and what the record cannot establish

Both runs used the blind-core-v3 scenario suite, participant model identifiers openai/gpt-5.4 and anthropic/claude-sonnet-5, and observer identifier openai/gpt-5.4. These are historical identifiers from the notes, not today’s model availability or the current Agentzia configuration.

The six scenarios covered direct evaluation, category discovery, competitive comparison, trust research, purchase decision, and implementation research. Both runs completed all six. These are different research tasks, not six interchangeable buyers or six unprompted discovery queries.

The source notes record citation-level evidence, but no exact browser replay. Search results and model responses could vary between runs. The downloadable notes reproduce the recorded counts and provenance; they are not the complete raw transcripts.

What the two reports recorded

Selection remained one of six. The later report recorded more scenarios as understood, trusted, and considered. We show counts because these six scenarios are too small and heterogeneous to represent a market-wide recommendation rate.

Historical report counts · August 28, 2026 · Six completed scenarios in each run
Recorded signalBeforeAfter
Discovered5 of 65 of 6
Understood4 of 65 of 6
Trusted0 of 62 of 6
Considered3 of 64 of 6
Selected1 of 61 of 6

Why the unchanged result matters

The historical summary score rose from 42 to 53. Had we used only that score, the edit could have looked like a clear success. The selection count tells a more limited story: the second report found improvements in intermediate signals without more final selections.

The follow-up notes say four participants found the pricing page, while direct evaluation still reported missing pricing on the homepage. Trust and purchase research continued to raise missing operator and legal information. That pointed to specific unresolved questions rather than a need for more confident marketing copy.

Even “discovered: 5 of 6” needs context. It aggregates several kinds of scenarios, some directed at Agentzia. The unbranded category-discovery scenario still missed the product. It would be misleading to describe this as appearing in five of six category searches.

What this does not prove

There was no randomized control, no repeated matched baseline, and no isolated single-page intervention. The records do not establish a causal lift in trust, a reduction in uncertainty across all buyers, or any change in search rankings, website traffic, or sales.

These notes describe the August 28 website and the scoring definitions in that early suite. They should not be read as a current audit of Agentzia’s policies or compared numerically with the later four-profile sample reports.

The defensible observation is narrow: these two completed reports recorded unchanged selection despite higher intermediate counts. That is worth reporting even without a positive headline.

How we would design a stronger follow-up

Choose one recurring objection, state the expected effect in advance, and change the page that should answer it. Save the published version and the exact questions. Repeat the baseline and post-edit observations with compatible configuration, keeping a separate unchanged comparison where feasible.

Track whether the updated evidence was encountered, whether the objection remains, and whether selection changes. Report failed attempts and every planned question. Reusing stored responses is useful for recovery but does not create independent observations.

That stronger experiment has not been completed for this historical comparison. The worksheet below provides a place to plan one without promising an outcome.

Measure the decision you actually care about.

Read a complete current sample report, then test your own product. Separate discovery, trial recommendations, and product choice.

Explore a complete report ↗