The question: would clearer buying information change the answer?
On August 28, 2026, we used an early version of Agentzia to research our own website. The first report described a product that was often understood but difficult to verify: pricing, proof, setup detail, and operator information were incomplete.
We then added explicit free and paid usage terms, an Agentzia-on-Agentzia research write-up, and canonical MCP setup documentation. We also narrowed unsupported market claims. A second run evaluated the changed site on the same day.
This is our own historical product research, not an independent customer study. Several things changed together. The record contains one run before and one after, with six scenarios in each. It does not establish that a particular page change caused an outcome.
What stayed comparable—and what the record cannot establish
Both runs used the blind-core-v3 scenario suite, participant model identifiers openai/gpt-5.4 and anthropic/claude-sonnet-5, and observer identifier openai/gpt-5.4. These are historical identifiers from the notes, not today’s model availability or the current Agentzia configuration.
The six scenarios covered direct evaluation, category discovery, competitive comparison, trust research, purchase decision, and implementation research. Both runs completed all six. These are different research tasks, not six interchangeable buyers or six unprompted discovery queries.
The source notes record citation-level evidence, but no exact browser replay. Search results and model responses could vary between runs. The downloadable notes reproduce the recorded counts and provenance; they are not the complete raw transcripts.
What the two reports recorded
Selection remained one of six. The later report recorded more scenarios as understood, trusted, and considered. We show counts because these six scenarios are too small and heterogeneous to represent a market-wide recommendation rate.
| Recorded signal | Before | After |
|---|---|---|
| Discovered | 5 of 6 | 5 of 6 |
| Understood | 4 of 6 | 5 of 6 |
| Trusted | 0 of 6 | 2 of 6 |
| Considered | 3 of 6 | 4 of 6 |
| Selected | 1 of 6 | 1 of 6 |
Why the unchanged result matters
The historical summary score rose from 42 to 53. Had we used only that score, the edit could have looked like a clear success. The selection count tells a more limited story: the second report found improvements in intermediate signals without more final selections.
The follow-up notes say four participants found the pricing page, while direct evaluation still reported missing pricing on the homepage. Trust and purchase research continued to raise missing operator and legal information. That pointed to specific unresolved questions rather than a need for more confident marketing copy.
Even “discovered: 5 of 6” needs context. It aggregates several kinds of scenarios, some directed at Agentzia. The unbranded category-discovery scenario still missed the product. It would be misleading to describe this as appearing in five of six category searches.
What this does not prove
There was no randomized control, no repeated matched baseline, and no isolated single-page intervention. The records do not establish a causal lift in trust, a reduction in uncertainty across all buyers, or any change in search rankings, website traffic, or sales.
These notes describe the August 28 website and the scoring definitions in that early suite. They should not be read as a current audit of Agentzia’s policies or compared numerically with the later four-profile sample reports.
The defensible observation is narrow: these two completed reports recorded unchanged selection despite higher intermediate counts. That is worth reporting even without a positive headline.
How we would design a stronger follow-up
Choose one recurring objection, state the expected effect in advance, and change the page that should answer it. Save the published version and the exact questions. Repeat the baseline and post-edit observations with compatible configuration, keeping a separate unchanged comparison where feasible.
Track whether the updated evidence was encountered, whether the objection remains, and whether selection changes. Report failed attempts and every planned question. Reusing stored responses is useful for recovery but does not create independent observations.
That stronger experiment has not been completed for this historical comparison. The worksheet below provides a place to plan one without promising an outcome.