Agentzia tests Agentzia / August 28, 2026

Three runs changed the product.

We gave Agentzia its own landing page, changed the deployed site from the traces, and ran the same six-intent protocol again. Two complete panels form the comparison. A third incomplete panel records what the reporting system needed to fix.

Run cost after controls$2.38baseline $11.76
Measured cost reduction79.8%same six-intent suite
Stable consideration5/6five in both complete runs

Complete-panel comparison

SignalVersion 01Version 02Change
Overall score6757-10
Discovered5/65/60
Understood5/65/60
Trusted0/60/60
Considered5/65/60
Selected5/62/6-3
Total measured cost$11.76$2.38-79.8%
Participant input1,329,970575,882-56.7%

The selected count moved from five to two because three low-risk trials were classified as shortlisted. Discovery, understanding, and consideration stayed at five of six. The journey shows the stable behavior that the aggregate score alone hides.

The deployed versions

01 / current-model baseline
released in PR 09

The product was understood when named, while unbranded category discovery missed it. One comparison journey consumed most of the run cost.

What shipped next. Bound search-result payloads, capped fetched-page context, added a category page, exposed selective MCP report reads, and added credit-free trace reuse for observer retries.

02 / after cost and discovery controls
released in PR 10

Discovery, understanding, and consideration stayed at five of six. The selected count moved because three low-risk trials were classified as shortlisted, which made the stable journey more useful than the aggregate score.

What shipped next. Published the complete model and data-provider matrix, added structured product metadata, and placed panel scope, trace fidelity, date, suite, and interpretation directly above every report.

03 / report-integrity check
released in PR 11

An incomplete panel still produced a confident-looking score.

What shipped next. Added a first-class degraded run state, displayed the returned and expected panel size before the score, and instructed users to exclude incomplete panels from version comparisons.

Excluded measurement

Degraded / 4 of 6

Two ChatGPT Free participant calls returned no answer. Four completed traces remained useful, but the 64 score and 75% recommendation rate were excluded from comparisons because the panel was incomplete.

$2.49 measured cost

Latest complete journey

Discovered5/683%
Understood5/683%
Trusted0/60%
Considered5/683%
Selected2/633%

What changed our view

The traces exposed product behavior and pipeline behavior.

Stable Five participants found, understood, and considered Agentzia in both complete runs.

Variable Selection moved as agents interpreted a free trial as either a choice or a shortlist. The report now emphasizes the full decision journey.

Cost Bounded fetched-page context cut participant input by 56.7% while preserving model-directed search and page choice.

Missing Unbranded discovery missed Agentzia in every run. Category indexing remains an open distribution task.

Quality The final run returned four participants. Its traces are retained, its panel is labeled degraded, and its score is excluded from the trend.

Method and reproducibility

Stable requests. Current models. Preserved evidence.

Suite agent-surfaces-v4: direct evaluation, discovery, comparison, trust, purchase, and implementation.

Panel ChatGPT Free, ChatGPT Paid, Gemini, and Claude surfaces using Luna, Sol, Gemini 3.7 Flash, and Opus.

Observer OpenAI GPT-5.6 Sol · medium reasoning classifies the preserved participant records after each decision.

Trace Citation-level records, model usage, cache usage, tool counts, duration, and cost are persisted per participant.

Comparison Only complete six-of-six panels enter the version trend. Dynamic models and indexes make the study auditable rather than bit-for-bit replayable.

Cost scope Run totals include participant and observer model charges. Firecrawl credit pricing is tracked separately by its provider.

Start an observation

Run the same decision panel against your page.

Test your page free