Test a direct product question and an unprompted discovery question. Then inspect what ChatGPT searched, which sources and competitors shaped the answer, and why your product was selected or passed over.
Five practical steps One repeatable test
The short answer
Ask the way a customer asks. Record much more than the final sentence.
A useful recommendation test starts with a realistic buying situation. ChatGPT should be free to research the market, decide which products belong in the option set, and reach its own conclusion.
The result becomes actionable when you can connect the decision to the searches, sources, product description, trust concerns, alternatives, and pages behind it.
Choose a question a customer would naturally ask before buying. Include their goal, constraints, and context. Keep your preferred answer out of the prompt so the assistant can form its own option set.
Customer context02
Ask about the product directly
Run a known-product question such as, “I am considering Product X. Is it a good choice for this use case?” This reveals existing understanding, trust, objections, and the alternatives ChatGPT introduces.
Named product03
Run an unprompted discovery question
Describe the category or problem without naming the product. This shows whether ChatGPT discovers the brand naturally and which competitors occupy the same customer need.
Unprompted search04
Preserve the path to the answer
Record the query, cited sources, reported page visits, products considered, product description, trust concerns, and final recommendation. The reasoning is what makes the result useful to a marketing or product team.
Evidence capture05
Change the relevant page and retest
Improve the category language, use-case explanation, comparisons, pricing clarity, evidence, or page linking supported by the result. Repeat the same request after release and compare the decision and its supporting evidence.
Matched retest
A reusable test set
Four questions reveal different parts of the decision.
Replace the bracketed details with the customer, problem, category, and constraints that matter to your product.
Situation
Example request
What it reveals
Known product
I am considering [product] for [use case]. Would you recommend it?
Perception, trust, objections
Category discovery
What are the best [category] products for [customer]?
Unprompted discovery and shortlist
Problem led
I need to [solve problem]. What should I use?
Problem association and category fit
Competitive choice
Which product would you choose for [situation], and why?
Final selection and tradeoffs
What a real result looks like
8/8 discovery paths found itDuolingo public sample
Every agent found Duolingo. They still disagreed on whether it was enough for serious learners.
Its brand was unmistakable. Evidence for advanced speaking and writing outcomes shaped the harder decisions.
The useful finding is the disagreement. Discovery was strong, while evidence about advanced learning outcomes shaped whether the product fit a more demanding customer.
Clarify category and problem language, publish focused use-case pages, improve internal links, and earn relevant third-party citations.
Misunderstood
The product entered with the wrong story
State what the product does in plain language, connect features to customer situations, and make the strongest differentiators easy to extract.
Concerned
Trust weakened the recommendation
Add verifiable proof, transparent pricing, company information, policies, limitations, and the evidence required for the specific claim.
Outcompeted
Another product fit the request better
Explain who the product serves best, where it differs, and why those differences matter for the customer's exact constraints.
Common questions
Build a recommendation test you can repeat.
Can I check this by asking ChatGPT myself?
Yes. Use realistic customer questions and preserve the complete answer, sources, alternatives, and decision. A repeatable test becomes more useful when the same situations are run across assistants and again after a product or website change.
Does one ChatGPT answer prove how every customer will be advised?
One answer is a single observation. Recommendation behavior can vary with the model, available search tools, prompt, region, and time. Use several representative situations and compare repeated tests under the same setup.
What counts as a recommendation?
A recommendation occurs when the assistant selects the product for the customer's stated situation. Mentions, consideration, trial suggestions, and final selection should be recorded separately because they represent different levels of intent.
What should I change if ChatGPT does not recommend my product?
Follow the evidence. Common improvements include clearer category language, specific use cases, transparent pricing, focused comparison pages, credible proof, accessible policies, and stronger links between the pages the assistant reached and the information it missed.
Should I test Claude, Gemini, and Grok too?
Test the assistants your customers use. Running matched situations across ChatGPT, Claude, Gemini, and Grok reveals where discovery and recommendation behavior differs by client.