The name appeared. The product was misidentified.
The discovery question did not surface the product. In a named usage-rules question, two answers assigned it to another product category. The diagnosis separated name mentions, recognition and accuracy, and identified conflicting help guidance.
01 / FINDINGS
Findings
The need did not lead to the product
All three discovery answers offered other options without mentioning the target. The report compared those options with the intended need rather than relying only on mention counts.
Naming the product did not prevent misclassification
The audience question produced generic or inaccurate descriptions. In the usage-rules question, two answers used the wrong product category and one explicitly could not identify the product. No official product source was returned in this sample.
Help guidance contained conflicting versions
A separate review found an older help page inconsistent with current guidance on part of the usage behavior. The conflicting page was not returned in this run, so it cannot be established as the cause of the model errors.
02 / QUESTIONS & OUTCOMES
Question themes and diagnosis outcomes
Below are anonymized summaries of question intent, not the original prompts or answers. Q1 was unbranded; Q2 and Q3 named the target product to test accuracy. Full prompts and results were delivered privately to the customer. Counts and classifications are unchanged.
| Question theme / intent | GPT-4.1 mini | Claude Haiku 4.5 | Gemini 2.5 Flash |
|---|---|---|---|
| Q1 · Unbranded discovery Find lightweight tools for a group's everyday collaboration needs. | Not mentioned | Not mentioned | Not mentioned |
| Q2 · Intended audience Assess which users and situations the target product is designed for. | Generic guidance | Inaccurate description | Generic guidance |
| Q3 · Usage rules Understand identity, device and state-management rules in everyday use. | Wrong product identity | Product not identified | Wrong product identity |
A mention does not establish a recommendation or an accurate description. Named questions and unprompted discovery are interpreted separately.
03 / RECOMMENDATIONS
Recommended priorities
Reconcile help guidance first
Reconcile older help pages, current product guidance and interface instructions, with clear steps for each usage situation.
State the product category on key pages
Use consistent product naming, audience and purpose on the homepage and help pages. Explain technical terms within the actual use case.
Connect the need to a getting-started example
Show a short path from first use to outcome and make the explanatory text directly readable. Retest unbranded discovery and named accuracy separately.
04 / SCOPE
Test scope
All nine answers completed, sampled once per question and model using shared OpenRouter Exa search. Generic advice, inaccurate claims, wrong categories and non-identification were recorded separately. No human consumer-interface or functional testing was performed, and no optimization effect has been established.
NiubiGEO completed this report using model API tests. Human web and app testing can be arranged separately, with human testing supplied by NiubiStar.