Recognized by name, absent from the shortlist.
Neither substantive shortlist included the client's product. Named questions produced product descriptions that mixed in outdated information. NiubiGEO identified separate discovery and accuracy issues.
01 / FINDINGS
Findings
Absent from both substantive shortlists
Both substantive Q4 answers listed other options without the client's product. Q1 requested comparison criteria, so its absence was not treated as a recommendation failure.
Recognition did not ensure current information
Against the official materials reviewed at the time, NiubiGEO found that some named answers mixed older descriptions with current capabilities. Recognition alone did not ensure accuracy.
Public comparison evidence needed more detail
Reviewed materials included tests and technical documentation, but not a complete public result set covering the performance conditions in scope. This does not imply an absence of internal testing.
02 / QUESTIONS & OUTCOMES
Question themes and diagnosis outcomes
The following are anonymized intent summaries, not the prompts sent to the models. Results refer to saved original tests: Q1 and Q4 were unbranded; Q2 and Q3 named the product to check factual accuracy. No customer-supplied answers were included.
| Question theme / intent | GPT-6 Astra | Claude Fable 5.1 | Gemini 3.8 Flash |
|---|---|---|---|
| Q1 · Unbranded · Comparison criteria Identify the factors for comparing application protection options. | Not mentioned | Declined · unscored | Not mentioned |
| Q2 · Named · Feature understanding Check whether the model distinguishes current capabilities and usage limits accurately. | Recognized when named | Declined · unscored | Recognized when named |
| Q3 · Named · Adoption review Review the conditions a team should verify before using the product. | Recognized when named | Declined · unscored | Recognized when named |
| Q4 · Unbranded · Product shortlist Find candidate tools for performance requirements and evidence for comparing them. | Not mentioned | Declined · unscored | Not mentioned |
A mention does not establish a recommendation or an accurate description. Named questions and unprompted discovery are interpreted separately.
03 / RECOMMENDATIONS
Recommended priorities
Consolidate current-version information
Use a clear product entry point for version, maturity, feature status and usage conditions, with older materials pointing to current documentation.
Publish performance results readers can verify
Document environments, settings, measurement methods, results and limits around real selection needs so buyers can compare suitability.
Evaluate discovery and accuracy separately
After content updates, repeat the original tests to check unbranded discovery and outdated information in named answers separately.
04 / SCOPE
Test scope
Twelve API requests with web search enabled returned eight substantive answers and four provider declines; declines were not scored. Some answers returned no retrieval record, so actual retrieval is not established for those responses. Each question-model pair was sampled once, not a stable recommendation rate. No human interface or software performance testing was conducted; recommendations are not completed improvements or measured gains.
NiubiGEO completed this report using model API tests. Human web and app testing can be arranged separately, with human testing supplied by NiubiStar.