Version 1.0 · 3 August 2026
Research Methodology
How we frame prompts, record answer-system observations, code citations and communicate uncertainty.
Research scope
AI Search Asia observes outputs available through named generative answer-system product surfaces. We do not claim to inspect model weights, training corpora, retrieval indexes or stable universal rankings. Every finding is bounded by interface version, account state, geography, language, collection time and the buyer journey framed in the prompt register.
This scope is deliberate. Enterprise readers need reproducible evidence about what buyers actually see—not speculative claims about invisible internals.
Protocol overview
- Frame: define a buyer decision, market, material claim and failure modes before writing prompts.
- Register: freeze prompt families, permitted paraphrase variants, exclusion rules and journey-family assignments.
- Collect: record timestamps, visible product labels, session conditions, answer text, cited URLs and collection failures.
- Code: classify source, answer role, claim support, local fit and reviewer confidence using a versioned codebook.
- Review: double-code a calibration subset; adjudicate material disagreement; update codebook prospectively if needed.
- Release: publish sample description, missingness, limitations, version identifier and correction route.
Prompt design rules
- Prompts mimic documented enterprise research questions—not leading brand queries or hidden marketing copy
- Material constraints (jurisdiction, sector, scale) are stated in the prompt when they matter to the decision
- Paraphrase variants are registered in advance; ad hoc prompt shopping during collection is prohibited
- Follow-up prompts, where used, are scripted to represent realistic multi-step research—not trick chains
Collection standards
Quality controls
Redirects are resolved before domain coding. Duplicate citations are retained at answer level but deduplicated for source-diversity summaries. Inaccessible sources are labelled, not assumed to support a claim. Reviewer uncertainty is a data field rather than silently forced into a binary category. Machine translation may assist reviewer navigation but does not create new evidence fields.
Calibration targets and agreement metrics will be published with the first baseline. Until then, this page describes intended controls—not completed inter-rater statistics.
Reproducibility limits
Generative systems are non-deterministic and interfaces change without notice. A repeated prompt may produce a different answer without an underlying trend in your evidence base. We therefore use frozen collection windows, report observation counts and dispersion, and avoid causal language unless the design supports it.
Readers should treat single screenshots—ours or anyone else’s—as anecdotes. Comparative claims require the registered sample described in the baseline plan.
Dataset publication
Dataset schema will only be emitted for an actual downloadable, described dataset marked up according to our public schema. Planned samples and narrative pages are not labelled as Dataset. Where redistribution of verbatim answers is restricted by platform terms or copyright, we may publish derived citation records, coding tables and a field dictionary instead of raw answer text.
Versioning and corrections
Material protocol changes receive a new version number, apply prospectively by default, and are logged on the corrections page. Editorial typos may be fixed silently. Coding definition changes trigger re-calibration before comparative release.