BEN / 006

Pre-registration note

Benchmark Baseline & Planned Sample

The sampling frame, coding rules and release gates proposed for the first benchmark cycle. No measured results are presented.

This is a pre-registration, not a result

As of 3 August 2026, AI Search Asia has not released a measured Asia AI Search Benchmark baseline. The tables and labels below define the intended sample and expected data fields so that future changes cannot be quietly made after seeing favourable outcomes.

Pre-registration is a discipline borrowed from empirical research: fix the main design decisions before observation, then let readers judge whether the published work matches the plan. For enterprise AI visibility, that matters because it is easy to assemble persuasive anecdotes—curated prompts, favourable screenshots, selective markets—without documenting what was excluded or why.

No measured results on this page. Everything below describes what we intend to collect and how we intend to code it. Figures describe design targets, not completed observations.

Planned observation unit

One observation unit consists of a prompt identifier, market context, language, answer-system product surface, collection timestamp, clean-session protocol, answer text, visible citations and collection notes. Account state, model label and retrieval mode will be recorded when the interface exposes them.

A single observation is sufficient to describe what one interface returned at one moment. Comparative statements require a declared sample, coding rules and uncertainty treatment. We treat failed or blocked runs as data—logged, not silently replaced—because absence of an answer is often as informative as presence of a citation.

Proposed sample matrix

MarketsHong Kong, Singapore, Japan, Taiwan
Journey familiesCategory discovery, supplier comparison, risk validation, implementation research
LanguagesEnglish plus market-relevant Traditional Chinese or Japanese pathways
Prompts per stratumRegistered primary prompts with permitted paraphrase variants; exact counts frozen at registration
RepetitionsMultiple runs per prompt within a window to document non-determinism; not averaged into a single “true” answer
CadenceOne frozen collection window per quarter, subject to feasibility review

Journey family definitions

Each journey family maps to a distinct enterprise uncertainty. Prompts within a family share a decision frame but vary surface wording within registered limits.

A

Category discovery

Buyer seeks viable options in a category with constraints (sector, scale, geography). Tests which sources establish the shortlist and whether local filters are respected.

B

Supplier comparison

Buyer compares named or implied vendors on capabilities, risk and fit. Tests entity resolution, claim support and role clarity (manufacturer vs distributor vs group company).

C

Risk validation

Buyer investigates compliance, security, financial stability or incident history. Tests whether answers cite primary evidence or rely on secondary summaries.

D

Implementation research

Buyer evaluates deployment, integration, support and total cost. Tests whether answers distinguish marketing claims from operational documentation.

Planned data fields

Each observation record will capture enough context for independent review without assuming access to proprietary model internals.

IdentityPrompt ID, journey family, market, language, product surface label, collector ID, timestamp (UTC)
SessionClean-session protocol version, account state, disclosed model or mode label, geo or locale setting if exposed
AnswerVerbatim response text or permitted derivative where platform terms restrict redistribution
CitationsSource URL, registrable domain, citation position, source class, accessibility status at collection time
CodingOrganisation mention, answer role, claim-support status, local-market applicability, language match, reviewer confidence
QualityCollection notes, failure reason, exclusion flag, adjudication reference if reviewers disagreed

Coding definitions (draft)

  • Source class: first-party corporate, editorial media, institutional or statutory, academic, community or forum, aggregator or comparison site, other (defined in codebook)
  • Answer role: recommendation, example, cited source, comparison candidate, caution, background mention, none
  • Claim support: fully supported, partially supported, unsupported, not applicable (no material claim made)
  • Local applicability: clearly applicable, ambiguous, clearly inapplicable, not tested by prompt
  • Reviewer confidence: high, medium, low—with mandatory note when low

What will not be inferred

The study will not treat citations as sales impact, conflate one product surface with an underlying model, or claim population-wide market share. Scores, if introduced, will include a formula, missing-data treatment and sensitivity note before publication. We will not retroactively redefine journey families to improve a vendor’s appearance, and we will not publish market-level conclusions from a single interface or a single collection day.

Release decision tree

If calibration coding fails to reach acceptable agreement, we will revise the codebook and re-calibrate—not publish interim rankings. If an interface changes materially mid-window, we will split the window or exclude affected runs with explanation. If collection quality is insufficient across markets, the outcome will be a documented methods update and a revised registration—not a partial baseline marketed as complete.

Related reading: benchmark overview, research methodology and the research note on why we pre-register.