RESEARCH / 021

Benchmark

Research Note: Why We Pre-register the Benchmark

The release gates that separate a planned sample from a defensible baseline.

Planning document. No measured benchmark results are included. This note explains design discipline—not findings.

Answer-first: what pre-registration buys readers

Pre-registration fixes the main decisions before observation: markets, journey families, languages, prompt rules, coding definitions, exclusion criteria and release gates. Readers can later judge whether published work matches the plan—or whether the plan quietly moved after someone saw preliminary answers.

For enterprise AI visibility, credibility is fragile. A vendor deck can always show a favourable screenshot. A research programme should show its sampling frame instead.

Preventing outcome-led design

Outcome-led design appears when prompts, markets or scoring rules shift toward answers that flatter a sponsor, a headline narrative or an internal stakeholder. Even well-intentioned teams drift: drop “noisy” prompts, add paraphrases that happen to retrieve better sources, redefine journey families to exclude awkward categories.

Pre-registration does not forbid learning—it forces learning to appear as a versioned protocol change with prospective application, not as retroactive polish.

What is frozen before the first collection window

Journey familiesCategory discovery, supplier comparison, risk validation, implementation research—with definitions
Market contextsHong Kong, Singapore, Japan, Taiwan strata and language pathways
Prompt registerPrimary prompts and permitted paraphrase variants; no ad hoc additions mid-window
Variation rulesRepetition count, follow-up script limits, interface list
Collection windowsStart/end timestamps; rules for interface changes mid-window
Coding definitionsSource class, answer role, claim support, applicability, confidence
Quality gatesCalibration agreement thresholds, failure handling, minimum completeness by market

Release gates (summary)

  • Prompt registration frozen and versioned publicly before collection opens
  • Independent double-coding on a calibration subset with adjudication log
  • Documented exclusion of personalised, failed or ambiguous sessions
  • Missingness report and interface-change annotation
  • Limitations note and correction route published with any summary table

Full detail lives on the planned baseline sample page and in the methodology.

What happens when quality is insufficient

If calibration fails, we revise the codebook and re-calibrate—not publish interim league tables. If a market stratum lacks usable observations after exclusions, we report shortfall rather than extrapolate. If an interface changes materially mid-window, we split or exclude affected runs with explanation. The outcome may be a methods update—not a partial baseline marketed as complete.

What pre-registration does not promise

It does not make generative systems deterministic. It does not freeze vendor product roadmaps or search interfaces. It does not turn observation into causal proof of commercial impact. It simply ensures that when we say “we tested X,” X was defined before we knew the answers.

Review the benchmark overview for programme intent and the corrections log for future protocol version history.