Benchmark
Research Note: Why We Pre-register the Benchmark
The release gates that separate a planned sample from a defensible baseline.
Answer-first: what pre-registration buys readers
Pre-registration fixes the main decisions before observation: markets, journey families, languages, prompt rules, coding definitions, exclusion criteria and release gates. Readers can later judge whether published work matches the plan—or whether the plan quietly moved after someone saw preliminary answers.
For enterprise AI visibility, credibility is fragile. A vendor deck can always show a favourable screenshot. A research programme should show its sampling frame instead.
Preventing outcome-led design
Outcome-led design appears when prompts, markets or scoring rules shift toward answers that flatter a sponsor, a headline narrative or an internal stakeholder. Even well-intentioned teams drift: drop “noisy” prompts, add paraphrases that happen to retrieve better sources, redefine journey families to exclude awkward categories.
Pre-registration does not forbid learning—it forces learning to appear as a versioned protocol change with prospective application, not as retroactive polish.
What is frozen before the first collection window
Release gates (summary)
- Prompt registration frozen and versioned publicly before collection opens
- Independent double-coding on a calibration subset with adjudication log
- Documented exclusion of personalised, failed or ambiguous sessions
- Missingness report and interface-change annotation
- Limitations note and correction route published with any summary table
Full detail lives on the planned baseline sample page and in the methodology.
What happens when quality is insufficient
If calibration fails, we revise the codebook and re-calibrate—not publish interim league tables. If a market stratum lacks usable observations after exclusions, we report shortfall rather than extrapolate. If an interface changes materially mid-window, we split or exclude affected runs with explanation. The outcome may be a methods update—not a partial baseline marketed as complete.
What pre-registration does not promise
It does not make generative systems deterministic. It does not freeze vendor product roadmaps or search interfaces. It does not turn observation into causal proof of commercial impact. It simply ensures that when we say “we tested X,” X was defined before we knew the answers.
Review the benchmark overview for programme intent and the corrections log for future protocol version history.