RESEARCH / 019

Methods

Field Note: What Counts as an AI Answer Observation?

A practical definition of the unit of observation for reproducible enterprise AI visibility research.

Answer-first definition

An AI answer observation is a structured record of what a named generative answer interface returned for a registered prompt at a documented time—not a screenshot filed without context, and not a brand mention scraped from an unknown query.

Without a stable unit of observation, teams compare incomparable anecdotes. Marketing counts “mentions”; product counts “correct specs”; legal counts “compliance statements”—all from different prompts, accounts and dates. This field note defines the minimum record that lets independent readers ask: would the same protocol produce a comparable row?

Why a screenshot is not a dataset

Screenshots lose session state, product-surface labels, paraphrase history and failure modes. They encourage cherry-picking: one favourable answer becomes a deck, one embarrassing answer becomes a crisis. A useful observation binds a registered prompt to market context, language, product surface, timestamp, session conditions, answer text, visible sources and collector notes.

Scope limit. One observation describes one interface response. It does not establish market share, durable ranking, model-wide behaviour or commercial effect.

Required record fields

Prompt identityPrompt ID, journey family, verbatim text, language, market context
InterfaceProduct name, disclosed model or mode label, account state, locale if exposed
SessionClean-session protocol version, collector ID, UTC timestamp
OutputAnswer text or permitted derivative; cited URLs with positions; organisation mentions
QualityFailure reason if any; exclusion flag; free-text collector notes

At minimum, retain enough context to reproduce the attempt. Record failures rather than silently replacing them—refusals, empty answers and blocked sessions are observations about interface behaviour.

Recommended coding extensions

When observations feed comparative work, extend the record with codebook fields described in the research methodology:

  • Source class and registrable domain after redirect resolution
  • Answer role: recommendation, example, cited source, comparison candidate, caution, background
  • Claim-support status relative to cited sources
  • Local-market applicability where the prompt embeds jurisdiction or availability constraints
  • Reviewer confidence with mandatory notes when low

Common anti-patterns

01

Vanity prompts

“Tell me about [Brand]” without a buyer decision frame. Produces mention counts without decision relevance.

02

Moving goalposts

Changing prompts after seeing answers. Destroys comparability and invites outcome-led design.

03

Surface conflation

Treating consumer chat, search-augmented answers and enterprise copilots as one homogeneous “AI.”

04

Silent deduplication

Dropping failed runs or “weird” answers without documenting exclusion rules.

What one observation can support

Descriptive statements about that interface, prompt and moment—for example, which domains appeared, whether a statutory source was cited, or whether an entity name matched the registered organisation. Aggregated patterns require a declared sample, coding rules and uncertainty treatment, as set out in the planned baseline.

Use this field note alongside the full research methodology and the enterprise governance guide when standing up an internal observation log.