Docs

Evidence and trust

The fields that say how far a Search result or an eval grade can be trusted, what each one means when it is null, and a policy for acting on them in code.

Every Search result and every eval grade carries fields that say how much it can be trusted. They are the part of a response your code should branch on before it uses the rest. This page collects them in one place; the Search and Evals API references carry the full detail.

Three rules hold across all of them:

  • Null is not zero. A field that could not be measured comes back null or with an unavailable status, never as 0. A zero is a measurement.
  • A refusal is not a fail. When a grade could not be made honestly, the run says not_applicable and gives a reason. Nothing was assessed, so there is no score to read into.
  • Quotes are retrieved, never written. Every customer quote in a response is a real utterance from your corpus. The model cites ids; the server swaps them back to the stored text and drops any id it does not recognise.

On a Search result

FieldWhat it tells youAct on it
mode_ranWhich of the four lanes answered. Lanes read different stores at different grains.Compare two results only when they ran on the same lane.
corpus.store, corpus.grainThe physical store (warehouse, pgvector, theme_index) and what one row counts.A theme is a cluster, not an utterance. Do not quote it as one person's words.
corpus.empty_reasonWhy an empty result is empty.read_failed is an error: retry. no_match means nothing was said. weak_matches_only means rephrase.
corpus.speaker_scopeWhose speech the store can hold: all, or external_only (customer side only).Do not report an internal-vs-external split off an external_only store.
corpus.coverage.basisWhether the time-range numbers describe the whole match set, a truncated slice, or nothing (unavailable).Read this before any date or recency claim.
corpus.truncatedA cap cut the result.Do not treat the rows as the population.
compiled.sqlThe query as you asked for it, before the gate added your tenant filter.Keep it as a record of intent.
detail.internal.statusOn the plain-language lane, whether the SQL ran (ok) or not (failed).failed during a burst can be the warehouse-read budget. Back off before rewording.

On an eval grade

FieldWhat it tells youAct on it
headline.submitted.checks (passed / total)How many rubric lines your own copy cleared. The only number on the report about your writing.Report this fraction.
gate.passedThe send/hold bit for the copy you submitted. boolean or null.false holds. null abstains: pass through, do not block.
not_applicable_reasonWhy a run graded nothing (empty_corpus, not_outreach, evidence_unavailable, …). overall_score is null beside it.Drop refused runs from averages. Branch on the reason.
Quote tierWhose voice a quote is: account (one named company), segment (a cohort), corpus (a recurring pattern), external (the public web).A claim may not reach past its tier. "You told us" needs account.
corpus_onlyThe judge had nothing but corpus patterns to cite.Any account-specific claim in the graded copy was reaching past its evidence.
confidence.levelhigh, moderate or low, with reasons (thin evidence, an errored leg, a truncated prompt).Hold a moderate or low score loosely. It never changes the score.
lift_reportableWhether the before/after difference is larger than the judge's own run-to-run spread.Do not quote lift when this is false.
usage on each artifactas_provided, simulated_specimen, reusable_prompt, illustration_only.Never send an illustration_only message. Keep the reusable_prompt.
evidence_scope, candidate_scopeWhether the evidence and the graded rewrite were pinned from another run.Attribute a difference between two runs only when both say pinned.

What a grade is, and is not

A grade is a structured second opinion on writing, grounded in cited quotes. Each rubric line is a binary pass or fail, so on a five-line rubric the headline score takes exactly six values: 1, 1.8, 2.6, 3.4, 4.2 and 5. It does not predict reply rate or conversion.

One run is one draw. Grading the same text twice can move the score by about two rubric lines, which is why comparing two runs withholds a delta that falls inside that spread. To measure a small change, grade each version several times and compare medians. The current bars and spreads are on Models and limits.

A policy for acting on a result

A starting point for code that acts without a person in the loop. Tighten it for anything with real cost.

SituationDo
Search: empty_reason is read_failed, or detail.internal.status is failedRetry with backoff. Do not report "nothing found".
Search: coverage.basis is unavailable, or truncated is trueUse the rows, but do not state a population, a date range or a recency.
Search: speaker_scope is external_onlyQuote it as customer voice only.
Eval: gate.passed is trueProceed.
Eval: gate.passed is falseHold. Regenerate with the improved prompt from a full run, then grade again.
Eval: gate is null or gate.passed is nullAbstain and pass through. Surface not_applicable_reason to a person.
Eval: confidence.level is low, or corpus_only is trueProceed only with a person reviewing the copy.

The patterns that use these fields end to end are on Patterns.